Search PubMedSearch

SEARCH · Search PubMed

Results for “genetic barcoding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A set of genetic tools for use in Clostridioides difficile and related species.

The Clostridia are a phylogenetically diverse group of anaerobic, spore-forming bacteria that include species of medical, veterinary and industrial importance. The last two decades have seen major advances in our understanding of Clostridial biology despite the difficulties of anaerobic microbiology and the challenges associated with limited genetic tools. Effort has largely focused on the human pathogen Clostridioides difficile, but many of the methods developed have also proven useful in other species. Here, we present a collection of new genetic tools, including an array of promoters of varying strength, that we have characterized in C. difficile, the food spoilage bacterium Clostridium sporogenes and industrially important Clostridium saccharoperbutylacetonicum. We also present a set of modular plasmids that allow expression of proteins with a variety of tags, including for protein purification and fluorescence microscopy and a method for genetic barcoding of C. difficile to facilitate competitive index experiments. We make these tools available in the hope that they will prove useful to the community in support of our growing understanding of these important bacteria.

Clostridioides difficile

In vivo genome editing of central nervous system SIV reservoirs in ART-suppressed rhesus macaques.

Latent human immunodeficiency virus type 1 (HIV-1) reservoirs in the central nervous system (CNS) may sustain viral persistence and neuroinflammation contributing to HIV-associated neurocognitive disorders (HAND) despite suppressive ART. AAV9-delivered CRISPR has successfully edited SIV proviral DNA in peripheral tissues with acceptable safety profiles, but the extent of in vivo genome editing in the brain remains unclear. Using SIV-infected rhesus macaques, we mapped intact proviral DNA across CNS regions and tested systemic AAV9-CRISPR-Cas9 targeting conserved sites within Ψ packaging signal and Gag region. Ten adult rhesus macaques were infected with genetically barcoded SIVmac239, suppressed with ART, then randomized to receive intravenous AAV9-SaCas9 with dual gRNAs (Ψ + Gag) or a Cas9-only control. At necropsy after viral rebound, SIV genomes were detected in multiple brain regions as well as lymphoid tissues, confirming the CNS as a persistent reservoir during ART. Barcode analysis revealed region-specific patterns consistent with compartmentalized CNS persistence. In CRISPR-treated animals, proviral editing was measurable across anatomically distinct CNS sites. These findings demonstrate that intact and potentially replication-competent virus persists in the primate brain under ART and that systemic AAV9-CRISPR can reach and edit proviral DNA in this sanctuary, supporting genome editing as a strategy toward durable remission of CNS reservoirs.

ART

Paired Single-Cell Transcriptome and DNA Barcode Detection in Zebrafish Using ScarTrace.

ScarTrace is a CRISPR/Cas9-based genetic lineage tracing method that allows for uniquely barcoding the DNA of single cells at a target GFP sequence during developing zebrafish embryos. Single cells from barcoded adult zebrafish can be isolated from various tissues (e.g., marrow, brain, eyes, fins), and their transcriptome and barcode sequences are captured by single-cell cDNA amplification and genomic DNA nested PCR, respectively. Computationally, cell type and barcode identification permit clone tracing and lineage tree reconstruction of tissues to unravel fate decisions during embryogenesis.

Animals

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

A genetic atlas for the butterflies of continental Canada and United States.

Multi-locus genetic data for phylogeographic studies is generally limited in geographic and taxonomic scope as most studies only examine a few related species. The strong adoption of DNA barcoding has generated large datasets of mtDNA COI sequences. This work examines the butterfly fauna of Canada and United States based on 13,236 COI barcode records derived from 619 species. It compiles i) geographic maps depicting the spatial distribution of haplotypes, ii) haplotype networks (minimum spanning trees), and iii) standard indices of genetic diversity such as nucleotide diversity (π), haplotype richness (H), and a measure of spatial genetic structure (GST). High intraspecific genetic diversity and marked spatial structure were observed in the northwestern and southern North America, as well as in proximity to mountain chains. While species generally displayed concordance between genetic diversity and spatial structure, some revealed incongruence between these two metrics. Interestingly, most species falling in this category shared their barcode sequences with one at least other species. Aside from revealing large-scale phylogeographic patterns and shedding light on the processes underlying these patterns, this work also exposed cases of potential synonymy and hybridization.

Animals

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs = 0.14, Ht = 0.217, I = 0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst = 0.319) and moderate gene flow (Nm = 1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated ΔK using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K = 2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

Identification of a robust promoter in mouse and human hepatocytes by in vivo biopanning of a barcoded AAV library.

Recombinant adeno-associated viruses (AAVs) are leading vectors for in vivo human gene therapy. An integral vector element is promoters, which control transgene expression in either a ubiquitous or cell-type-selective manner. Identifying optimal capsid-promoter combinations is challenging, especially when considering on- versus off-target expression. Here, we report a pipeline for in vivo promoter biopanning in AAV building on our AAV capsid barcoding technology and illustrate its potential by screening 53 promoters in 16 murine tissues using an AAV9 vector. Surprisingly, the 2.2-kb human glial fibrillary acidic protein (GFAP) promoter was the top hit in the liver, where it outperformed robust benchmarks such as the human α-1-antitrypsin promoter or the clinically used liver-specific promoter 1 (LP1). Analysis of hepatic cell populations revealed preferred GFAP promoter activity in hepatocytes. Notably, the GFAP promoter also surpassed the LP1 and cytomegalovirus promoters in human hepatocytes engrafted in an immune-deficient mouse. These findings establish the GFAP promoter as an exciting alternative for research and clinical applications requiring efficient and specific transgene expression in hepatocytes. Our pipeline expands the arsenal of technologies for high-throughput in vivo screening of viral vector components and is compatible with capsid barcoding, facilitating the combinatorial interrogation of complex AAV libraries.

Dependovirus

Molecular identification of Hymenopteran insects collected by using Malaise traps from Hazarganji Chiltan National Park Quetta, Pakistan.

The order Hymenoptera holds great significance for humans, particularly in tropical and subtropical regions, due to its role as a pollinator of wild and cultivated flowering plants, parasites of destructive insects and honey producers. Despite this importance, limited attention has been given to the genetic diversity and molecular identification of Hymenopteran insects in most protected areas. This study provides insights into the first DNA barcode of Hymenopteran insects collected from Hazarganji Chiltan National Park (HCNP) and contributes to the global reference library of DNA barcodes. A total of 784 insect specimens were collected using Malaise traps, out of which 538 (68.62%) specimens were morphologically identified as Hymenopteran insects. The highest abundance of species of Hymenoptera (133/538, 24.72%) was observed during August and least in November (16/538, 2.97%). Genomic DNA extraction was performed individually from 90/538 (16.73%) morphologically identified specimens using the standard phenol-chloroform method, which were subjected separately to the PCR for their molecular confirmation via the amplification of cytochrome c oxidase subunit 1 (cox1) gene. The BLAST analyses of obtained sequences showed 91.64% to 100% identities with related sequences and clustered phylogenetically with their corresponding sequences that were reported from Australia, Bulgaria, Canada, Finland, Germany, India, Israel, and Pakistan. Additionally, total of 13 barcode index numbers (BINs) were assigned by Barcode of Life Data Systems (BOLD), out of which 12 were un-unique and one was unique (BOLD: AEU1239) which was assigned for Anthidium punctatum. This indicates the potential geographical variation of Hymenopteran population in HCNP. Further comprehensive studies are needed to molecularly confirm the existing insect species in HCNP and evaluate their impacts on the environment, both as beneficial (for example, pollination, honey producers and natural enemies) and detrimental (for example, venomous stings, crop damage, and pathogens transmission).

Humans

Comparative Analysis of Chloroplast Genomes Reveals Molecular Evolution and Phylogenetic Relationships in Fraxinus (Fraxinus mandshurica).

Fraxinus mandshurica (Manchurian ash) is an ecologically and economically valuable hardwood tree native to Northeast Asia, yet its genomic resources remain limited. We assembled its complete chloroplast (cp) genome (155,559 bp) using hybrid PacBio and Illumina sequencing and performed comparative, phylogenetic, and evolutionary analyses. The cp genome exhibits a typical quadripartite structure encoding 132 gene copies, comprising 114 unique genes (80 protein-coding, 30 tRNA, and 4 rRNA genes), with 18 genes duplicated in the inverted repeat (IR) regions. Simple sequence repeat analysis revealed dominance of mononucleotide A/T repeats. Phylogenetic analysis of 53 complete cp genomes strongly supported the monophyly of Oleaceae and resolved F. mandshurica as sister to the North American F. nigra, consistent with previously proposed Miocene intercontinental dispersal scenarios between East Asia and North America. Most protein-coding genes were under strong purifying selection (Ka/Ks << 1), whereas petB, rpl2, and several ndh genes showed elevated Ka/Ks values that are suggestive of altered selective constraint but are based on very few substitutions and are therefore not, on their own, evidence of positive selection. Nucleotide diversity (Pi) analysis identified 15 hypervariable intergenic spacers (mean Pi = 0.067), among which trnM-CAU-rps14, ndhJ-ndhK, and petL-petG represent promising candidate barcode regions requiring further validation. This study provides a high-quality, fully annotated cp genome of F. mandshurica and a valuable genomic resource for future phylogenetic, population genetic, and conservation studies of this important genus.

Fraxinus

Identification and full genome sequencing of previously unknown sandfly-borne phleboviruses using a newly established capture-based next-generation sequencing approach.

Sandfly-borne phleboviruses cause febrile illness and neuroinvasive disease in humans. While infections are reported in the Mediterranean region, the discovery of previously unknown phleboviruses in sandflies from Kenya suggests a wider geographic distribution. Detection and characterization of novel phleboviruses are often hindered by low-quality and low-viral-load samples. We developed a capture-based target enrichment next-generation sequencing approach that showed a 99%-100% fold enrichment of viral genomes from primary material and provides a robust tool for generating complete genomes of both known and previously unknown viruses. From a collection of 15,652 sandflies in Kenya, we recovered seven complete coding sequences of Embossos, Bogoria, and Kiborgoch viruses, and of two previously unknown phleboviruses, which were named Sosoik and Shable viruses. Sosoik virus shared 83% amino acid identity in its RdRp gene with that of Bogoria virus, while Shable virus shared ca. 88% amino acid identity with viruses of the Salehabad serocomplex. Additionally, a reassortant of Shable virus was detected that possessed an M segment from an undescribed Ponticelli-like virus. DNA barcoding of blood-fed sandflies revealed several potentially novel Sergentomyia species and evidence of host-feeding on humans, livestock, and reptiles, suggesting possibilities for zoonotic transmission. Overall, our findings increase the known genetic diversity of Old World sandfly-borne phlebovirus species from 18 to 25 (by 38.9%), including the detection of viruses from all pathogenic sandfly-borne phlebovirus serocomplexes in East Africa, opening new horizons in disease ecology research.IMPORTANCEKnowledge of the genetic diversity of circulating pathogens is crucial for providing appropriate diagnostics and disease management. This study established a novel capture-based target enrichment next-generation sequencing approach that enabled the near-complete viral genome recovery from primary samples, while native NGS yielded negative or poor-quality results. In addition to the five recently discovered sandfly-borne phleboviruses in Kenya, two previously unknown phleboviruses were detected in sandflies from the same region. The viruses were detected in several sandfly species, which showed diverse host-feeding behaviors, including mixed feeding on humans and chickens. The study significantly advances the understanding of sandfly-borne phleboviruses by uncovering their broader geographic distribution and genetic diversity, particularly in East Africa, highlighting the importance of expanding surveillance efforts beyond traditionally studied regions.

Phlebovirus

Dual plasmepsin IX and X inhibitors are refractory to development of resistance.

Artemisinin-based combination therapies (ACTs) remain the cornerstone of malaria treatment, but emerging resistance threatens their efficacy. The potential for the development of drug resistance against plasmepsin X (PMX)-selective inhibitors and dual plasmepsin IX/X (PMIX/X) inhibitors was investigated in Plasmodium falciparum. A series of PMX-selective (WM4, WM76, WM92) and PMIX/X dual inhibitors (WM382, WM09, WM42) were characterised for potency against parasite growth and enzyme inhibition. In vitro selection experiments showed that all compounds had a high barrier to resistance, although parasites with reduced sensitivity to PMX&#x2011;selective inhibitors could still be selected. Resistance mechanisms involved pmx gene amplification and point mutations (D245N, S315P, S359P, I363L) that alter inhibitor binding. Recombinant expression and Michaelis-Menten kinetics demonstrated that these mutations impair drug binding whilst preserving PMX catalytic function. Reverse genetics confirmed that introducing these mutations into the pmx gene resulted in decreased potency of the inhibitors. In this study, resistance to the PMIX/X dual inhibitors evaluated here could not be selected, despite prolonged selection pressure. Antimalarial Resistome Barcoding (AReBar) assays confirmed the absence of pre-existing resistance to either inhibitor class. Critically, PMIX/X dual inhibitors maintained efficacy against parasites with decreased sensitivity to PMX-selective compounds. These findings demonstrate that dual PMIX/X inhibitors present a substantially higher barrier to resistance than PMX-selective inhibitors, informing antimalarial drug development strategies and highlighting dual-target inhibition as a promising approach to mitigate resistance risks.

Aspartic Acid Endopeptidases

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis

Environmental Release of Genetically Intervened Microorganisms: Towards a New Narrative.

The deliberate release of genetically engineered microorganisms for environmental applications has remained largely blocked since the early days of recombinant DNA technology, when limited ecological knowledge, lack of success stories and public apprehension shaped a culture of caution and restrictive regulation. Despite profound advances in microbial ecology, synthetic biology and genetic design, current frameworks still rely on outdated assumptions and legacy regulations that equate engineered microbes with inherent danger and demand unrealistic forms of absolute containment. This review examines how laboratory-trained microorganisms exist on a continuum with naturally evolved life, and that their risks are neither categorically different nor greater. Rather than pursuing unachievable containment, governance should shift towards traceability, stewardship and long-term monitoring through genomic barcodes, digital twins and transparent oversight. The vision moves from domination and control to care and partnership recognizing engineered microbes as live amendments capable of restoring degraded ecosystems. Achieving this transformation requires new terminology, phased field-trial frameworks, improved scaling methods, and the integration of epistemological perspectives that emphasize reciprocity and coexistence with nature. Reframing biotechnology in this way could finally unlock the capacity of engineered microorganisms to contribute responsibly and effectively to planetary repair in an era of escalating environmental crises.

Microorganisms, Genetically-Modified

CROPseq-multi: a universal solution for multiplexed perturbation in high-content pooled CRISPR screens.

Forward genetic screens seek to dissect complex biological systems by systematically perturbing genetic elements and observing the resulting phenotypes. While standard screening methodologies introduce individual perturbations, multiplexing perturbations improves the performance of single-target screens and enables combinatorial screens for the study of genetic interactions. Current tools for multiplexing perturbations are limited by technical challenges and do not offer compatibility across diverse screening methodologies, including enrichment, single-cell sequencing, and optical pooled screens. Here, we report the development of CROPseq-multi (CSM), a CROPseq1-inspired lentiviral system to multiplex Streptococcus pyogenes (Sp) Cas9-based perturbations with versatile readout compatibility and high performance for both perturbation and barcode identification. CSM has equivalent per-guide activity to CROPseq and low lentiviral recombination frequencies. Dual-guide CSM libraries are constructed in a single, facile molecular cloning step that facilitates the use of unique molecular identifiers. CSM is compatible with enrichment screening methodologies, single-cell RNA-sequencing readouts, and optical pooled screens. For optical pooled screens, an optimized and multiplexed in situ detection protocol improves barcode counts 10-fold (for mRNA detection), enables detection of recombination events, and reduces the number of sequencing cycles required for decoding by 3-fold relative to CROPseq. CROPseq-multi-v2 (CSMv2) adds compatibility for detection methods based on T7 RNA polymerase in vitro transcription2-5. CSM provides a single system for CRISPR screens that is compatible with individual and combinatorial perturbations, diverse SpCas9-based perturbation technologies, and multiple high-content, single-cell phenotypic readouts.

CRISPR Cas9

Directed evolution of engineered virus-like particles with improved production and transduction efficiencies.

Engineered virus-like particles (eVLPs) are promising vehicles for transient delivery of proteins and RNAs, including gene editing agents. We report a system for the laboratory evolution of eVLPs that enables the discovery of eVLP variants with improved properties. The system uses barcoded guide RNAs loaded within DNA-free eVLP-packaged cargos to uniquely label each eVLP variant in a library, enabling the identification of desired variants following selections for desired properties. We applied this system to mutate and select eVLP capsids with improved eVLP production properties or transduction efficiencies in human cells. By combining beneficial capsid mutations, we developed fifth-generation (v5) eVLPs, which exhibit a 2-4-fold increase in cultured mammalian cell delivery potency compared to previous-best v4 eVLPs. Analyses of v5 eVLPs suggest that these capsid mutations optimize packaging and delivery of desired ribonucleoprotein cargos rather than native viral genomes and substantially alter eVLP capsid structure. These findings suggest the potential of barcoded eVLP evolution to support the development of improved eVLPs.

Humans

Emerging Principles in Spatial Functional Genomics.

Spatial transcriptomic and proteomic atlases have enabled mapping of gene programs within intact tissues, but these measurements remain largely descriptive and do not define the mechanisms controlling tissue biology. Pooled CRISPR screening provides scalable causal interrogation of gene function but remains largely confined to dissociated systems that lack spatial context. In vivo spatial functional genomics (SFG) bridges these approaches by integrating genetic perturbations with in situ transcriptomic and proteomic readouts to measure gene function within intact tissue ecosystems. By preserving spatial organization, SFG enables interpretation of perturbations through effects on cell-cell interactions, diffusible signals, multicellular niches, and tissue architecture. Here, we outline key design axes of SFG: perturbation strategy, barcoding strategy, and phenotypic readout. We discuss computational challenges, including spatial autocorrelation, neighborhood dependence, and context-aware null modeling, and highlight how SFG reveals non-cell-autonomous, architecture-dependent mechanisms of gene function, advancing toward predictive models of tissue organization and gene function.

Genomics

Efficient and multiplexed somatic genome editing with Cas12a mice.

Somatic genome editing in mouse models has increased our understanding of the in vivo effects of genetic alterations. However, existing models have a limited ability to create multiple targeted edits, hindering our understanding of complex genetic interactions. Here we generate transgenic mice with Cre-regulated and constitutive expression of enhanced Acidaminococcus sp. Cas12a (enAsCas12a), which robustly generates compound genotypes, including diverse cancers driven by inactivation of trios of tumour suppressor genes or an oncogenic translocation. We integrate these modular CRISPR RNA (crRNA) arrays with clonal barcoding to quantify the size and number of tumours with each array, as well as the impact of varying the guide number and position within a four-guide array. Finally, we generate tumours with inactivation of all combinations of nine tumour suppressor genes and find that the fitness of triple-knockout genotypes is largely explainable by one- and two-gene effects. These Cas12a alleles will enable further rapid creation of disease models and high-throughput investigation of coincident genomic alterations in vivo.

Animals