Search PubMedSearch

SEARCH · Search PubMed

Results for “barcoding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The Use of eDNA Metabarcoding to Detect and Identify Phytophthora in Water Samples.

We describe a protocol to amplify DNA barcodes of known and unknown taxa of Phytophthora and related plant pathogenic oomycetes from a range of environments. The methods focus on sampling pathogen propagules from water using in situ sampling and filtration equipment and buffers that enable efficient storage and DNA extraction for later downstream processing.

Phytophthora

A cell-state axis underlying colonization in carcinomas with implications for metastasis risk prediction and interception.

Metastasis to the liver drives mortality in pancreatic ductal adenocarcinoma (PDAC), yet mechanisms of colonization remain unclear. Using genomic barcoding, we developed a clonal competition model under immune surveillance, isolating murine PDAC subclones with high or low liver-colonization potential. Combined transcriptome and chromatin-accessibility analyses revealed a distinct "metastatic-potential axis," separate from the normal-to-PDAC and classical-basal axes. We established "MetScore" as a biomarker of this axis. MetScore distinguishes metastases from primary PDAC tumors in patients, predicts outcomes beyond classical-basal classifications, and generalizes across carcinoma subtypes, suggesting conserved colonization mechanisms. High-MetScore PDAC cells preferentially occupy immune cell-enriched niches, suggesting they remodel the metastatic microenvironment. Functional screening identified c-Fos as a positive mediator of colonization and a candidate anti-metastatic target. Collectively, we identify a cell-state axis underpinning PDAC liver colonization, introduce MetScore as a broadly applicable biomarker, and nominate actionable targets for peri-operative therapeutic intervention.

Animals

Metschnikowia maris comb. nov., a large-spored yeast species endemic to Serra do Mar Atlantic Rainforest biome, Sao Paulo State, Brazil.

Two yeast isolates from passion flowers were sampled in the southern part of the Serra do Mar Atlantic Rainforest in Sao Paulo State, Brazil. Barcode sequencing and mating experiments showed them to be representatives of Metschnikowia matae var. maris, thus originally named due to the availability of only a single isolate and uncertainties regarding reproductive isolation. The two new isolates being of the complementary mating type to the previously known strain, intravarietal crosses were performed. They yielded a preponderance of two-spored asci, unlike crosses with M. matae var. matae, which led to largely sterile asci. We therefore elevate the variety maris to the rank of species, with the name Metschnikowia maris comb. nov. The holotype is UFMG-CM-Y397T (MATα). Strain UFMG-CM-Y7613A (MAT a) is designated as allotype. The new combination is registered as MB 859665.

Brazil

A set of genetic tools for use in Clostridioides difficile and related species.

The Clostridia are a phylogenetically diverse group of anaerobic, spore-forming bacteria that include species of medical, veterinary and industrial importance. The last two decades have seen major advances in our understanding of Clostridial biology despite the difficulties of anaerobic microbiology and the challenges associated with limited genetic tools. Effort has largely focused on the human pathogen Clostridioides difficile, but many of the methods developed have also proven useful in other species. Here, we present a collection of new genetic tools, including an array of promoters of varying strength, that we have characterized in C. difficile, the food spoilage bacterium Clostridium sporogenes and industrially important Clostridium saccharoperbutylacetonicum. We also present a set of modular plasmids that allow expression of proteins with a variety of tags, including for protein purification and fluorescence microscopy and a method for genetic barcoding of C. difficile to facilitate competitive index experiments. We make these tools available in the hope that they will prove useful to the community in support of our growing understanding of these important bacteria.

Clostridioides difficile

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S

High baseline PD-1+ CD8 T Cells and TIGIT+ CD8 T Cells in circulation associated with response to PD-1 blockade in patients with non-small cell lung cancer.

Blockade of PD-1 or its ligand PD-L1 with antibodies revolutionized treatment for stage III and IV non-small cell lung cancer (NSCLC) since FDA approval in 2015. However, resistance to PD-1/PD-L1 blockade remains a challenge, highlighting the need for biomarkers. This study analyzed 36 stage III and IV NSCLC patients, classified as responders or non-responders by iRECIST criteria. Peripheral blood mononuclear cells collected at baseline and post-treatment were examined for surface and intracellular markers via flow cytometry. CITE sequencing of CD8 T cells from three patients and plasma ctDNA analysis from 13 patients was performed using an ultrasensitive barcoding and next-generation sequencing method. Phenotypic analysis of CD8 T cells revealed higher TIGIT and PD-1 expression at baseline in responders compared to non-responders. Long-term responders (> 21 months) exhibited increased TCF-1+PD-1+ CD8 T cell frequencies relative to shorter-term responders (> 15 months) and non-responders. CITE sequencing revealed intrinsic differences in immune regulation pathways between responders and non-responders. Finally, non-responders showed elevated and increasing ctDNA levels post-treatment, correlating with declining TCF-1+PD-1+ CD8 T cells. Our data suggests combining CD8 T cell analysis with ctDNA dynamics could identify promising biomarkers for monitoring clinical response and treatment efficacy to PD-1/PD-L1 blockade in NSCLC.

Humans

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs = 0.14, Ht = 0.217, I = 0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst = 0.319) and moderate gene flow (Nm = 1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated ΔK using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K = 2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

Integrating hotspot dynamics and centers of diversity: a review of Indo-Australian Archipelago biogeographic evolution and conservation.

The Indo-Australian Archipelago (IAA) is the world's preeminent marine biodiversity hotspot, distinguished by its exceptional species richness in tropical shallow waters. This biodiversity has spurred extensive research into its evolutionary and biogeographic origins. Two prominent theoretical frameworks dominate explanations for the IAA's biodiversity: the "centers-of hypotheses" and the "hopping hotspot hypothesis". The "centers-of hypotheses" posits that specific regions serve as key sources of IAA biodiversity, either through the accumulation and overlap of species from external areas or via elevated rates of local speciation. In contrast, the "hopping hotspot hypothesis" asserts that biodiversity hotspots are dynamic, shifting across geological timescales in response to tectonic and environmental changes. This review synthesizes these contrasting perspectives into an integrated framework, the "Dynamic Centers Hypothesis," which proposes that as biodiversity hotspots migrate over time, the IAA's role in generating and sustaining biodiversity has evolved, with varying contributions from different sources dominating distinct historical phases. By synthesizing the evidence for both hypotheses and incorporating recent findings, including fossil and phylogeography data, we propose the "Dynamic Centers Hypothesis" as a comprehensive and unifying explanation for the IAA's biodiversity. The review further explores biogeographic delineation, aligning tropical marine realms with the IAA's evolutionary trajectory, from its Tethyan roots to its modern Indo-West Pacific dominance. Looking forward, advances in DNA barcoding and genomics are uncovering vast cryptic diversity, revolutionizing our comprehension of IAA phylogeographic history. These discoveries underscore the imperative for a multidimensional conservation framework, integrating phylogenetic, and functional diversity, to preserve this biodiversity hotspot amid escalating global change.

Biogeography

Cloning and validating systems for high throughput molecular recording.

Molecular recording technologies record and store information about cellular history. Lineage tracing is one form of molecular recording and produces information describing cellular trajectories during mammalian development, differentiation and maintenance of adult stem cell niches, and tumor evolution. Our molecular recorder technology utilizes CRISPR-Cas9 barcode editing to generate mutations in genomically integrated, engineered DNA cassettes, which are read out by single-cell RNA sequencing and used to produce high-resolution lineage trees. Here, we describe optimized cloning and validation procedures to construct the molecular recorder lineage tracing system. We include information on considerations of technology design, cloning procedures, the generation of lineage tracing cell lines, and time course experiments to assess their performance.

Cloning, Molecular

Systematic discovery of pathogen effector functions across human pathogens and pathways.

Pathogens deploy effector proteins to exploit host cell biology, and most effector open reading frames (ORFs) are rapidly evolving and lack functional annotation. We developed the effector ORFeome (eORFeome), a scalable functional genomics platform encompassing 3,835 effector ORFs from diverse viruses, bacteria, and parasites. High-throughput barcoded screens across nuclear factor κB (NF-κB), apoptosis, p53, cGAS-STING, and major histocompatibility complex class I (MHC class I) pathways revealed novel pathway-modulating functions for hundreds of uncharacterized eORFs, unexpected activities of known effectors, and distinct pathway-specific functions encoded by single ORFs. Illustrating the power of this approach, we identified HHV6A U14 as a p53 antagonist, HHV7 U21 as a dual-function STING antagonist and MHC-I antigen display inhibitor, and adenoviral 13.6K/i-leader protein as a de novo-evolved TAP inhibitor that suppresses MHC-I display. These results establish a general framework for systematic effector annotation, uncover new mechanisms of host-pathogen interaction across kingdoms, and highlight pathogen effectors as a versatile toolkit for rewiring and probing human cellular pathways.

Humans

Multiome Perturb-seq unlocks scalable discovery of integrated perturbation effects on the transcriptome and epigenome.

Single-cell CRISPR screens link genetic perturbations to transcriptional states, but high-throughput methods connecting these induced changes to their regulatory foundations are limited. Here, we introduce Multiome Perturb-seq, extending single-cell CRISPR screens to simultaneously measure perturbation-induced changes in gene expression and chromatin accessibility. We apply Multiome Perturb-seq in a CRISPRi screen of 13 chromatin remodelers in human RPE-1 cells, achieving efficient assignment of sgRNA identities to single nuclei via an improved method for capturing barcode transcripts from nuclear RNA. We organize expression and accessibility measurements into coherent programs describing the integrated effects of perturbations on cell state, finding that ARID1A and SUZ12 knockdowns induce programs enriched for developmental features. Modeling of perturbation-induced heterogeneity connects accessibility changes to changes in gene expression, highlighting the value of multimodal profiling. Overall, our method provides a scalable and simply implemented system to dissect the regulatory logic underpinning cell state. A record of this paper's transparent peer review process is included in the supplemental information.

Humans

Single-cell glycome and transcriptome profiling enabled by a library of anti-glycan antibodies.

Glycans play critical roles in cellular processes and clinical applications, but they remain difficult to study due to a shortage of well-characterized anti-glycan reagents and high-throughput technologies for glycome profiling, especially ones capable of single-cell resolution. To meet these needs, we generated a database of 650 anti-glycan antibody sequences, recombinantly expressed a library of 154 antibodies, and extensively characterized their binding properties using glycan microarrays. In addition to providing valuable information and resources for the field, the sequence database and microarray data also enabled development of "Glycomic-seq" (Glycome profiling via multiplexed immunoglobulins combined with sequencing), a DNA-barcoded anti-glycan antibody platform that enables high-throughput, single-cell profiling of both RNA and cell-surface glycan expression. Using Glycomic-seq, we profiled two isogenic colorectal cancer cell lines. The results revealed various glycans associated with cancer stem cells and metastasis, demonstrating the power of integrating glycomic information with multi-omic efforts to discover biomarkers and therapeutic targets.

Polysaccharides

Emerging Principles in Spatial Functional Genomics.

Spatial transcriptomic and proteomic atlases have enabled mapping of gene programs within intact tissues, but these measurements remain largely descriptive and do not define the mechanisms controlling tissue biology. Pooled CRISPR screening provides scalable causal interrogation of gene function but remains largely confined to dissociated systems that lack spatial context. In vivo spatial functional genomics (SFG) bridges these approaches by integrating genetic perturbations with in situ transcriptomic and proteomic readouts to measure gene function within intact tissue ecosystems. By preserving spatial organization, SFG enables interpretation of perturbations through effects on cell-cell interactions, diffusible signals, multicellular niches, and tissue architecture. Here, we outline key design axes of SFG: perturbation strategy, barcoding strategy, and phenotypic readout. We discuss computational challenges, including spatial autocorrelation, neighborhood dependence, and context-aware null modeling, and highlight how SFG reveals non-cell-autonomous, architecture-dependent mechanisms of gene function, advancing toward predictive models of tissue organization and gene function.

Genomics

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

Phage Immunoprecipitation and Sequencing-a Versatile Technique for Mapping the Antibody Reactome.

Characterizing the antibody reactome for circulating antibodies provide insight into pathogen exposure, allergies, and autoimmune diseases. This is important for biomarker discovery, clinical diagnosis, and prognosis of disease progression, as well as population-level insights into the immune system. The emerging technology phage display immunoprecipitation and sequencing (PhIP-seq) is a high-throughput method for identifying antigens/epitopes of the antibody reactome. In PhIP-seq, libraries with sequences of defined lengths and overlapping segments are bioinformatically designed using naturally occurring proteins and cloned into phage genomes to be displayed on the surface. These libraries are used in immunoprecipitation experiments of circulating antibodies. This can be done with parallel samples from multiple sources, and the DNA inserts from the bound phages are barcoded and subjected to next-generation sequencing for hit determination. PhIP-seq is a powerful technique for characterizing the antibody reactome that has undergone rapid advances in recent years. In this review, we comprehensively describe the history of PhIP-seq and discuss recent advances in library design and applications.

Humans

ChemPerturb-seq screen identifies a small molecule cocktail enhancing human beta cell survival after subcutaneous transplantation.

Traditional chemical screens have focused on a single assay per screen, making them labor intensive and costly. Here, we combined a chemical screen with single-cell RNA sequencing (scRNA-seq) to perform Chemical Perturb-seq (ChemPerturb-seq), enabling a systematic analysis of the molecular changes of human beta cells upon individual small molecule treatments. Using this platform, we performed an in vivo barcoded screen and discovered a small molecule cocktail, including beta-lipotropin 61-91, insulin growth factor-1, and prostaglandin E2, with which preconditioning human beta cells and primary islets significantly enhanced function and survival when transplanted subcutaneously to female, but not to male, mice. We identified two additional molecules, serotonin and histamine, that promote islet function when transplanted subcutaneously to male mice using ChemPerturb-seq. Such small molecule cocktails could be applied to improve the current FDA-approved islet transplantation procedure. Finally, we developed an artificial intelligence (AI)-powered website, ChemPerturbDB, which provides user-friendly open access analysis of the extensive ChemPerturb-seq dataset.

Humans