Search PubMedSearch

SEARCH · Search PubMed

Results for “genetic barcoding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

30 records · Page 2Linked to original sources

Whole-genome sequences reveal zygotic composition in chimeric twins.

While most dizygotic twins have a dichorionic placenta, rare cases of dizygotic twins with a monochorionic placenta have been reported. The monochorionic placenta in dizygotic twins allows in utero exchange of embryonic cells, resulting in chimerism in the twins. In practice, this chimerism is incidentally identified in mixed ABO blood types or in the presence of cells with a discordant sex chromosome. Here, we applied whole-genome sequencing to one triplet and one twin family to precisely understand their zygotic compositions, using millions of genomic variants as barcodes of zygotic origins. Peripheral blood showed asymmetrical contributions from two sister zygotes, where one of the zygotes was the major clone in both twins. Single-cell RNA sequencing of peripheral blood tissues further showed differential contributions from the two sister zygotes across blood cell types. In contrast, buccal tissues were pure in genetic composition, suggesting that in utero cellular exchanges were confined to the blood tissues. Our study illustrates the cellular history of twinning during human development, which is critical for managing the health of chimeric individuals in the era of genomic medicine.

Humans

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software

Modeling homologous chromosome recognition via nonspecific interactions.

In many organisms, most notably Drosophila, homologous chromosomes associate in somatic cells, a phenomenon known as somatic pairing, which takes place without double strand breaks or strand invasion, thus requiring some other mechanism for homologs to recognize each other. Several studies have suggested a "specific button" model, in which a series of distinct regions in the genome, known as buttons, can associate with each other, mediated by different proteins that bind to these different regions. Here, we use computational modeling to evaluate an alternative "button barcode" model, in which there is only one type of recognition site or adhesion button, present in many copies in the genome, each of which can associate with any of the others with equal affinity. In this model, buttons are nonuniformly distributed, such that alignment of a chromosome with its correct homolog, compared with a nonhomolog, is energetically favored; since to achieve nonhomologous alignment, chromosomes would be required to mechanically deform in order to bring their buttons into mutual register. By simulating randomly generated nonuniform button distributions, many highly effective button barcodes can be easily found, some of which achieve virtually perfect pairing fidelity. This model is consistent with existing literature on the effect of translocations of different sizes on homolog pairing. We conclude that a button barcode model can attain highly specific homolog recognition, comparable to that seen in actual cells undergoing somatic homolog pairing, without the need for specific interactions. This model may have implications for how meiotic pairing is achieved.

Animals

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180 bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve > 20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total = 84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans

Multiome Perturb-seq unlocks scalable discovery of integrated perturbation effects on the transcriptome and epigenome.

Single-cell CRISPR screens link genetic perturbations to transcriptional states, but high-throughput methods connecting these induced changes to their regulatory foundations are limited. Here, we introduce Multiome Perturb-seq, extending single-cell CRISPR screens to simultaneously measure perturbation-induced changes in gene expression and chromatin accessibility. We apply Multiome Perturb-seq in a CRISPRi screen of 13 chromatin remodelers in human RPE-1 cells, achieving efficient assignment of sgRNA identities to single nuclei via an improved method for capturing barcode transcripts from nuclear RNA. We organize expression and accessibility measurements into coherent programs describing the integrated effects of perturbations on cell state, finding that ARID1A and SUZ12 knockdowns induce programs enriched for developmental features. Modeling of perturbation-induced heterogeneity connects accessibility changes to changes in gene expression, highlighting the value of multimodal profiling. Overall, our method provides a scalable and simply implemented system to dissect the regulatory logic underpinning cell state. A record of this paper's transparent peer review process is included in the supplemental information.

Humans

DNA Extraction Optimisation for Minute Land Snails of Vertigo Müller, 1773 (Gastropoda: Vertiginidae): A Comparative Evaluation of Six Methods, Including a Non-Destructive Shell-Preserving Protocol.

No systematic comparison of DNA extraction strategies exists for minute Vertiginidae (shell height <&#x2009;3&#x2009;mm), a group posing a dual analytical challenge: extremely low tissue input and co-purified PCR-inhibitory mucus. For legally protected species, an additional requirement to preserve the shell voucher further constrains available protocols. Using Vertigo antivertigo as the model species, we compared six approaches applied to specimens preserved in 96% ethanol (n&#x2009;=&#x2009;10 per method): two HotSHOT alkaline-lysis protocols (destructive and non-destructive shell-preserving variants), a modified CTAB protocol supplemented with PVP-40 and DTT, and three commercial silica-column kits (GeneJET Genomic, DNeasy Blood & Tissue, QIAamp DNA Micro). DNA yields were quantified by QuantiFluor fluorometry, and PCR performance was subsequently assessed across four loci (COI barcode, COI mini-barcode, ITS1, ITS2). DNeasy Blood & Tissue produced the highest fluorometric concentrations; QIAamp DNA Micro and CTAB&#x2009;+&#x2009;PVP-40 gave intermediate values. The shell-preserving HotSHOT variant yielded lower concentrations but improved A260/230 ratios. BSA and trehalose supplementation increased PCR success in inhibition-prone HotSHOT extracts from 70% to 100%. ITS1 Sanger sequencing of three Vertigo species listed in Annex II of the EU Habitats Directive, all extracted with the shell-preserving protocol, confirmed species-level identification (99.8%-100% BLASTn identity; mean Phred Q&#x2009;>&#x2009;51). The shell-preserving non-destructive HotSHOT protocol yields sequenceable DNA from protected Vertiginidae while retaining the morphological voucher, making it the preferred option for conservation-genetic monitoring. The practical decision framework documented here-integrating voucher preservation, amplification robustness and per-sample cost-has broad applicability to other minute terrestrial gastropods processed in large-scale biodiversity surveys.

Habitats Directive

Integrative multiomic approaches reveal ZMAT3 and p21 as conserved hubs in the p53 tumor suppression network.

TP53, the most frequently mutated gene in human cancer, encodes a transcriptional activator that induces myriad downstream target genes. Despite the importance of p53 in tumor suppression, the specific p53 target genes important for tumor suppression remain unclear. Recent studies have identified the p53-inducible gene Zmat3 as a critical effector of tumor suppression, but many questions remain regarding its p53-dependence, activity across contexts, and mechanism of tumor suppression alone and in cooperation with other p53-inducible genes. To address these questions, we used Tuba-seqUltra somatic genome editing and tumor barcoding in a mouse lung adenocarcinoma model, combinatorial in vivo CRISPR/Cas9 screens, meta-analyses of gene expression and Cancer Dependency Map data, and integrative RNA-sequencing and shotgun proteomic analyses. We established Zmat3 as a core component of p53-mediated tumor suppression and identified Cdkn1a as the most potent cooperating p53-induced gene in tumor suppression. We discovered that ZMAT3/CDKN1A serve as near-universal effectors of p53-mediated tumor suppression that regulate cell division, migration, and extracellular matrix organization. Accordingly, combined Zmat3-Cdkn1a inactivation dramatically enhanced cell proliferation and migration compared to controls, akin to p53 inactivation. Together, our findings place ZMAT3 and CDKN1A as hubs of a p53-induced gene program that opposes tumorigenesis across various cellular and genetic contexts.

Animals

A rapid CRISPR-based nanodroplet assay enables direct clinical identification of mycobacteria species.

The global incidence and mortality of nontuberculous mycobacterial infections have risen sharply with population aging. In some regions, they are now surpassing Mycobacterium tuberculosis complex infections, imposing a substantial clinical and economic burden. Because nontuberous mycobacteria exhibit species-level heterogeneity and require prolonged culture for identification, their diagnosis remains slow and is frequently inaccurate. Here, we describe a multiplexed clustered regularly interspaced short palindromic repeats (CRISPR)-assisted nanodroplet differential identification (CANDI) diagnostic platform that integrates species-agnostic target amplification with species-specific CRISPR-associated protein 12a (Cas12a) detection in fluorescence-barcoded nanodroplets. By spatially compartmentalizing CRISPR reactions into color-encoded nanodroplets, CANDI overcomes the multiplexing limitations of conventional CRISPR diagnostics and enables simultaneous interrogation of multiple mycobacterial targets in a single assay. We designed a 16-plex panel that distinguishes 15 clinically relevant Mycobacterium species and subspecies. CANDI achieved high analytical sensitivity and accurate discrimination in samples containing coinfections with multiple species or subspecies. When applied to 230 clinical specimens, including sputum, tracheal aspirates, and other respiratory fluids, CANDI delivered subspecies-level results within 3.5 hours, achieving 97.08% sensitivity and 99.7% specificity relative to culture-based identification. By combining multiplexed, high-specificity CRISPR detection with scalable droplet-based engineering, CANDI has the potential to overcome the culture dependency of current diagnostics and enable species- and subspecies-level identification across the genetically complex Mycobacterium genus, offering a clinically adaptable framework for rapid, precision diagnosis of mycobacterial infections.

Humans

Fluctuating DNA methylation tracks cancer evolution at clinical scale.

Cancer development and response to treatment are evolutionary processes1,2, but characterizing evolutionary dynamics at a clinically meaningful scale has remained challenging3. Here we develop a new methodology called EVOFLUx, based on natural DNA methylation barcodes fluctuating over time4, that quantitatively infers evolutionary dynamics using only a bulk tumour methylation profile as input. We apply EVOFLUx to 1,976 well-characterized lymphoid cancer samples spanning a broad spectrum of diseases and show that initial tumour growth rate, malignancy age and epimutation rates vary by orders of magnitude across disease types. We measure that subclonal selection occurs only infrequently within bulk samples and detect occasional examples of multiple independent primary tumours. Clinically, we observe faster initial tumour growth in more aggressive disease subtypes, and that evolutionary histories are strong independent prognostic factors in two series of chronic lymphocytic leukaemia. Using EVOFLUx for phylogenetic analyses of aggressive Richter-transformed chronic lymphocytic leukaemia samples detected that the seed of the transformed clone existed decades before presentation. Orthogonal verification of EVOFLUx inferences is provided using additional genetic data, including long-read nanopore sequencing, and clinical variables. Collectively, we show how widely available, low-cost bulk DNA methylation data precisely measure cancer evolutionary dynamics, and provides new insights into cancer biology and clinical behaviour.

Humans

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

Deep learning guided programmable design of Escherichia coli core promoters from sequence architecture to strength control.

Core promoters are essential regulatory elements that control transcription initiation, but accurately predicting and designing their strength remains challenging due to complex sequence-function relationships and the limited generalizability of existing AI-based approaches. To address this, we developed a modular platform integrating rational library design, predictive modelling, and generative optimization into a closed-loop workflow for end-to-end core promoter engineering. Conserved and spacer region of core promoters exert distinct effects on transcriptional strength, with the former driving large-scale variation and the latter enabling finer gradation. Based on this insight, Mutation-Barcoding-Reverse Sequencing approach was used and constructed a synthetic promoter library comprising 112&#xa0;955 variants with minimal redundancy and a 16&#xa0;226-fold expression range. A Transformer-based model trained on this dataset achieved a Pearson correlation of 0.87 with experimentally measured promoter strengths. When combined with a conditional diffusion model, the system enabled de novo generation of promoter sequences with defined strengths, achieving a design-to-measurement correlation of 0.95 and maintaining high accuracy (R&#xa0;=&#xa0;0.93) across varied sequence contexts. The designed promoters consistently preserved their intended strength gradients, demonstrating robust plug-and-play functionality. This work establishes a scalable and extensible platform (www.yudenglab.com) for deep learning-guided programmable design of Escherichia&#xa0;coli core promoters, enabling precise transcriptional control.

Promoter Regions, Genetic

Getting to the Core of the Matter-Assessing the Role of Replication in Metabarcoding-Based sedaDNA.

Replication is central to most experimental and sampling designs, increasing inferential power and capturing fine-scale data heterogeneity. However, its importance remains poorly evaluated in some ecological and evolutionary settings. This is the case of metabarcoding studies using DNA recovered from sedimentary archives, in which biological signals integrate ecological information through depositional and burial processes, yet are commonly inferred from a single sediment core per site. Here, we evaluated the effect of different types of replication using sedimentary DNA metabarcoding data from two genetic markers (mitochondrial COI and nuclear 18S) using a nested sampling design. The design included three intertidal sites, three spatially separated sediment cores per site (biological replicates), two sediment horizons per core, and eight PCR (technical) replicates per sediment sample. Variance partitioning showed that site identity and sediment age group together explained >&#x2009;70% of the variation in beta diversity, indicating that among-site spatial and stratigraphic differences were the dominant drivers of community composition. PERMANOVA likewise identified non-significant effects of biological replication. Among PCR replicates from the same sediment sample, richness varied substantially, whereas Shannon diversity was more consistent. Despite this variability, differences in community composition among technical replicates remained smaller than those associated with biological replication or site identity, indicating a limited influence on broader ecological patterns. Community composition was highly similar among replicate cores within sites, consistent with stratigraphic coherence. These results indicate limited within-site heterogeneity and suggest that, under stratigraphically coherent conditions, increasing biological replication may provide little additional information, whereas enhancing technical replication and stratigraphic resolution can improve ecological inference from sedimentary DNA metabarcoding datasets.

DNA Barcoding, Taxonomic