Search PubMedSearch

SEARCH · Search PubMed

Results for “DNA analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Microbial DNA analysis of paired blood-bronchoalveolar lavage fluid in post-HSCT patients with pneumonia implying application conditions of blood as a surrogate in pathogen detection.

BACKGROUND: Blood testing aids pneumonia diagnosis, but its effectiveness varies. Given the invasiveness of bronchoalveolar lavage fluid (BALF) sampling versus blood testing's simplicity, this study investigates when blood can reliably substitute for BALF in detecting microbial presence, especially for pathogens. RESULTS: Metagenomic sequencing was performed on paired BALF-blood samples from 21 post-HSCT immunocompromised (ICP) and 21 immunocompetent (ICT) patients. The ICP cohort was expanded to 62 for biomarker validation. Host responses were profiled via metatranscriptomics (30 BALF samples). Microbial alpha and beta diversity differed significantly between blood and BALF in ICP, but not ICT, patients. ICP patients' BALF contained a greater diversity and abundance of microbes. A higher proportion of microbial DNA sequences in ICP patients' blood was also present in their BALF, suggesting a potentially more permeable alveolar-capillary barrier. Related genes (e.g., NABA CORE MATRISOME, extracellular matrix organization, cell-cell adhesion) were downregulated. Upregulated pathways like VEGFA-VEGFR2 signaling and Rho GTPases suggested increased vascular permeability. In ICP patients, 419 microbial sequences in blood indicated their presence in the lower respiratory tract with > 70% certainty. CONCLUSION: Host immune status significantly influences blood-BALF microbial diversity differences. Shared blood-BALF microbial DNA sequences show potential for aiding pneumonia pathogen diagnosis, offering a novel biomarker identification approach.

Humans

Private detection of relatives in forensic genomics using homomorphic encryption.

BACKGROUND: Forensic analysis heavily relies on DNA analysis techniques, notably autosomal Single Nucleotide Polymorphisms (SNPs), to expedite the identification of unknown suspects through genomic database searches. However, the uniqueness of an individual's genome sequence designates it as Personal Identifiable Information (PII), subjecting it to stringent privacy regulations that can impede data access and analysis, as well as restrict the parties allowed to handle the data. Homomorphic Encryption (HE) emerges as a promising solution, enabling the execution of complex functions on encrypted data without the need for decryption. HE not only permits the processing of PII as soon as it is collected and encrypted, such as at a crime scene, but also expands the potential for data processing by multiple entities and artificial intelligence services. METHODS: This study introduces HE-based privacy-preserving methods for SNP DNA analysis, offering a means to compute kinship scores for a set of genome queries while meticulously preserving data privacy. We present three distinct approaches, including one unsupervised and two supervised methods, all of which demonstrated exceptional performance in the iDASH 2023 Track 1 competition. RESULTS: Our HE-based methods can rapidly predict 400 kinship scores from an encrypted database containing 2000 entries within seconds, capitalizing on advanced technologies like Intel AVX vector extensions, Intel HEXL, and Microsoft SEAL HE libraries. Crucially, all three methods achieve remarkable accuracy levels (ranging from 96% to 100%), as evaluated by the auROC score metric, while maintaining robust 128-bit security. These findings underscore the transformative potential of HE in both safeguarding genomic data privacy and streamlining precise DNA analysis. CONCLUSIONS: Results demonstrate that HE-based solutions can be computationally practical to protect genomic privacy during screening of candidate matches for further genealogy analysis in Forensic Genetic Genealogy (FGG).

Humans

MetaChrome: An Open-Source, User-Friendly Tool for Automated Metaphase Chromosome Analysis.

DNA Fluorescence In Situ Hybridization (FISH) is an essential technique to study chromosome biology and genetics, enabling precise visualization of specific genomic loci to study structural abnormalities, gene mapping, and chromosomal rearrangements. High-Throughput Imaging (HTI) can automate the analysis of DNA-FISH chromosome images, but the accurate and automated segmentation of mitotic chromosomes and simultaneous colocalization of FISH signals remains a challenge. While several commercial automated karyotyping tools partially solve these issues, open-source software that effectively combines robust chromosome segmentation with comprehensive colocalization analysis capabilities remains necessary. To address this unmet need, we developed MetaChrome, an open-source software platform built around a graphical user interface and explicitly designed for automated metaphase chromosome analysis. MetaChrome leverages fine-tuned deep learning models to automate metaphase chromosome segmentation, together with colocalization analysis of chromosome-specific FISH probes and immunofluorescent-labeled proteins. Importantly, MetaChrome achieves enhanced segmentation accuracy compared to traditional image processing methods by adopting a Cellpose segmentation model fine-tuned with manually annotated metaphase chromosome datasets. The fine-tuned model ensures precise assignment of DNA-FISH spots to individual chromosomes in an automated manner. This facilitates rapid identification of chromosomal abnormalities, reduces human error, and advances high-throughput chromosome analysis workflows, addressing a key bottleneck in chromosome biology research.

Chromosome segmentation

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

DNA Methylation Analysis by Bisulfite Pyrosequencing of Mouse Embryonic Fibroblasts with Reprogramming Enhanced by Thyroid Hormones.

DNA methylation is a widely studied epigenetic mark which in mammals involves the incorporation of a methyl group to the fifth carbon of cytosines, mainly those belonging to CpG dinucleotides. It has been linked to context-dependent regulatory functions ranging from gene and repetitive DNA silencing to gene body transcriptional activity. Because of its important roles during embryonic development and cell differentiation, DNA methylation can be used to track cell reprogramming by measuring the methylation levels of pluripotency-associated factors. In this scenario, bisulfite pyrosequencing is a simple, robust, and widely used technique which allows for the quantification of DNA methylation levels at small, specific regions of the genome. It involves the amplification and biotin tagging of bisulfite-converted DNA. Single amplified strands are then purified using streptavidin and finally pyrosequenced using a sequencing primer. Thus, it is an ideal method for the quantitative profiling of specific genomic regions, with applications ranging from biomarker discovery and epigenetic clock tracking to omic validation studies.

Animals

Extraction, Purification, and Next-Generation Sequencing (NGS) Analysis of DNA and RNA from Formalin-Fixed and Paraffin-Embedded (FFPE) Tissue.

Formalin fixed paraffin embedded (FFPE) tissues have long been used for immunohistological analyses. FFPE tissues can be stored at room temperature for several years enabling analyses to be performed later. Ease of storage and transport makes these tissues an attractive source of biological material. However, formalin fixation results in chemical modifications of proteins and nucleic acids that poses a major challenge to any type of analysis. Recovery of nucleic acids for quantitative assays is rendered difficult due to degradation resulting from fixation and long-term storage, producing low usable yields. Extensive efforts in the last 20 years have led to significant improvements in use of FFPE tissues for DNA and RNA analyses and resulted in development of sensitive assays for a wide range of applications, including next-generation sequencing. In this chapter, we describe the optimization of methods for sequential extraction of DNA and RNA from FFPE tissue and subsequent preparation of DNA-seq and RNA-seq libraries for use with the Illumina platform using commercially available reagents/kits.

Paraffin Embedding

PathwayVote: an R package for robust pathway enrichment analysis for DNA methylation data using a consensus-based voting framework.

MOTIVATION: Pathway enrichment analysis is commonly used to interpret epigenomewide association studies, yet conventional methods often rely on arbitrary thresholds and simplified CpG-gene mappings, making them sensitive to analytical choices and unable to fully leverage CpG-gene relationships Recent advances in expression quantitative trait methylation (eQTM) studies offer a rich resource to refine these mappings, but are rarely utilized in DNA methylation enrichment pipelines. RESULTS: We developed PathwayVote, an R package that implements a voting-based consensus approach and leverages eQTM data to identify robustly enriched pathways. PathwayVote reduces dependence on arbitrary cutoffs and improves sensitivity and reproducibility of enrichment results. AVAILABILITY AND IMPLEMENTATION: PathwayVote is freely available on GitHub (https://github.com/YinanZheng/PathwayVote) under the GPL-3 license and CRAN: https://CRAN.R-project.org/package=PathwayVote. The version of the code corresponding to this manuscript has been archived on Zenodo (https://doi.org/10.5281/zenodo.17209507).

Humans

clusIBD: Robust Detection of Identity-by-descent Segments Using Unphased Genetic Data from Poor-quality Samples.

The detection of identity-by-descent (IBD) segments is widely used to infer relatedness in many fields, including forensics and ancient DNA analysis. However, existing methods are often ineffective for poor-quality DNA samples. Here, we propose a method, clusIBD, which can robustly detect IBD segments using unphased genetic data with a high rate of genotyping error. We evaluated and compared the performance of clusIBD with that of IBIS, TRUFFLE, and IBDseq using simulated data, artificial poor-quality materials, and ancient DNA samples. The results show that clusIBD outperforms these existing tools and could be used for kinship inference in fields such as ancient DNA analysis and criminal investigation. clusIBD is publicly available at GitHub (https://github.com/Ryan620/clusIBD/) and BioCode (https://ngdc.cncb.ac.cn/biocode/tool/BT007882).

Humans

Genetic Identification of Burned Human Remains: A Systematic Review.

Background/Objectives: DNA-based identification of degraded human remains represents a major challenge in forensic science, particularly in cases involving burned, fragmented, or commingled bodies. Advances in forensic genetics have expanded the analytical capabilities for such samples; however, the effectiveness of different approaches and their integration within Disaster Victim Identification (DVI) workflows remain heterogeneous. This systematic review aims to critically evaluate current evidence on DNA-based identification of degraded remains, focusing on methodological strategies, emerging genomic technologies, and DVI applications, while integrating laboratory evidence and operational forensic practice into a structured analytical framework. Methods: A systematic literature search was conducted in Scopus and Web of Science from database inception to 5 June 2026, following PRISMA 2020 guidelines. Eligible studies included original research addressing DNA analysis of degraded, thermally altered, or highly compromised human remains in forensic or DVI contexts. After a multistep screening process involving title/abstract and full-text evaluation, 37 studies were included. Data were extracted and organized into three thematic categories: (i) core DNA analysis, (ii) advanced molecular technologies, and (iii) DVI case applications. Results: The findings demonstrate that DNA recovery from degraded remains is influenced by thermal exposure, tissue type, and sampling strategy. Teeth and dense cortical bone consistently provide higher DNA yield. While autosomal STR profiling remains the primary analytical approach, its limitations in highly degraded samples are mitigated through the complementary use of mitochondrial DNA (mtDNA), Y-chromosome STRs (Y-STRs), and SNP markers, together with advanced sequencing technologies such as massively parallel sequencing (MPS). Emerging technologies, including rapid DNA systems and predictive models based on macroscopic indicators, significantly enhance efficiency and success rates. DVI studies report identification rates exceeding 90-95% when multidisciplinary and structured workflows are applied. The evidence further supports a flexible triage-based analytical strategy, in which marker selection is guided by tissue preservation and degradation level. Conclusions: DNA-based identification of degraded human remains has evolved into an adaptive, multi-level forensic process. Successful outcomes rely on the integration of optimized sampling, hierarchical genetic analysis, and coordinated DVI strategies. The findings support a triage-based framework that links tissue selection, degradation assessment, and analytical methodology to maximize identification success. Future developments should focus on predictive models, advanced genomic tools, and standardized workflows to further improve identification in challenging forensic scenarios.

Humans

Comprehensive analysis of DNA methylome and transcriptome reveals the epigenetic regulation of nitric oxide treatment in delaying apricot fruit senescence.

Apricot produces climacteric fruit, which are perishable after harvest. To elucidate the regulatory role of NO treatment through DNA methylation in post-harvest senescence, apricot fruits were treated with 0.2 mmol/L sodium nitroprusside (SNP) solution for 10 min, with distilled water treatment serving as the control. Treated fruits were then stored at 25°C and 80% relative humidity. Changes in appearance quality, physiological parameters, metabolome profiles, transcriptome dynamics, and DNA methylation patterns were analyzed before and after storage. Results showed that NO treatment delayed apricot softening, increased flavonoid metabolite accumulation, and reduced lipid and abscisic acid accumulation, with these effects correlated to the expression of specific genes and transcription factors. This work reveals the epigenetic regulatory mechanism underlying NO treatment delaying ripening and senescence. Further analysis revealed that the transcription levels of ACO, PAL, UFGT-like, NCED1, PP2C, MYB21, CCoAOMT-like, CYP707A, and ZNF7-like were all correlated with DNA methylation. This indicates that SNP treatment can lead to large changes in DNA methylation levels in apricot fruits, and that the differences in gene transcription levels are associated with the occurrence of hypomethylation and hypermethylation. Collectively, these findings establish an epigenetic framework for post-harvest regulation of apricot fruit, revealing DNA methylation-mediated freshness preservation mechanisms.

DNA Methylation

Chronic disease in a 15th-century skeleton from the first European settlement of the Canary Islands: Early evidence in the colonial Atlantic expansion (San Marcial de Rubicón, Lanzarote).

OBJECTIVE: This study seeks to evaluate morphological changes in a skeleton recovered from San Marcial de Rubicón, the earliest permanent European settlement in the Canary Islands. MATERIALS: The individual derives from a primary inhumation dated to the early fifteenth century. METHODS: Macroscopic observation was combined with conventional radiography, computed tomography, and mitochondrial DNA analysis. A systematic differential diagnosis considered metabolic, infectious, inflammatory, neoplastic, and degenerative conditions. RESULTS: The individual under investigation is an adult male with evidence of diffuse cortical thickening, periosteal new bone formation, heterogeneous radiodensity, cranial diploic expansion, long-bone bowing, severe degenerative joint disease, sacroiliac ankylosis, and elongated thoracic vertebral defects. Mitochondrial DNA analysis identified haplogroup X2c1, consistent with European maternal ancestry. CONCLUSIONS: The overall pattern is most consistent with polyostotic Paget disease of bone, although coexisting axial ankylosis and vertebral defects complicate the interpretation. These additional lesions are insufficient to support an alternative primary diagnosis. SIGNIFICANCE: This case provides an early extra-European archaeological example of Paget disease in the context of Atlantic colonial expansion. Rather than simply extending the geographic record of the disease, it shows how chronic skeletal conditions with strong European clinical and archaeological associations may be identified in frontier populations formed through mobility, settlement and colonial interaction. LIMITATIONS: The diagnosis is based on a single individual, and nuclear DNA data were insufficient to assess genetic susceptibility. SUGGESTIONS FOR FURTHER RESEARCH: Further radiological, genomic, and isotopic analyses of early colonial skeletal assemblages are needed to evaluate chronic disease, mobility, and biological diversity in Atlantic frontier populations.

Male

Genome-wide DNA methylation analysis revealed epigenetic mechanism underlying end-stage renal disease.

End-stage renal disease (ESRD) remains a major clinical challenge with high morbidity and mortality, and its molecular mechanisms, particularly those shared among diverse primary kidney diseases during progression to ESRD, have not been studied. Here we conduct a large-scale two-stage epigenome-wide association study of ESRD in two independent cohorts consisting of 704 controls and 1031 ESRD cases. We identify 52 ESRD-associated differentially methylated CpG positions (ESRD DMPs) showing consistent association between the two cohorts and across diverse kidney diseases, implicating 144 candidate genes enriched in inflammatory and immune pathways. Five of the 52 DMPs are associated with ESRD complications, and seven with renal function decline in early-stage chronic kidney disease, demonstrating their potential as prognostic biomarkers for ESRD and its complications. Our findings highlight inflammation, immune dysregulation, and renal fibrosis as shared epigenetic drivers of ESRD progression, and identify biomarkers with potential utility for risk stratification and therapeutic intervention.

Humans

Estimating the sensitivity of genomic newborn screening for treatable inherited metabolic disorders.

PURPOSE: Over 30 research groups and companies are exploring newborn screening using genomic sequencing (NBSeq), but the sensitivity of this approach is not well understood. METHODS: We identified individuals with treatable inherited metabolic disorders (IMDs) and ascertained the proportion whose DNA analysis revealed explanatory deleterious variants (EDVs). We examined variables associated with EDV detection and estimated the sensitivity of DNA-first NBSeq. We further predicted the annual rate of true-positive and false-negative NBSeq results in the United States for several conditions on the Recommended Uniform Screening Panel. RESULTS: We identified 635 individuals with 80 unique IMDs. In univariate analyses, Black race (OR = 0.37, 95% CI: 0.16-0.89, P = .02) and public insurance (OR = 0.60, 95% CI: 0.39-0.91, P = .02) were less likely to be associated with finding EDVs. Had all individuals been screened with NBSeq, the sensitivity would have been 80.3%. We estimated that between 0 and 649.9 cases of Recommended Uniform Screening Panel IMDs would be missed annually by NBSeq in the United States. CONCLUSION: The overall sensitivity of NBSeq for treatable IMDs is estimated at 80.3%. That sensitivity will likely be lower for Black infants and those who are on public insurance.

Humans

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

Whole genome and exome sequencing of pancreatic neuroendocrine tumour to investigate PRRT response.

Patients with pancreatic neuroendocrine tumours (PNETs) often have similar baseline clinical characteristics, including grade and molecular imaging phenotype, yet have highly variable responses to peptide receptor radionuclide therapy (PRRT). To identify genomic alterations and mutational patterns associated with PRRT treatment response and acquired somatic changes following PRRT exposure, whole genome or exome sequencing was applied to 40 PNET samples from 32 patients, including eight paired pre- or post-PRRT samples. The genomic profile of tumours reflected the known mutational landscape of PNET with MEN1 (34%), ATRX/DAXX (47%) alterations and a recurrent pattern of aneuploidy (38%) detected. A recurrent PSIP1::TBL1X fusion of unknown function was also identified in four tumours. The disease control rate following PRRT using RECIST1.1 and molecular imaging criteria was 88% (28/32). No mutational features were found to be statistically associated with progression-free survival. There was no significant increase in tumour mutational burden in the post-PRRT tumours, nor recurrent emergent mutational changes in cancer driver genes to explain progression to higher-grade disease, when observed. However, a small indel signature (ID8) previously associated with DNA damage repair by non-homologous end joining (NHEJ) was higher in PRRT-exposed compared with PRRT-naive samples (23.8 vs 4.8%, respectively; P < 0.001). Thus, comprehensive DNA analysis of pancreatic NETs did not identify biomarkers predictive of PRRT response nor evidence for high-level PRRT-induced genomic instability or hypermutation, yet mutation signature analysis supports NHEJ as being important for DNA repair and survival of neuroendocrine cells following exposure to beta-particle radiation.

Humans

Exploring the potential of genetic analysis in historical blood spots for patients with iodine-deficient goiter and thyroid carcinomas in Switzerland and Germany (1929-1989).

Iodine deficiency-induced goiter continues to be a global public health concern, with varying manifestations based on geography, patient's age, and sex. To gain insights into clinical occurrences, a retrospective study analyzed medical records from patients with iodine deficiency-induced goiter or thyroid cancer who underwent surgery at the Community Hospital in Riehen, Switzerland, between 1929 and 1989. Despite today's adequate iodine supplementation, a significant risk for iodine-independent goiter remains in Switzerland, suggesting that genetic factors, among others, might be involved. Thus, a pilot study exploring the feasibility of genetic analysis of blood spots from these medical records was conducted to investigate and enhance the understanding of goiter development, potentially identify genetic variations, and explore the influence of dietary habits and other environmental stimuli on the disease.Blood prints from goiter patients' enlarged organs were collected per decade from medical records. These prints had been made by pressing, drawing, or tracing (i.e., pressed and drawn) the removed organs onto paper sheets. DNA analysis revealed that its yields varied more between the prints than between years. A considerable proportion of the samples exhibited substantial DNA degradation unrelated to sample collection time and DNA mixtures of different contributors. Thus, each goiter imprint must be individually evaluated and cannot be used to predict the success rate of genetic analysis in general. Collecting a large sample or the entire blood ablation for genetic analysis is recommended to mitigate potential insufficient DNA quantities. Researchers should also consider degradation and external biological compounds' impact on the genetic analysis of interest, with the dominant contributor anticipated to originate from the patient's blood.

Humans

Pituitary Neuroendocrine Tumor or Pituitary Adenoma? Let's Ask the Epigenome!

The introduction of the term pituitary neuroendocrine tumor (PitNET) to replace pituitary adenoma has sparked a versatile debate among experts. The controversy surrounding this nomenclature change includes the question of whether these tumors' biological identity truly corresponds to neuroendocrine tumors. In this meta-analysis, DNA methylation data were interrogated to clarify whether the old or new nomenclature more accurately reflects the epigenome of these tumors. Publicly available DNA methylation data of 100 NETs, 100 PitNETs/adenomas, and 100 adenomas of various origins and lineages were compiled from 18 different publications. Epigenomic signatures characteristic of NETs and adenomas were defined and compared to those of PitNETs/adenomas. Promoter CpG methylation levels were investigated for hallmarks of cellular differentiation. Comparative DNA methylation analyses demonstrated that all 100 PitNETs/adenomas aligned more closely with NETs than with adenomas. Focusing on promoter-associated CpGs moreover confirmed robust epigenomic features associated with neuroendocrine differentiation in PitNETs/adenomas. These findings indicate that&#xa0;PitNETs/adenomas resemble NETs rather than adenomas on the epigenomic level&#xa0;and support PitNET as the biologically more accurate term. Of note, appropriately addressing the broad spectrum of clinical behaviors in these tumors remains a critical issue in the current pituitary tumor classification framework and nomenclature.

Humans

A probabilistic generative model for quantification of DNA modifications enables analysis of demethylation pathways.

We present a generative model, Lux, to quantify DNA methylation modifications from any combination of bisulfite sequencing approaches, including reduced, oxidative, TET-assisted, chemical-modification assisted, and methylase-assisted bisulfite sequencing data. Lux models all cytosine modifications (C, 5mC, 5hmC, 5fC, and 5caC) simultaneously together with experimental parameters, including bisulfite conversion and oxidation efficiencies, as well as various chemical labeling and protection steps. We show that Lux improves the quantification and comparison of cytosine modification levels and that Lux can process any oxidized methylcytosine sequencing data sets to quantify all cytosine modifications. Analysis of targeted data from Tet2-knockdown embryonic stem cells and T cells during development demonstrates DNA modification quantification at unprecedented detail, quantifies active demethylation pathways and reveals 5hmC localization in putative regulatory regions.

5-Methylcytosine