Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “spike-in”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Improved spike-in normalization clarifies the relationship between active histone modifications and transcription.

Spike-in normalization enables quantitative analysis of chromatin immunoprecipitation sequencing (ChIP-seq) signal. Here we introduce a robust dual spike-in normalization approach for ChIP-seq (ChIP-wrangler), optimize parameters and verify its accuracy in quantifying changes in ChIP-seq signal and detecting technical artifacts. We use ChIP-wrangler to revisit recent claims that active histone marks depend on transcription. We show that acute depletion of RNA polymerase II (RNAPII) has a modest impact on H3K27ac levels, with only 6% of peaks significantly changing after RNAPII depletion, indicating that histone acetylation maintenance is not entirely dependent on ongoing transcription. Promoters and enhancers are differentially affected, with 82% of decreasing acetylation peaks located at promoter-distal elements with enhancer-related motifs. ChIP-wrangler provides increased rigor and 'guardrails' for successful spike-in normalization and, as applied here, refines the understanding of crosstalk between RNAPII activity and transcription-associated histone marks.

Histones↗

Combining Data Independent Acquisition With Spike-In SILAC (DIA-SiS) Improves Proteome Coverage and Quantification.

Data-independent acquisition (DIA) is increasingly preferred over data-dependent acquisition due to its higher throughput and fewer missing values. Whereas data-dependent acquisition often uses stable isotope labeling to improve quantification, DIA mostly relies on label-free approaches. Efforts to integrate DIA with isotope labeling include chemical methods like mass differential tags for relative and absolute quantification and dimethyl labeling, which, while effective, complicate sample preparation. Stable isotope labeling by amino acids in cell culture (SILAC) achieves high labeling efficiency through the metabolic incorporation of heavy labels into proteins in vivo. However, the need for metabolic incorporation limits the direct use in clinical scenarios and certain high-throughput experiments. Spike-in SILAC (SiS) methods use an externally generated heavy sample as an internal reference, enabling SILAC-based quantification even for samples that cannot be directly labeled. Here, we combine DIA-SiS, leveraging the robust quantification of SILAC without the complexities associated with chemical labeling. We developed DIA-SiS and rigorously assessed its performance with mixed-species benchmark samples on bulk and single cell-like amount level. We demonstrate that DIA-SiS substantially improves proteome coverage and quantification compared to label-free approaches and reduces incorrectly quantified proteins. Additionally, DIA-SiS proves effective in analyzing proteins in low-input formalin-fixed paraffin-embedded tissue sections. DIA-SiS combines the precision of stable isotope-based quantification with the simplicity of label-free sample preparation, facilitating simple, accurate, and comprehensive proteome profiling.

Isotope Labeling↗

ChIP-Rx: Arabidopsis Chromatin Profiling Using Quantitative ChIP-Seq.

Chromatin immunoprecipitation followed by deep sequencing (ChIP-seq) is widely used to probe the chromatin landscape of transcription factors, chromatin components, and associated proteins. Conventional ChIP normalization procedures robustly allow estimating differences in local enrichment across genomic regions. Yet, inter-sample comparisons can be biased by technical variability and biological differences. This is notably the case when samples display large differences in the abundance of the target protein or its enrichment at chromatin. For example, epigenome defects are improperly detected or quantified upon large-effect genetic or chemical inhibition of chromatin modifiers. To circumvent these caveats and robustly determine biological variations while minimizing technical variability, ChIP adaptations using an external reference have flourished. Here, we describe a step-by-step protocol employing a reference exogenous chromatin (ChIP-Rx) that allows absolute comparisons of epigenome variations in Arabidopsis samples displaying drastic differences in chromatin mark abundance. In contrast to the originally published ChIP-Rx approach, which assumes that exogenous spike-in references are constant across samples, the method detailed here involves the sequencing of each input sample to account for technical variability in initial reference chromatin contents. We also report a detailed computational workflow with an accompanying Github resource to help in calculating spike-in normalization factors, applying them to normalize epigenome tracks, and performing spike-in normalized inter-sample differential analyses. We propose two ways of computing the spike-in factor: a classically used method based on raw counts and a noise-corrected method using peak detection on the exogenous genome.

Arabidopsis↗

Exploration, normalization, and summaries of high density oligonucleotide array probe level data.

In this paper we report exploratory analyses of high-density oligonucleotide array data from the Affymetrix GeneChip system with the objective of improving upon currently used measures of gene expression. Our analyses make use of three data sets: a small experimental study consisting of five MGU74A mouse GeneChip arrays, part of the data from an extensive spike-in study conducted by Gene Logic and Wyeth's Genetics Institute involving 95 HG-U95A human GeneChip arrays; and part of a dilution study conducted by Gene Logic involving 75 HG-U95A GeneChip arrays. We display some familiar features of the perfect match and mismatch probe (PM and MM) values of these data, and examine the variance-mean relationship with probe-level data from probes believed to be defective, and so delivering noise only. We explain why we need to normalize the arrays to one another using probe level intensities. We then examine the behavior of the PM and MM using spike-in data and assess three commonly used summary measures: Affymetrix's (i) average difference (AvDiff) and (ii) MAS 5.0 signal, and (iii) the Li and Wong multiplicative model-based expression index (MBEI). The exploratory data analyses of the probe level data motivate a new summary measure that is a robust multi-array average (RMA) of background-adjusted, normalized, and log-transformed PM values. We evaluate the four expression summary measures using the dilution study data, assessing their behavior in terms of bias, variance and (for MBEI and RMA) model fit. Finally, we evaluate the algorithms in terms of their ability to detect known levels of differential expression using the spike-in data. We conclude that there is no obvious downside to using RMA and attaching a standard error (SE) to this quantity using a linear model which removes probe-specific affinities.

Algorithms↗

Cerebrospinal fluid microbiome revisited: no evidence of resident bacteria in archived samples.

UNLABELLED: DNA from oral bacteria has been detected in the cerebrospinal fluid (CSF) of patients with Alzheimer's disease and related dementias (AD/ADRD). We hypothesized that examination of archived CSF samples from donors with variable cognitive status would reveal evidence of a resident microbiome. 176 CSF samples harvested from community-dwelling individuals (77% between 61 and 80 years old) were analyzed; 57% originated from donors with impaired cognitive status. DNA was extracted after adding microbial spike-in controls, and libraries were prepared and sequenced on an Illumina-MiSeq platform. 16S rRNA sequences were processed, and a taxonomic classification was performed. Spike-in bacteria were consistently detected, and Streptococcus pneumoniae was found in a positive control sample from a patient with bacterial meningitis. However, very few reads mapping to other bacterial taxa were detected across samples, suggesting a negligible bacterial content consistent with occasional contamination or sequencing errors. CSF is a privileged, sterile environment that does not harbor a resident microbiome in elderly people with various morbidities, including AD/ADRD. IMPORTANCE: Recent reports have suggested that DNA from oral bacteria has been found in the cerebrospinal fluid (CSF) of patients with Alzheimer's disease and related dementias (AD/ADRD). We hypothesized that examination of archived CSF samples from donors with variable cognitive status would reveal evidence of a resident microbiome. We thus analyzed 176 CSF samples harvested from community-dwelling individuals including donors with impaired cognitive status. While our findings suggested the presence of occasional bacterial contamination, they provided no evidence of a resident microbiome. We thus conclude that the CSF is indeed a privileged, sterile environment that does not harbor a resident microbiome in elderly people with various morbidities including AD/ADRD.

Humans↗

Unbiased characterization of high-density oligonucleotide microarrays using probe-level statistics.

Affymetrix GeneChips are being used increasingly for quantitative monitoring of gene expression in a variety of biological systems. Depending on the experiment, the analysis of Affymetrix results can have several different goals ranging from calculation of signal strength for a variety of inter-gene comparisons to the determination of which genes show significant differential expression between sample conditions. There have been several proposed methods for precise quantification of expression signal with promising results; however the question of what constitutes a significant change between replicate groups still remains. We have designed a method which performs statistical analysis on the differential expression of genes in the Affymetrix GeneChip system at the probe level in order to bypass the assumptions made in other analysis techniques. Validation using both spike-in data and real experimental data proves the method is effective at isolating differentially expressed genes statistically, thereby eliminating the need for arbitrary restrictions such as fold change. Application to an existing neural stem cell data set demonstrates the method's applicability to highly complex systems and its ability to detect very low expression differences (<1.2-fold change), providing resolution which may be of significant interest in neural systems such as this.

Animals↗

No receptor-binding domain adaptation detected in within-host H5N1 surveillance of 4,559 US dairy outbreak sequences.

BACKGROUND: The 2024-2026 US H5N1 clade 2.3.4.4b dairy cattle outbreak has been characterised primarily through consensus-level phylogenetics. Whether mammalian-adaptation variants are emerging at sub-consensus frequencies within infected hosts, particularly at the haemagglutinin receptor-binding domain (RBD), remains unknown because no systematic within-host variant analysis of the public sequencing corpus has been performed. METHODS: We conducted a pre-registered, corpus-wide intrahost single-nucleotide variant (iSNV) analysis of all publicly available H5N1 cattle, feline-spillover, and retail-milk sequences on the NCBI Sequence Read Archive (4559 samples across 7 BioProjects). A dual-caller concordance pipeline (iVar&#xa0;+&#xa0;LoFreq) with empirically determined allele frequency (AF) threshold (3%, set via four-criterion validation including synthetic spike-in controls) was applied to an 11-site Tier 1 mammalian-adaptation panel spanning the polymerase complex, haemagglutinin RBD, and accessory proteins. Within-host nucleotide diversity was compared across host categories. RESULTS: The HA RBD sites Q226L and G228S (H3 numbering) showed zero detections across >4300 adequately sequenced samples at all AF thresholds tested (1-5%), despite the pipeline detecting other non-synonymous variants at these exact codon positions (upper 95% CI for prevalence: 0.08%). Seven of eleven adaptation sites carried statistically significant iSNV signals after Bonferroni correction (corrected &#x3b1;&#x202f;=&#x202f;0.00417), though all at low prevalence (&#x2264;2.95%). Genotype stratification showed that most polymerase-site detections reflected genotype structure rather than within-host emergence: the apparent PB2 631&#x202f;L&#x2192;M "reversion" was largely the ancestral avian state of the D1.1 genotype (20 of 23 detections), which never acquired the 631L mammalian adaptation, with only two genuine sub-consensus events in the B3.13 background, while consensus-level PB2 701N was a fixed feature of the D1.1 genotype (10 of 14 detections) rather than independent sub-consensus emergence. Cattle exhibited significantly higher within-host nucleotide diversity than feline-spillover samples (&#x3c0;&#x202f;=&#x202f;1.59&#x202f;&#xd7;&#x202f;10-4 vs 6.11&#x202f;&#xd7;&#x202f;10-5; Kruskal-Wallis p&#x202f;=&#x202f;6.6&#x202f;&#xd7;&#x202f;10-15), a finding that persisted after depth-matching (p&#x202f;=&#x202f;4.6&#x202f;&#xd7;&#x202f;10-5); this may reflect prolonged mammary-gland infection, though sampling differences and host biology cannot be excluded. CONCLUSIONS: We did not detect HA receptor-switching adaptation (the acquisition of human-type &#x3b1;2,6 receptor binding via Q226L/G228S) at any tested allele frequency in the US dairy H5N1 outbreak. Sub-consensus mammalian-adaptation signals exist at polymerase-complex sites but at low prevalence, are genotype-structured rather than independently recurrent, and require functional characterisation before informing risk assessment.

Dairy cattle↗

Gene expression profiles in end-stage human idiopathic dilated cardiomyopathy: altered expression of apoptotic and cytoskeletal genes.

Dilated cardiomyopathy is now the leading cause of cardiovascular morbidity and mortality. While the molecular basis of this disease remains uncertain, evidence is emerging that gene expression profiles of left ventricular myocardium isolated from failing versus nonfailing patients differ dramatically. In this study, we use high-density oligonucleotide microarrays with approximately 22000 probes to characterize differences in the expression profiles further. To facilitate interpretation of experimental data, we evaluate algorithms for normalization of hybridization data and for computation of gene expression indices using a control spike-in data set. We then use these methods to identify statistically significant changes in the expression levels of genes not previously implicated in the molecular phenotype of heart failure. These regulated genes take part in diverse cellular processes, including transcription, apoptosis, sarcomeric and cytoskeletal function, remodeling of the extracellular matrix, membrane transport, and metabolism.

Algorithms↗

Standardization of protocols in cDNA microarray analysis.

Systematic variations can occur at various steps of a cDNA microarray experiment and affect the measurement of gene expression levels. Accepted standards integrated into every cDNA microarray analysis can assess these variabilities and aid the interpretation of cDNA microarray experiments from different sources. A universally applicable approach to evaluate parameters such as input and output ratios, signal linearity, hybridization specificity and consistency across an array, as well as normalization strategies, is the utilization of exogenous control genes as spike-in and negative controls. We suggest that the use of such control sets, together with a sufficient number of experimental repeats, in-depth statistical analysis and thorough data validation should be made mandatory for the publication of cDNA microarray data.

DNA, Complementary↗

Improved Detection of Differentially Abundant Proteins through FDR-Control of Peptide-Identity-Propagation.

The goal of proteomics is to identify and quantify peptides and proteins within a biological sample. Almost all algorithms for the identification of peptides in LC-MS/MS data employ two steps: peptide/spectrum matching and peptide-identity-propagation (PIP), also known as match-between-runs. PIP can routinely account for up to 40% of all results, with that proportion rising as high as 75% in single-cell proteomics. Unlike peptide identities derived through peptide/spectrum matches, for which error estimation has been strictly enforced for decades, peptide identities derived through PIP have not historically been subject to statistical evaluation. As an indispensable component of label-free quantification, PIP needs a statistically rigorous method for estimating its false-discovery rate (FDR). We present a method for FDR control of PIP, called PIP-ECHO, and devise a rigorous protocol for evaluating FDR control of any PIP method. Using three different benchmark data sets, we evaluate PIP-ECHO alongside the PIP procedures implemented by FlashLFQ, IonQuant, and MaxQuant. These analyses show that only PIP-ECHO can accurately control the FDR of PIP at 1% across all data sets. When analyzing a spike-in data set, PIP-ECHO increases both the accuracy and sensitivity of differential expression analysis, yielding substantially more differentially abundant proteins than either MaxQuant or IonQuant.

Proteomics↗

Evaluating the analytical validity of circulating tumor DNA sequencing assays for precision oncology.

Circulating tumor DNA (ctDNA) sequencing is being rapidly adopted in precision oncology, but the accuracy, sensitivity and reproducibility of ctDNA assays is poorly understood. Here we report the findings of a multi-site, cross-platform evaluation of the analytical performance of five industry-leading ctDNA assays. We evaluated each stage of the ctDNA sequencing workflow with simulations, synthetic DNA spike-in experiments and proficiency testing on standardized, cell-line-derived reference samples. Above 0.5% variant allele frequency, ctDNA mutations were detected with high sensitivity, precision and reproducibility by all five assays, whereas, below this limit, detection became unreliable and varied widely between assays, especially when input material was limited. Missed mutations (false negatives) were more common than erroneous candidates (false positives), indicating that the reliable sampling of rare ctDNA fragments is the key challenge for ctDNA assays. This comprehensive evaluation of the analytical performance of ctDNA assays serves to inform best practice guidelines and provides a resource for precision oncology.

Circulating Tumor DNA↗

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem↗

Chrom-Sig: de-noising 1D genomic profiles by signal processing methods.

MOTIVATION: Modern genomic research is driven by next-generation sequencing experiments such as ChIP-seq, CUT&Tag, and CUT&RUN that generate coverage files for transcription factor binding, as well as ATAC-seq that yield coverage files for chromatin accessibility. Due to the inherent technical noise present in the experimental protocols, researchers need statistically rigorous and computationally efficient methods to extract true biological signal from a mixture of signal and noise. However, existing approaches are often computationally demanding or require input or spike-in controls. RESULTS: We developed Chrom-Sig, a Python package to quickly de-noise 1D genomic coverage tracks by computing the empirical null distribution without prior assumptions or experimental controls. When tested on 19 ChIP-seq, CUT&RUN, ATAC-seq, and snATAC-seq datasets, Chrom-Sig can effectively decompose the data into signal and noise components. Notably, Chrom-Sig performs de-noising and peak calling in 1-2&#x2009;h using around 20&#xa0;GB of memory. The de-noised signal corroborates with biologically meaningful results: CTCF CUT&RUN data retained a high percentage of peaks overlapping CTCF binding motifs, while ATAC-seq and RNA Polymerase II data were enriched in enhancers and promoters. We envision Chrom-Sig to be a versatile and general tool for current and future genomic technologies. AVAILABILITY AND IMPLEMENTATION: Chrom-Sig is publicly available on GitHub (https://github.com/minjikimlab/chromsig) and Zenodo (doi: 10.5281/zenodo.17488772) under the MIT licence.

Genomics↗

A benchmark for Affymetrix GeneChip expression measures.

MOTIVATION: The defining feature of oligonucleotide expression arrays is the use of several probes to assay each targeted transcript. This is a bonanza for the statistical geneticist, who can create probeset summaries with specific characteristics. There are now several methods available for summarizing probe level data from the popular Affymetrix GeneChips, but it is difficult to identify the best method for a given inquiry. RESULTS: We have developed a graphical tool to evaluate summaries of Affymetrix probe level data. Plots and summary statistics offer a picture of how an expression measure performs in several important areas. This picture facilitates the comparison of competing expression measures and the selection of methods suitable for a specific investigation. The key is a benchmark data set consisting of a dilution study and a spike-in study. Because the truth is known for these data, we can identify statistical features of the data for which the expected outcome is known in advance. Those features highlighted in our suite of graphs are justified by questions of biological interest and motivated by the presence of appropriate data.

Algorithms↗

GenePicker: replicate analysis of Affymetrix gene expression microarrays.

UNLABELLED: GenePicker allows efficient analysis of Affymetrix gene expression data performed in replicate, through definition of analysis schemes, data normalization, t-test/ANOVA, Change-Fold Change-analysis and yields lists of differentially expressed genes with high confidence. Comparison of noise and signal analysis schemes allows determining a signal-to-noise ratio in a given experiment. Change Call, Fold Change and Signal mean ratios are used in the analysis. While each parameter alone yields gene lists that contain up to 30% false positives, the combination of these parameters nearly eliminates the false positives as verified by northern blotting, quantitative PCR in numerous independent experiments as well as by the analysis of spike-in data. AVAILABILITY: http://www.ifom-firc.it/RESEARCH/Appl_Bioinfo/tools.html. SUPPLEMENTARY INFORMATION: http://www.ifom-firc.it/RESEARCH/Appl_Bioinfo/tools.html.

Algorithms↗

Identifying differentially expressed genes from microarray experiments via statistic synthesis.

MOTIVATION: A common objective of microarray experiments is the detection of differential gene expression between samples obtained under different conditions. The task of identifying differentially expressed genes consists of two aspects: ranking and selection. Numerous statistics have been proposed to rank genes in order of evidence for differential expression. However, no one statistic is universally optimal and there is seldom any basis or guidance that can direct toward a particular statistic of choice. RESULTS: Our new approach, which addresses both ranking and selection of differentially expressed genes, integrates differing statistics via a distance synthesis scheme. Using a set of (Affymetrix) spike-in datasets, in which differentially expressed genes are known, we demonstrate that our method compares favorably with the best individual statistics, while achieving robustness properties lacked by the individual statistics. We further evaluate performance on one other microarray study.

Algorithms↗

Use of within-array replicate spots for assessing differential expression in microarray experiments.

MOTIVATION: Spotted arrays are often printed with probes in duplicate or triplicate, but current methods for assessing differential expression are not able to make full use of the resulting information. The usual practice is to average the duplicate or triplicate results for each probe before assessing differential expression. This results in the loss of valuable information about genewise variability. RESULTS: A method is proposed for extracting more information from within-array replicate spots in microarray experiments by estimating the strength of the correlation between them. The method involves fitting separate linear models to the expression data for each gene but with a common value for the between-replicate correlation. The method greatly improves the precision with which the genewise variances are estimated and thereby improves inference methods designed to identify differentially expressed genes. The method may be combined with empirical Bayes methods for moderating the genewise variances between genes. The method is validated using data from a microarray experiment involving calibration and ratio control spots in conjunction with spiked-in RNA. Comparing results for calibration and ratio control spots shows that the common correlation method results in substantially better discrimination of differentially expressed genes from those which are not. The spike-in experiment also confirms that the results may be further improved by empirical Bayes smoothing of the variances when the sample size is small. AVAILABILITY: The methodology is implemented in the limma software package for R, available from the CRAN repository http://www.r-project.org

Algorithms↗

Detecting differential gene expression with a semiparametric hierarchical mixture method.

Mixture modeling provides an effective approach to the differential expression problem in microarray data analysis. Methods based on fully parametric mixture models are available, but lack of fit in some examples indicates that more flexible models may be beneficial. Existing, more flexible, mixture models work at the level of one-dimensional gene-specific summary statistics, and so when there are relatively few measurements per gene these methods may not provide sensitive detectors of differential expression. We propose a hierarchical mixture model to provide methodology that is both sensitive in detecting differential expression and sufficiently flexible to account for the complex variability of normalized microarray data. EM-based algorithms are used to fit both parametric and semiparametric versions of the model. We restrict attention to the two-sample comparison problem; an experiment involving Affymetrix microarrays and yeast translation provides the motivating case study. Gene-specific posterior probabilities of differential expression form the basis of statistical inference; they define short gene lists and false discovery rates. Compared to several competing methodologies, the proposed methodology exhibits good operating characteristics in a simulation study, on the analysis of spike-in data, and in a cross-validation calculation.

Algorithms↗