Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Signature”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

A gene-expression signature to predict survival in breast cancer across independent data sets.

Prognostic signatures in breast cancer derived from microarray expression profiling have been reported by two independent groups. These signatures, however, have not been validated in external studies, making clinical application problematic. We performed microarray expression profiling of 135 early-stage tumors, from a cohort representative of the demographics of breast cancer. Using a recently proposed semisupervised method, we identified a prognostic signature of 70 genes that significantly correlated with survival (hazard ratio (HR): 5.97, 95% confidence interval: 3.0-11.9, P = 2.7e-07). In multivariate analysis, the signature performed independently of other standard prognostic classifiers such as the Nottingham Prognostic Index and the 'Adjuvant!' software. Using two different prognostic classification schemes and measures, nearest centroid (HR) and risk ordering (D-index), the 70-gene classifier was also found to be prognostic in two independent external data sets. Overall, the 70-gene set was prognostic in our study and the two external studies which collectively include 715 patients. In contrast, we found that the two previously described prognostic gene sets performed less optimally in external validation. Finally, a common prognostic module of 29 genes that associated with survival in both our cohort and the two external data sets was identified. In spite of these results, further studies that profile larger cohorts using a single microarray platform, will be needed before prospective clinical use of molecular classifiers can be contemplated.

Breast Neoplasms↗

Signature currents: a patch-clamp method for determining the selectivity of ion-channel blockers in isolated cardiac myocytes.

BACKGROUND: We describe a simple method using membrane potential ramps for rapidly determining the ion-channel selectivity of drugs that affect action-potential duration in isolated cardiac myocytes. The method allows the simultaneous assay of compounds on a number of ionic currents in a single cardiac cell. METHODS: Trains of membrane potential ramps were applied from -90 to +70 mV at 0.33 Hz to obtain a consistent "signature current," in which the major individual currents involved in the cardiac action potential could be easily identified. Confirmatory experiments were performed using known inhibitors of these currents. RESULTS: The identities of the currents in the signature were established by varying the concentrations of extracellular cations and by adding known ion channel blockers to superfusion solutions. Inhibition of each current had a characteristic and reproducible effect on the overall signature current. CONCLUSIONS: The consistent current signature in the presence and absence of blockers suggests that this method could be used for tertiary electrophysiological evaluation of compounds, eg, in a drug discovery program focusing on antiarrhythmic agents. The ability to assay for secondary effects of novel compounds against multiple currents in the target cell type is convenient and avoids the artefacts associated with using artificial expression systems.

Action Potentials↗

The prognostic role of a gene signature from tumorigenic breast-cancer cells.

BACKGROUND: Breast cancers contain a minority population of cancer cells characterized by CD44 expression but low or undetectable levels of CD24 (CD44+CD24-/low) that have higher tumorigenic capacity than other subtypes of cancer cells. METHODS: We compared the gene-expression profile of CD44+CD24-/low tumorigenic breast-cancer cells with that of normal breast epithelium. Differentially expressed genes were used to generate a 186-gene "invasiveness" gene signature (IGS), which was evaluated for its association with overall survival and metastasis-free survival in patients with breast cancer or other types of cancer. RESULTS: There was a significant association between the IGS and both overall and metastasis-free survival (P<0.001, for both) in patients with breast cancer, which was independent of established clinical and pathological variables. When combined with the prognostic criteria of the National Institutes of Health, the IGS was used to stratify patients with high-risk early breast cancer into prognostic categories (good or poor); among patients with a good prognosis, the 10-year rate of metastasis-free survival was 81%, and among those with a poor prognosis, it was 57%. The IGS was also associated with the prognosis in medulloblastoma (P=0.004), lung cancer (P=0.03), and prostate cancer (P=0.01). The prognostic power of the IGS was increased when combined with the wound-response (WR) signature. CONCLUSIONS: The IGS is strongly associated with metastasis-free survival and overall survival for four different types of tumors. This genetic signature of tumorigenic breast-cancer cells was even more strongly associated with clinical outcomes when combined with the WR signature in breast cancer.

Breast↗

Mutational analysis of a fatty acyl-coenzyme A synthetase signature motif identifies seven amino acid residues that modulate fatty acid substrate specificity.

Fatty acyl-CoA synthetase (fatty acid:CoA ligase, AMP-forming; EC 6.2.1.3) catalyzes the formation of fatty acyl-CoA by a two-step process that proceeds through the hydrolysis of pyrophosphate. In Escherichia coli this enzyme plays a pivotal role in the uptake of long chain fatty acids (C12-C18) and in the regulation of the global transcriptional regulator FadR. The E. coli fatty acyl-CoA synthetase has remarkable amino acid similarities and identities to the family of both prokaryotic and eukaryotic fatty acyl-CoA synthetases, indicating a common ancestry. Most notable in this regard is a 25-amino acid consensus sequence, DGWLHTGDIGXWXPXGXLKIIDRKK, common to all fatty acyl-CoA synthetases for which sequence information is available. Within this consensus are 8 invariant and 13 highly conserved amino acid residues in the 12 fatty acyl-CoA synthetases compared. We propose that this sequence represents the fatty acyl-CoA synthetase signature motif (FACS signature motif). This region of fatty acyl-CoA synthetase from E. coli, 431NGWLHTGDIAVMDEEGFLRIVDRKK455, contains 17 amino acid residues that are either identical or highly conserved to the FACS signature motif. Eighteen site-directed mutations within the fatty acyl-CoA synthetase structural gene (fadD) corresponding to this motif were constructed to evaluate the contribution of this region of the enzyme to catalytic activity. Three distinct classes of mutations were identified on the basis of growth characteristics on fatty acids, enzymatic activities using cell extracts, and studies using purified wild-type and mutant forms of the enzyme: 1) those that resulted in either wild-type or nearly wild-type fatty acyl-CoA synthetase activity profiles; 2) those that had little or no enzyme activity; and 3) those that resulted in lowering and altering fatty acid chain length specificity. Among the 18 mutants characterized, 7 fall in the third class. We propose that the FACS signature motif is essential for catalytic activity and functions in part to promote fatty acid chain length specificity and thus may compose part of the fatty acid binding site within the enzyme.

Amino Acid Sequence↗

Biophysical characterization of the signature domains of thrombospondin-4 and thrombospondin-2.

The signature domain of thrombospondins consists of tandem epidermal growth factor-like modules, 13 calcium-binding repeats, and a lectin-like module. Although very similar, the signature domains of thrombospondin-1 and -2 differ in several potentially important ways from the domains of thrombospondin-3, -4, and -5. We have compared matching recombinant segments representing the signature domains of thrombospondin-2 and -4. In the presence of 2 mM CaCl2, the far UV circular dichroism spectra of thrombospondin-2 and -4 constructs contain a strong negative band at 202 nm, but only the thrombospondin-2 construct has a band at 216 nm. Chelation of calcium shifted the negative bands to lower magnitudes. Titrations of the spectra demonstrated lower cooperativity and affinity for binding of calcium to thrombospondin-4 compared with thrombospondin-2. Atomic absorption spectroscopy demonstrated that the thrombospondin-4 constructs bind seven less calcium than the thrombospondin-2 construct at 0.6 mM CaCl2. In 2 mM CaCl2, the near UV circular dichroism spectra of thrombospondin-2, but not thrombospondin-4, contain a positive band at 292 nm that disappears upon calcium chelation. Intrinsic fluorescence spectra for both proteins were also sensitive to calcium, but the changes were simpler and more marked for thrombospondin-2 than for thrombospondin-4. In differential scanning calorimetry, the thrombospondin-2 construct melted in two distinct transitions at 53.5 and 81.8 degrees C, whereas the first transition for thrombospondin-4 constructs was observed at 63.5 degrees C. Thus, the studies revealed significant differences between the signature domains of thrombospondin-2 and thrombospondin-4 in calcium binding, fine structure, and inter-modular interactions.

Amino Acid Motifs↗

Identification of proteomic signatures of exposure to marine pollutants in mussels (Mytilus edulis).

Bivalves and especially mussels are very good indicators of marine and estuarine pollution, and so they have been widely used in biomonitoring programs all around the world. However, traditional single parameter biomarkers face the problem of high sensitivity to biotic and abiotic factors. In our study, digestive gland peroxisome-enriched fractions of Mytilus edulis (L., 1758) were analyzed by DIGE and MS. We identified several proteomic signatures associated with the exposure to several marine pollutants (diallyl phthalate, PBDE-47, and bisphenol-A). Animals collected from North Atlantic Sea were exposed to the contaminants independently under controlled laboratory conditions. One hundred and eleven spots showed a significant increase or decrease in protein abundance in the two-dimensional electrophoresis maps from the groups exposed to pollutants. We obtained a unique protein expression signature of exposure to each of those chemical compounds. Moreover a set of proteins composed a proteomic signature in common to the three independent exposures. It is remarkable that the principal component analysis of these spots showed a discernible separation between groups, and so did the hierarchical clustering into four classes. The 14 proteins identified by MS participate in alpha- and beta-oxidation pathways, xenobiotic and amino acid metabolism, cell signaling, oxyradical metabolism, peroxisomal assembly, respiration, and the cytoskeleton. Our results suggest that proteomic signatures could become a valuable tool to monitor the presence of pollutants in field experiments where a mixture of pollutants is often present. Further studies on the identified proteins could provide crucial information to understand possible mechanisms of toxicity of single xenobiotics or mixtures of them in marine ecosystems.

Animals↗

Use of the genomic signature in bacterial classification and identification.

In this study we investigated the correlation between dinucleotide relative abundance values (the genomic signature) obtained from bacterial whole-genome sequences and two parameters widely used for bacterial classification, 16S rDNA sequence similarity and DNA-DNA hybridisation values. Twenty-eight completely sequenced bacterial genomes were included in the study. The correlation between the genomic signature and DNA-DNA hybridisation values was high and taxa that showed less than 30% DNA-DNA binding will in general not have dinucleotide relative abundance dissimilarity (delta*) values below 40. On the other hand, taxa showing more than 50% DNA-DNA binding will not have delta* values higher than 17. Our data indicate that the overall correlation between genomic signature and 16S rDNA sequence similarity is low, except for closely related organisms (16S rDNA similarity >94%). Statistical analysis of delta* values between different subgroups of the Proteobacteria indicate that the beta- and gamma-Proteobacteria are more closely related to each other than to the other subgroups of the Proteobacteria and that the alpha- and epsilon-Proteobacteria form clearly separate subgroups. Using the genomic signature we have also predicted DNA-DNA binding values for fastidious or unculturable endosymbionts belonging to the genera Rickettsia, Wigglesworthia and Buchnera.

Bacteria↗

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software↗

hypeR-GEM: connecting metabolite signatures to enzyme-coding genes via genome-scale metabolic models.

MOTIVATION: Enrichment analysis is a cornerstone of "omics" data interpretation, enabling researchers to connect analysis results to biological processes and generate testable hypotheses. Enrichment analysis in metabolomics poses distinct challenges for interpretation and multi-omics integration due to the lack of well-defined and consistent connections to well-curated gene-centered biological knowledge repositories. To address these challenges, we developed hypeR-GEM, a methodology and associated R package that adapts gene set enrichment analysis to metabolomics. hypeR-GEM leverages genome-scale metabolic models (GEMs) to infer reaction-based links between metabolites and enzyme-coding genes, enabling the mapping of metabolite signatures to gene signatures and their subsequent annotation via gene set enrichment analysis. RESULTS: We validated hypeR-GEM using paired metabolomics-proteomics and metabolomics-transcriptomics datasets by assessing whether genes mapped from metabolites significantly overlapped with differentially expressed proteins or transcripts. We further evaluated whether pathways enriched via hypeR-GEM-mapped genes corresponded to those derived from paired proteomic or transcriptomic data. In most datasets analyzed, both the predicted enzyme-coding genes and the associated enriched pathways showed significant concordance with independently derived omics signatures, supporting the utility and robustness of hypeR-GEM. Finally, we applied hypeR-GEM to the analysis of age-associated metabolic signatures from the New England Centenarian Study. The results revealed consistent enrichment of lipid-related pathways, aligning with the well-established role of lipid metabolism in aging, and highlighted additional pathways not captured in the metabolites' annotation, demonstrating hypeR-GEM's practical utility in a real-world use case. AVAILABILITY AND IMPLEMENTATION: The hypeR-GEM R package, documentation, and workflow examples are freely available at https://github.com/montilab/hypeR-GEM and archived at https://doi.org/10.5281/zenodo.20586748.

Metabolomics↗

Automated generation and refinement of protein signatures: case study with G-protein coupled receptors.

MOTIVATION: Previous work had established that it was possible to derive sparse signatures (essentially sequence-length motifs) by examining points of contact between residues in proteins of known three-dimensional (3D) structure. Many interesting protein families have very little tertiary structural information. Methods for deriving signatures using only primary and secondary-structural information were therefore developed. RESULTS: Two methods for deriving protein signatures using protein sequence information and predicted secondary structures are described. One method is based on a scoring approach, the other on the Genetic Algorithm (GA). The effectiveness of the method was tested on the superfamily of GPCRs and compared with the established hidden Markov model (HMM) method. The signature method is shown to perform well, detecting 68% of superfamily members before the first false positive sequence and detecting several distant relationships. The GA population was used to provide information on alignment regions of particular importance for selection of key residues.

Algorithms↗

CoPS: Comprehensive Peptide Signature database.

UNLABELLED: We present the development of a Comprehensive database of 12 076 invariant Peptide Signatures (CoPS) derived from 52 bacterial genomes with a minimum occurrence in at least seven organisms. These peptides were observed in functionally similar proteins and are distributed over nearly 1250 different functional proteins. The database provides function, structure and occurrence in biochemical pathways of the proteins containing these signature peptides. It houses additional information on the signature peptides, such as identical match in other motif/pattern (e.g. PROSITE, BLOCKS, PRINTS and Pfam) databases and the database of interacting proteins, human proteome and mutation effect on these signature peptides. There is a wide applicability of this database in the identification of critical functional residues in proteins. The database also facilitates the identification of folding nucleus/structural determinants in proteins and functional assignment to yet unknown proteins. We demonstrate functional assignment to 2605 hypothetical proteins in bacterial genomes and 112 unknown proteins in human using this database. AVAILABILITY: The database can be freely accessed through the following URL: http://203.195.151.46/copsv2/index.html or http://203.90.127.70/copsv2/index.html

Bacterial Proteins↗

GATHER: a systems approach to interpreting genomic signatures.

MOTIVATION: Understanding the full meaning of the biology captured in molecular profiles, within the context of the entire biological system, cannot be achieved with a simple examination of the individual genes in the signature. To facilitate such an understanding, we have developed GATHER, a tool that integrates various forms of available data to elucidate biological context within molecular signatures produced from high-throughput post-genomic assays. RESULTS: Analyzing the Rb/E2F tumor suppressor pathway, we show that GATHER identifies critical features of the pathway. We further show that GATHER identifies common biology in a series of otherwise unrelated gene expression signatures that each predict breast cancer outcome. We quantify the performance of GATHER and find that it successfully predicts 90% of the functions over a broad range of gene groups. We believe that GATHER provides an essential tool for extracting the full value from molecular signatures generated from genome-scale analyses. AVAILABILITY: GATHER is available at http://gather.genome.duke.edu/

Algorithms↗

An analysis of signatures of selective sweeps in natural populations of the house mouse.

Population and locus-specific reduction of variability of polymorphic loci could be an indication of positive selection at a linked site (selective sweep) and therefore point toward genes that have been involved in recent adaptations. Analysis of microsatellite variability offers a way to identify such regions and to ask whether they occur more often than expected by chance. We studied four populations of the house mouse (Mus musculus) to assess the frequency of such signatures of selective sweeps under natural conditions. Three samples represent the subspecies Mus m. dometicus [corrected] and came from Germany, France, and Cameroon. One sample came from Kazakhstan and constitutes a population of the subspecies Mus m. [corrected] musculus. Mitochondrial D-loop sequences from all animals confirm their respective assignments. Approximately 200 microsatellite loci were typed for up to 60 unrelated individuals from each population and evaluated for signs of selective sweeps on the basis of Schlötterer's ln RV and ln RH statistics. Our data suggest that there are slightly more signs of selective sweeps than would have been expected by chance alone in each of the populations and also highlights some of the statistical challenges faced in genome scans for detecting selection. Single-nucleotide polymorphism typing of one sweep signature in the M. m. domesticus populations around the beta-defensin 6 locus confirms a lowered nucleotide diversity in this region and limits the potential sweep region to about 20 kb. However, no amino acid exchange has occurred in the coding region when compared to M. m. musculus. If this sweep signature is due to a recent adaptation, it is expected that a regulatory change would have caused it. Our data provide a framework for conducting a systematic whole genome scan for signatures of selective sweeps in the mouse genome.

Amino Acid Sequence↗

ProTeus: identifying signatures in protein termini.

ProTeus (PROtein TErminUS) is a web-based tool for the identification of short linear signatures in protein termini. It is based on a position-based search method for revealing short signatures in termini of all proteins. The initial step in ProTeus development was to collect all signature groups (SIGs) based on their relative positions at the termini. The initial set of SIGs went through a sequential process of inspection and removal of SIGs, which did not meet the attributed statistical thresholds. The SIGs that were found significant represent protein sets with minimal or no overall sequence similarity besides the similarity found at the termini. These SIGs were archived and are presented at ProTeus. The SIGs are sorted by their strong correspondence to functional annotation from external databases such as GO. ProTeus provides rich search and visualization tools for evaluating the quality of different SIGs. A search option allows the identification of terminal signatures in new sequences. ProTeus (ver 1.2) is available at http://www.proteus.cs.huji.ac.il.

Databases, Protein↗

CRSD: a comprehensive web server for composite regulatory signature discovery.

Transcription factors (TFs) and microRNAs play important roles in the regulation of human gene expression, and the study of their combinatory regulations of gene expression is a new research field. We constructed a comprehensive web server, the composite regulatory signature database (CRSD), that can be applied in investigating complex regulatory behaviors involving gene expression signatures (GESs), microRNA regulatory signatures (MRSs) and TF regulatory signatures (TRSs). Six well-known and large-scale databases, including the human UniGene, mature microRNAs, putative promoter, TRANSFAC, pathway and Gene Ontology (GO) databases, were integrated to provide the comprehensive analysis in CRSD. Two new genome-wide databases, of MRSs and TRSs, were also constructed and further integrated into CRSD. To accomplish the microarray data analysis at one go, several methods, including microarray data pretreatment, statistical and clustering analysis, iterative enrichment analysis and motif discovery, were closely integrated in the web server, which has not been the case in previous studies. Our implementation showed that the published literature could demonstrate the results of genome-wide enrichment analysis. We conclude that CRSD is a powerful and useful bioinformatic web server and may provide new insights into gene regulation networks. CRSD and the online tutorial are publicly available at http://biochip.nchu.edu.tw/crsd1/.

3' Untranslated Regions↗

The use of signature sequences in different proteins to determine the relative branching order of bacterial divisions: evidence that Fibrobacter diverged at a similar time to Chlamydia and the Cytophaga-Flavobacterium-Bacteroides division.

The phylogenetic placement of the rumen bacterium Fibrobacter succinogenes was determined using a signature sequence approach that allows determination of the relative branching order of the major divisions among Bacteria [Gupta, R. S. (2000) FEMS Microbiol Rev 24, 367-402]. For this purpose, segments of the Hsp60 (groEL), Hsp70 (dnaK), CTP synthase and alanyl-tRNA synthetase genes, which are known to contain signature sequences that are useful for phylogenetic deterministic purposes, were cloned. Using degenerate oligonucleotide primers for highly conserved regions in these proteins, 1.4 kb, 0.75 kb, 401 bp and 171 bp fragments of the Hsp70, Hsp60, CTP synthase and alanyl-tRNA synthetase genes respectively were amplified by PCR, and these fragments were cloned and sequenced. These primers, because of their high degree of conservation, could also be used for cloning these genes from other bacterial species. The Hsp70 homologues from different Gram-negative bacteria contain a 21-23 aa insert that is not found in any Gram-positive bacteria. The presence of this insert in the F. succinogenes Hsp70 supports its placement within the Gram-negative group of bacteria. A conserved insert in F. succinogenes Hsp60 that is commonly present in all bacterial species, except various Gram-positive bacteria, Deinococcus-Thermus groups and green non-sulphur bacteria, provides evidence that F. succinogenes does not belong to these taxa. A particularly useful signature consisting of a 4 aa insert is found in Ala-tRNA synthetase. This insert is present in all proteobacterial homologues as well as in homologues from species belonging to the Chlamydia and Cytophaga-Flavobacterium- Bacteroides (CFB) groups, but it is not found in homologues from any other groups of bacteria. The presence of this insert in F. succinogenes Ala-tRNA synthetase provides evidence that this species is related to these groups. However, two other signatures in CTP synthase and Hsp70 proteins, that are distinctive of the proteobacterial species, are not present in the F. succinogenes homologues. These results provide evidence that F. succinogenes does not belong to the proteobacterial division and thus should be placed in a similar position as the Chlamydia and CFB groups of species.

Alanine-tRNA Ligase↗

Molecular signatures in protein sequences that are characteristic of cyanobacteria and plastid homologues.

Fourteen conserved indels (i.e. inserts or deletions) have been identified in 10 widely distributed proteins that appear to be characteristic of cyanobacterial species and are not found in any other group of bacteria. These signatures include three inserts of 6, 7 and 28 aa in the DNA helicase II (UvrD) protein, an 18-21 aa insert in DNA polymerase I, a 14 aa insert in the enzyme ADP-glucose pyrophosphorylase, a 3 aa insert in the FtsH protein, an 11-13 aa insert in phytoene synthase, a 5 aa insert in elongation factor-Tu, two deletions of 2 and 7 aa in ribosomal S1 protein, a 2 aa insert in the SecA protein, a 1 aa deletion and a 6 aa insert in the enzyme inosine-5'-monophosphate dehydrogenase and a 1 aa deletion in the major sigma factor. These signatures, which are flanked by conserved regions, provide molecular markers for distinguishing cyanobacterial taxa from all other bacteria and they should prove helpful in the identification of cyanobacterial species, simply on the basis of the presence or absence of these markers in the corresponding proteins. The signatures in six of these proteins (SecA, elongation factor-Tu, ADP-glucose pyrophosphorylase, phytoene synthase, FtsH and ribosomal S1 protein) are also commonly present in plastid homologues from plants and algae (chlorophytes, chromophytes and rhodophytes), indicating their specific relationship to cyanobacteria and supporting their endosymbiotic origin from these bacteria. In phylogenetic trees based on a number of these proteins (SecA, UvrD, DNA polymerase I, elongation factor-Tu) that were investigated, the available cyanobacterial homologues grouped together with high affinity (>95 % bootstrap value), supporting the view that the cyanobacterial phylum is monophyletic and that the identified signatures were introduced in a common ancestor of this group.

Adenosine Triphosphatases↗

Collateral mutagenesis funnels multiple sources of DNA damage into a ubiquitous mutational signature.

Mutations reflect the net effects of myriad types of damage, replication errors, and repair mechanisms, and thus are expected to differ across cell types with distinct exposures to mutagens, division rates, and cellular programs. Yet when mutations in humans are decomposed into a set of "signatures", one single base substitution signature, SBS5, is present across cell types and tissues, and predominates in post-mitotic neurons as well as male and female germlines [1-3]. The etiology of SBS5 is unknown. By modeling the processes by which mutations arise, we infer that SBS5 is the footprint of errors in DNA synthesis triggered by distinct types of DNA damage. Supporting this hypothesis, we find that SBS5 rates increase with signatures of endogenous and exogenous DNA damage in cancerous and non-cancerous cells and co-vary with repair rates along the genome as expected from model predictions. These analyses indicate that SBS5 captures the output of a "funnel", through which multiple sources of damage result in a similar mutation spectrum. As we further show, SBS5 mutations arise not only from translesion synthesis but also from DNA repair, suggesting that the signature reflects the occasional, shared use of a polymerase.

Journal Article↗