Search PubMed⌕ Search

Biomedical subjects

Dariusz Plewczynski

Publications and source records attributed to Dariusz Plewczynski.

11 recordsLinked to original sources

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin↗

The transcription factor BACH1 couples chromatin priming and repression to enable macrophage plasticity and adaptation.

Macrophage activation and tissue adaptation involve precise transcriptional control by lineage-determining transcription factors (LDTFs) and stimulus-dependent TFs. The heme-regulated transcriptional repressor BACH1 clusters with myeloid LDTFs in unstimulated macrophages, suggesting a role in shaping macrophage identity and function. We found that BACH1 bound to both inactive and active regulatory regions, including latent enhancers. BACH1 recruited the NuRD complex and had dual functions, establishing early chromatin accessibility while actively repressing transcription. Upon inflammatory stimulation, BACH1 rapidly redistributed in cis to nearby promoters, reshaping chromatin occupancy, motif specificity, and enhancer-promoter interactions. BACH1 constrained 3D chromatin architecture, limiting enhancer mobility and TF complex dynamics. In vivo, Bach1 deletion impaired macrophage polarization and tissue adaptation and limited resilience during systemic and regenerative inflammation. Thus, BACH1 acts as an early chromatin accessibility-priming factor while actively repressing transcription-a regulatory activity that can be defined as pioneer repression-thereby shaping the macrophage epigenome in response to inflammatory and tissue contexts.

Basic-Leucine Zipper Transcription Factors↗

The challenge of chromatin model comparison and validation: A project from the first international 4D Nucleome Hackathon.

The computational modeling of chromatin structure is highly complex due to the hierarchical organization of chromatin, which reflects its diverse biophysical principles, as well as inherent dynamism, which underlies its complexity. Chromatin structure modeling can be based on diverse approaches and assumptions, making it essential to determine how different methods influence the modeling outcomes. We conducted a project at the NIH-funded 4D Nucleome Hackathon on March 18-21, 2024, at The University of Washington in Seattle, USA. The hackathon provided an amazing opportunity to gather an international, multi-institutional and unbiased group of experts to discuss, understand and undertake the challenges of chromatin model comparison and validation. Here we give an overview of the current state of the 3D chromatin field and discuss our efforts to run and validate the models. We used distance matrices to represent chromatin models and we calculated Spearman correlation coefficients to estimate differences between models, as well as between models and experimental data. In addition, we discuss challenges in chromatin structure modeling that include: 1) different aspects of chromatin biophysics and scales complicate model comparisons, 2) large diversity of experimental data (e.g., population-based, single-cell, protein-specific) that differ in mathematical properties, heatmap smoothness, noise and resolutions complicates model validation, 3) expertise in biology, bioinformatics, and physics is necessary to conduct comprehensive research on chromatin structure, 4) bioinformatic software, which is often developed in academic settings, is characterized by insufficient support and documentation. We also emphasize the importance of establishing guidelines for software development and standardization.

Chromatin↗

Improved cohesin HiChIP protocol and bioinformatic analysis for robust detection of chromatin loops and stripes.

Chromosome Conformation Capture (3 C) methods, including Hi-C (a high-throughput variation of 3 C), detect pairwise interactions between DNA regions, enabling the reconstruction of chromatin architecture in the nucleus. HiChIP is a modification of the Hi-C experiment that includes a chromatin immunoprecipitation (ChIP) step, allowing genome-wide identification of chromatin contacts mediated by a protein of interest. In mammalian cells, cohesin protein complex is one of the major players in the establishment of chromatin loops. We present an improved cohesin HiChIP experimental protocol. Using comprehensive bioinformatic analysis, we show that a dual chromatin fixation method compared to the standard formaldehyde-only method, results in a substantially better signal-to-noise ratio, increased ChIP efficiency and improved detection of chromatin loops and architectural stripes. Additionally, we propose an automated pipeline called nf-HiChIP ( https://github.com/SFGLab/hichip-nf-pipeline ) for processing HiChIP samples starting from raw sequencing reads data and ending with a set of significant chromatin interactions (loops), which allows efficient and timely analysis of multiple samples in parallel, without requiring additional ChIP-seq experiments. Finally, using advanced approaches for biophysical modelling and stripe calling we generate accurate loop extrusion polymer models for a region of interest and provide a detailed picture of architectural stripes, respectively.

Chromatin↗

PDB-UF: database of predicted enzymatic functions for unannotated protein structures from structural genomics.

BACKGROUND: The number of protein structures from structural genomics centers dramatically increases in the Protein Data Bank (PDB). Many of these structures are functionally unannotated because they have no sequence similarity to proteins of known function. However, it is possible to successfully infer function using only structural similarity. RESULTS: Here we present the PDB-UF database, a web-accessible collection of predictions of enzymatic properties using structure-function relationship. The assignments were conducted for three-dimensional protein structures of unknown function that come from structural genomics initiatives. We show that 4 hypothetical proteins (with PDB accession codes: 1VH0, 1NS5, 1O6D, and 1TO0), for which standard BLAST tools such as PSI-BLAST or RPS-BLAST failed to assign any function, are probably methyltransferase enzymes. CONCLUSION: We suggest that the structure-based prediction of an EC number should be conducted having the different similarity score cutoff for different protein folds. Moreover, performing the annotation using two different algorithms can reduce the rate of false positive assignments. We believe, that the presented web-based repository will help to decrease the number of protein structures that have functions marked as "unknown" in the PDB file. AVAILABILITY: http://paradox.harvard.edu/PDB-UF and http://bioinfo.pl/PDB-UF.

Chromosome Mapping↗

Support-vector-machine classification of linear functional motifs in proteins.

Our algorithm predicts short linear functional motifs in proteins using only sequence information. Statistical models for short linear functional motifs in proteins are built using the database of short sequence fragments taken from proteins in the current release of the Swiss-Prot database. Those segments are confirmed by experiments to have single-residue post-translational modification. The sensitivities of the classification for various types of short linear motifs are in the range of 70%. The query protein sequence is dissected into short overlapping fragments. All segments are represented as vectors. Each vector is then classified by a machine learning algorithm (Support Vector Machine) as potentially modifiable or not. The resulting list of plausible post-translational sites in the query protein is returned to the user. We also present a study of the human protein kinase C family as a biological application of our method.

Databases, Genetic↗

Molecular modeling of phosphorylation sites in proteins using a database of local structure segments.

A new bioinformatics tool for molecular modeling of the local structure around phosphorylation sites in proteins has been developed. Our method is based on a library of short sequence and structure motifs. The basic structural elements to be predicted are local structure segments (LSSs). This enables us to avoid the problem of non-exact local description of structures, caused by either diversity in the structural context, or uncertainties in prediction methods. We have developed a library of LSSs and a profile--profile-matching algorithm that predicts local structures of proteins from their sequence information. Our fragment library prediction method is publicly available on a server (FRAGlib), at http://ffas.ljcrf.edu/Servers/frag.html . The algorithm has been applied successfully to the characterization of local structure around phosphorylation sites in proteins. Our computational predictions of sequence and structure preferences around phosphorylated residues have been confirmed by phosphorylation experiments for PKA and PKC kinases. The quality of predictions has been evaluated with several independent statistical tests. We have observed a significant improvement in the accuracy of predictions by incorporating structural information into the description of the neighborhood of the phosphorylated site. Our results strongly suggest that sequence information ought to be supplemented with additional structural context information (predicted with our segment similarity method) for more successful predictions of phosphorylation sites in proteins.

Amino Acid Sequence↗

AutoMotif server: prediction of single residue post-translational modifications in proteins.

UNLABELLED: The AutoMotif Server allows for identification of post-translational modification (PTM) sites in proteins based only on local sequence information. The local sequence preferences of short segments around PTM residues are described here as linear functional motifs (LFMs). Sequence models for all types of PTMs are trained by support vector machine on short-sequence fragments of proteins in the current release of Swiss-Prot database (phosphorylation by various protein kinases, sulfation, acetylation, methylation, amidation, etc.). The accuracy of the identification is estimated using the standard leave-one-out procedure. The sensitivities for all types of short LFMs are in the range of 70%. AVAILABILITY: The AutoMotif Server is available free for academic use at http://automotif.bioinfo.pl/

Algorithms↗

Integrated web service for improving alignment quality based on segments comparison.

BACKGROUND: Defining blocks forming the global protein structure on the basis of local structural regularity is a very fruitful idea, extensively used in description, and prediction of structure from only sequence information. Over many years the secondary structure elements were used as available building blocks with great success. Specially prepared sets of possible structural motifs can be used to describe similarity between very distant, non-homologous proteins. The reason for utilizing the structural information in the description of proteins is straightforward. Structural comparison is able to detect approximately twice as many distant relationships as sequence comparison at the same error rate. RESULTS: Here we provide a new fragment library for Local Structure Segment (LSS) prediction called FRAGlib which is integrated with a previously described segment alignment algorithm SEA. A joined FRAGlib/SEA server provides easy access to both algorithms, allowing a one stop alignment service using a novel approach to protein sequence alignment based on a network matching approach. The FRAGlib used as secondary structure prediction achieves only 73% accuracy in Q3 measure, but when combined with the SEA alignment, it achieves a significant improvement in pairwise sequence alignment quality, as compared to previous SEA implementation and other public alignment algorithms. The FRAGlib algorithm takes approximately 2 min. to search over FRAGlib database for a typical query protein with 500 residues. The SEA service align two typical proteins within circa approximately 5 min. All supplementary materials (detailed results of all the benchmarks, the list of test proteins and the whole fragments library) are available for download on-line at http://ffas.ljcrf.edu/darman/results/. CONCLUSIONS: The joined FRAGlib/SEA server will be a valuable tool both for molecular biologists working on protein sequence analysis and for bioinformaticians developing computational methods of structure prediction and alignment of proteins.

Computational Biology↗

Comparison of proteins based on segments structural similarity.

We present here a simple method for fast and accurate comparison of proteins using their structures. The algorithm is based on structural alignment of segments of Calpha chains (with size of 99 or 199 residues). The method is optimized in terms of speed and accuracy. We test it on 97 representative proteins with the similarity measure based on the SCOP classification. We compare our algorithm with the LGscore2 automatic method. Our method has the same accuracy as the LGscore2 algorithm with much faster processing of the whole test set, which is promising. A second test is done using the ToolShop structure prediction evaluation program and shows that our tool is on average slightly less sensitive than the DALI server. Both algorithms give a similar number of correct models, however, the final alignment quality is better in the case of DALI. Our method was implemented under the name 3D-Hit as a web server at http://3dhit.bioinfo.pl/ free for academic use, with a weekly updated database containing a set of 5000 structures from the Protein Data Bank with non-homologous sequences.

Algorithms↗

Assessing different classification methods for virtual screening.

How well do different classification methods perform in selecting the ligands of a protein target out of large compound collections not used to train the model? Support vector machines, random forest, artificial neural networks, k-nearest-neighbor classification with genetic-algorithm-optimized feature selection, trend vectors, naïve Bayesian classification, and decision tree were used to divide databases into molecules predicted to be active and those predicted to be inactive. Training and predicted activities were treated as binary. The database was generated for the ligands of five different biological targets which have been the object of intense drug discovery efforts: HIV-reverse transcriptase, COX2, dihydrofolate reductase, estrogen receptor, and thrombin. We report significant differences in the performance of the methods independent of the biological target and compound class. Different methods can have different applications; some provide particularly high enrichment, others are strong in retrieving the maximum number of actives. We also show that these methods do surprisingly well in predicting recently published ligands of a target on the basis of initial leads and that a combination of the results of different methods in certain cases can improve results compared to the most consistent method.

Algorithms↗