Search PubMed⌕ Search

PubMed · 14630660

PDB file parser and structure class implemented in Python.

Abstract

UNLABELLED: The biopython project provides a set of bioinformatics tools implemented in Python. Recently, biopython was extended with a set of modules that deal with macromolecular structure. Biopython now contains a parser for PDB files that makes the atomic information available in an easy-to-use but powerful data structure. The parser and data structure deal with features that are often left out or handled inadequately by other packages, e.g. atom and residue disorder (if point mutants are present in the crystal), anisotropic B factors, multiple models and insertion codes. In addition, the parser performs some sanity checking to detect obvious errors. AVAILABILITY: The Biopython distribution (including source code and documentation) is freely available (under the Biopython license) from http://www.biopython.org

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thomas Hamelryck, Bernard Manderick. 2003-11-22. PDB file parser and structure class implemented in Python.. https://doi.org/10.1093/bioinformatics%2Fbtg299

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Exploring peptide energy landscapes: a test of force fields and implicit solvent models.

A biased Monte Carlo-minimization/annealing conformational search was used to characterize five descriptions of the energy landscape for each of three model systems: the 20-residue "trp-cage" miniprotein, the 20-residue "BS1" peptide, and the 17-residue "U(1-17)T9D" peptide. The EEF1 and SASA energy landscapes were studied as well as those defined by using the GB/ACE implicit water model with one of three protein force fields: CHARMM19, CHARMM22, and CHARMM22/CMAP. The lowest-energy structures of the trp-cage and BS1 peptides found for the EEF1 landscape have main-chain root-mean-square deviations (rmsds) from the respective NMR structures of less than 2 A; for U(1-17)T9D, the deviation is less than 3 A using EEF1. The main-chain rmsd of the minimum-energy trp-cage conformation obtained for the GB/ACE/CHARMM22/CMAP landscape is less than 1 A. However, this energy function strongly favored helical structures for the two peptides shown by NMR to form beta-sheet structures. Brief annealing of the system following main-chain conformational changes was found to enhance the exploration of low-energy states. The thousands of simulations reported here suggest that the prediction of protein structure might be improved by the simultaneous use of a CMAP-like description of the main chain and an EEF1-like description of the solvent.

Computer Simulation↗