Search PubMed⌕ Search

PubMed · 16563023

Adding value to crystallographically-derived knowledge bases.

Abstract

A protocol for the partially automated computational investigation of crystal structure geometries of transition-metal complexes with unusual/outlier structural features has been developed for application in an e-science context. This protocol not only is envisaged as a part of knowledge base software packages such as Mogul but can also be used to further analyze the results of database searches. The issues arising from automating the initial input generation and DFT optimization of complexes have been examined and a procedure for extracting additional knowledge "value" from the computational results is described. Potential problems/weaknesses arising from the choice of computational approach and from errors in the crystal structure refinement are discussed. A range of likely outcomes of applying this protocol to database mining results is illustrated, with representative examples identified for tetracoordinate transition-metal complexes and ligand fragments (terminal chloride, monodentate phosphorus(III), and primary amine ligands) with unusual metal-ligand bond lengths.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Natalie Fey, Stephanie E Harris, Jeremy N Harvey, A Guy Orpen. Adding value to crystallographically-derived knowledge bases.. https://doi.org/10.1021/ci0504768

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Non-parametric estimation of the case fatality ratio with competing risks data: an application to Severe Acute Respiratory Syndrome (SARS).

For diseases with some level of associated mortality, the case fatality ratio measures the proportion of diseased individuals who die from the disease. In principle, it is straightforward to estimate this quantity from individual follow-up data that provides times from onset to death or recovery. In particular, in a competing risks context, the case fatality ratio is defined by the limiting value of the sub-distribution function, F(1)(t) = Pr(T infinity, where T denotes the time from onset to death (J = 1) or recovery (J = 2). When censoring is present, however, estimation of F(1)(infinity) is complicated by the possibility of little information regarding the right tail of F(1), requiring use of estimators of F(1)(t(*)) or F(1)(t(*))/(F(1)(t(*))+F(2)(t(*))) where t(*) is large, with F(2)(t) = Pr(T <or=t and J = 2) being the analogous sub-distribution function associated with recovery. With right censored data, the variability of such estimators increases as t(*) increases, suggesting the possibility of using estimators at lower values of t(*) where bias may be increased but overall mean squared error be smaller. These issues are investigated here for non-parametric estimators of F(1) and F(2). The ideas are illustrated on case fatality data for individuals infected with Severe Acute Respiratory Syndrome (SARS) in Hong Kong in 2003.

Computer Simulation↗