Search PubMed⌕ Search

Biomedical subjects

Stephanie E Harris

Publications and source records attributed to Stephanie E Harris.

4 recordsLinked to original sources

Development of a ligand knowledge base, part 1: computational descriptors for phosphorus donor ligands.

A prototype collection of knowledge on ligands in metal complexes, termed a ligand knowledge base (LKB), has been developed. This contribution describes the design of DFT-calculated descriptors for monodentate phosphorus(III) donor ligands in a range of representative complexes. Using the resulting data, a ligand space is mapped and predictive models are derived for metal complexes. Important characteristics, including chemical, computational and statistical robustness for the generation and exploitation of such an LKB are described. Chemical robustness ensures transferability of the descriptors, as well as comprehensive sampling of ligand space. To make the calculations amenable to automation in an e-science setting, a reliable, well-defined computational approach has been sought from which the descriptors can be readily extracted. The LKB has been explored with multivariate statistical methods. Principal component analysis (PCA) is used for the mapping of chemical space, projecting multiple descriptors into scatter plots which illustrate the clustering of chemically similar ligands. Interpretation of the resulting principal components in terms of established steric and electronic properties and the importance of its statistical robustness to variations in the ligand set are discussed. Multiple linear regression (MLR) models have been derived, demonstrating the versatility of the descriptors for modeling varied experimentally determined parameters (bond lengths, reaction enthalpies and bond-stretching frequencies). The importance of re-sampling methods for testing the robustness of predictions is highlighted. A strategy for the construction of a robust LKB suitable for the modeling of ligand and complex behavior is outlined based on these observations.

Journal Article↗

Retrieval of crystallographically-derived molecular geometry information.

The crystallographically determined bond length, valence angle, and torsion angle information in the Cambridge Structural Database (CSD) has many uses. However, accessing it by means of conventional substructure searching requires nontrivial user intervention. In consequence, these valuable data have been underutilized and have not been directly accessible to client applications. The situation has been remedied by development of a new program (Mogul) for automated retrieval of molecular geometry data from the CSD. The program uses a system of keys to encode the chemical environments of fragments (bonds, valence angles, and acyclic torsions) from CSD structures. Fragments with identical keys are deemed to be chemically identical and are grouped together, and the distribution of the appropriate geometrical parameter (bond length, valence angle, or torsion angle) is computed and stored. Use of a search tree indexed on key values, together with a novel similarity calculation, then enables the distribution matching any given query fragment (or the distributions most closely matching, if an adequate exact match is unavailable) to be found easily and with no user intervention. Validation experiments indicate that, with rare exceptions, search results afford precise and unbiased estimates of molecular geometrical preferences. Such estimates may be used, for example, to validate the geometries of libraries of modeled molecules or of newly determined crystal structures or to assist structure solution from low-resolution (e.g. powder diffraction) X-ray data.

Journal Article↗

Factors affecting d-block metal-ligand bond lengths: toward an automated library of molecular geometry for metal complexes.

Metal-ligand (M-L) bond lengths for a range of ligands (carboxylates, chlorides, pyridines, water, tertiary phosphines, and alkenes) and a variety of metals have been retrieved from the Cambridge Structural Database, CSD. Analysis of the factors which affect M-L bond lengths (for example, ligand coordination mode, oxidation state, metal coordination number and geometry, spin and Jahn-Teller effects, and ligand trans to M-L bond) shows that it is generally possible to subdivide the M-L data sets systematically to obtain better defined, unimodal, bond length distributions with means and sample standard deviations (SSDs) which reflect the nature of the bond in question. Typically, the SSDs for the M-L data sets can be reduced to 0.04-0.05 A by these methods. This work is an extension to tables of bond lengths in organometallic compounds and coordination complexes published in 1989. The importance of the factors which affect M-L bond lengths for particular metal-ligand groups are discussed. From the case studies reported, an algorithm is proposed by which compilation of a library of molecular geometry for metal complexes may be automated. The points that need to be considered to produce such a molecular library from the data stored in the CSD are discussed. The development of such a library would allow users to retrieve chemically well-defined geometric data rapidly and accurately. This should be of use, for example, to crystallographers and molecular modelers.

Journal Article↗

Adding value to crystallographically-derived knowledge bases.

A protocol for the partially automated computational investigation of crystal structure geometries of transition-metal complexes with unusual/outlier structural features has been developed for application in an e-science context. This protocol not only is envisaged as a part of knowledge base software packages such as Mogul but can also be used to further analyze the results of database searches. The issues arising from automating the initial input generation and DFT optimization of complexes have been examined and a procedure for extracting additional knowledge "value" from the computational results is described. Potential problems/weaknesses arising from the choice of computational approach and from errors in the crystal structure refinement are discussed. A range of likely outcomes of applying this protocol to database mining results is illustrated, with representative examples identified for tetracoordinate transition-metal complexes and ligand fragments (terminal chloride, monodentate phosphorus(III), and primary amine ligands) with unusual metal-ligand bond lengths.

Computer Simulation↗