Search PubMed⌕ Search

Biomedical subjects

Eleanor J Gardiner

Publications and source records attributed to Eleanor J Gardiner.

9 recordsLinked to original sources

A structural similarity analysis of double-helical DNA.

A database of the structural properties of all 32,896 unique DNA octamer sequences has been calculated, including information on stability, the minimum energy conformation and flexibility. The contents of the database have been analysed using a variety of Euclidean distance similarity measures. A global comparison of sequence similarity with structural similarity shows that the structural properties of DNA are much less diverse than the sequences, and that DNA sequence space is larger and more diverse than DNA structure space. Thus, there are many very different sequences that have very similar structural properties, and this may be useful for identifying DNA motifs that have similar functional properties that are not apparent from the sequences. On the other hand, there are also small numbers of almost identical sequences that have very different structural properties, and these could give rise to false-positives in methods used to identify function based on sequence alignment. A simple validation test demonstrates that structural similarity can differentiate between promoter and non-promoter DNA. Combining structural and sequence similarity improves promoter recall beyond that possible using either similarity measure alone, demonstrating that there is indeed information available in the structure of double-helical DNA that is not readily apparent from the sequence.

DNA↗

Sequence-dependent DNA structure: a database of octamer structural parameters.

We have constructed the potential energy surfaces for all unique tetramers, hexamers and octamers in double helical DNA, as a function of the two principal degrees of freedom, slide and shift at the central step. From these potential energy maps, we have calculated a database of structural and flexibility properties for each of these sequences. These properties include: the values of each of the six step parameters (twist roll, tilt, rise, slide and shift), for each step of the sequence; flexibility measures for both decrease and increase in each property value from the minimum energy conformation for the central step; and the deviation from the path of a hypothetical straight octamer. In an analysis of structural change as a function of sequence length, we observe that almost all DNA tends to B-DNA and becomes less flexible. A more detailed analysis of octamer properties has allowed us to determine the structural preferences of particular sequence elements. GGC and GCC sequences tend to confer bistability, low stability and a predisposition to A-form DNA, whereas AA steps strongly prefer B-DNA and inhibit A-structures. There is no correlation between flexibility and intrinsic curvature, but bent DNA is less stable than straight. The most difficult deformation is undertwisting. The TA step stands out as the most flexible sequence element with respect to decreasing twist and increasing roll. However, as with the structural properties, this behavior is highly context-dependent and some TA steps are very straight.

DNA↗

GAPDOCK: a Genetic Algorithm Approach to Protein Docking in CAPRI round 1.

As part of the first Critical Assessment of PRotein Interactions, round 1, we predict the structure of two protein-protein complexes, by using a genetic algorithm, GAPDOCK, in combination with surface complementarity, buried surface area, biochemical information, and human intervention. Among the five models submitted for target 1, HPr phosphocarrier protein (B. subtilis) and the hexameric HPr kinase (L. lactis), the best correctly predicts 17 of 52 interprotein contacts, whereas for target 2, bovine rotavirus VP6 protein-monoclonal antibody, the best model predicts 27 of 52 correct contacts. Given the difficult nature of the targets, these predictions are very encouraging and compare well with those obtained by other methods. Nevertheless, it is clear that there is a need for improved methods for distinguishing between "correct" and "plausible but incorrect" complexes.

Algorithms↗

Heuristics for similarity searching of chemical graphs using a maximum common edge subgraph algorithm.

Recently a method (RASCAL) for determining graph similarity using a maximum common edge subgraph algorithm has been proposed which has proven to be very efficient when used to calculate the relative similarity of chemical structures represented as graphs. This paper describes heuristics which simplify a RASCAL similarity calculation by taking advantage of certain properties specific to chemical graph representations of molecular structure. These heuristics are shown experimentally to increase the efficiency of the algorithm, especially at more distant values of chemical graph similarity.

Journal Article↗

Further development of reduced graphs for identifying bioactive compounds.

Reduced graphs provide summary representations of chemical structures. Here, a variety of different types of reduced graphs are compared in similarity searches. The reduced graphs are found to give comparable performance to Daylight fingerprints in terms of the number of active compounds retrieved. However, no one type of reduced graph is found to be consistently superior across a variety of different data sets. Consequently, a representative set of reduced graphs was chosen and used together with Daylight fingerprints in data fusion experiments. The results show improved performance in 10 out of 11 data sets compared to using Daylight fingerprints alone. Finally, the potential of using reduced graphs to build SAR models is demonstrated using recursive partitioning. An SAR model consistent with a published model is found following just two splits in the decision tree.

Data Display↗

A fourier fingerprint-based method for protein surface representation.

A crucial enabling technology for structural genomics is the development of algorithms that can predict the putative function of novel protein structures: the proposed functions can subsequently be experimentally tested by functional studies. Testable assignments of function can be made if it is possible to attribute a putative, or indeed probable, function on the basis of the shapes of the binding sites on the surface of a protein structure. However the comparison of the surfaces of 3D protein structures is a computationally demanding task. Here we present four surface representations that can be used locally to describe the global shape of specifically bounded local region models. The most successful of these representations is obtained by a Fourier analysis of the distribution of surface curvature on concentric spheres around a surface point and summarizes a 24 A diameter spherically clipped region of protein surface by a fingerprint of 18 Fourier amplitude values. Searching experiments using these fingerprints on a set of 366 proteins demonstrate that this provides an effective and an efficient technique for the matching of protein surfaces.

Fourier Analysis↗

Scaffold hopping using clique detection applied to reduced graphs.

Similarity-based methods for virtual screening are widely used. However, conventional searching using 2D chemical fingerprints or 2D graphs may retrieve only compounds which are structurally very similar to the original target molecule. Of particular current interest then is scaffold hopping, that is, the ability to identify molecules that belong to different chemical series but which could form the same interactions with a receptor. Reduced graphs provide summary representations of chemical structures and, therefore, offer the potential to retrieve compounds that are similar in terms of their gross features rather than at the atom-bond level. Using only a fingerprint representation of such graphs, we have previously shown that actives retrieved were more diverse than those found using Daylight fingerprints. Maximum common substructures give an intuitively reasonable view of the similarity between two molecules. However, their calculation using graph-matching techniques is too time-consuming for use in practical similarity searching in larger data sets. In this work, we exploit the low cardinality of the reduced graph in graph-based similarity searching. We reinterpret the reduced graph as a fully connected graph using the bond-distance information of the original graph. We describe searches, using both the maximum common induced subgraph and maximum common edge subgraph formulations, on the fully connected reduced graphs and compare the results with those obtained using both conventional chemical and reduced graph fingerprints. We show that graph matching using fully connected reduced graphs is an effective retrieval method and that the actives retrieved are likely to be topologically different from those retrieved using conventional 2D methods.

Algorithms↗

Genomic data analysis using DNA structure: an analysis of conserved nongenic sequences and ultraconserved elements.

Recent comparative studies of the human and mouse genomes have revealed sets of conserved nongenic sequences (CNGs) and sets of ultraconserved elements (UCEs). Both sets of sequences, which exhibit extremely high levels of conservation, extend over hundreds of bases and have no known function. Since there is no detectable sequence homology between paralogous CNGs or UCEs in either of the species, an alignment-free technique is needed for their analysis. We have previously compiled a database of the structural properties of all 32,896 unique DNA octamers, including information on stability, the minimum energy conformation, and flexibility. We have used Fourier techniques to analyze the UCEs and CNGs in terms of their octamer structural properties, to reveal structural correlations which may indicate possible functions for some of these sequences.

Animals↗

Structural DNA profiles: single sequence queries.

Structural DNA profiles use the structural properties of the constituent octamers either to observe any characteristics of a single sequence that are unusual (a single sequence query) or to visualize a pattern common to a set of sequences (a multiple sequence query). They are an aid in understanding structural reasons for functional DNA activity. Profiles that answer single sequence queries are introduced and Profile Manager (a software application developed to automate profile generation) is presented. Two sequences that are similar by their nucleotide composition but are known to be very different by structure are analyzed, resulting in useful illustrations that agree with the experimental nuclear magnetic resonance structures.

Animals↗