Search PubMed⌕ Search

Biomedical subjects

T R Ioerger

Publications and source records attributed to T R Ioerger.

6 recordsLinked to original sources

Determining protein structure from electron-density maps using pattern matching.

TEXTAL is an automated system for building protein structures from electron-density maps. It uses pattern recognition to select regions in a database of previously determined structures that are similar to regions in a map of unknown structure. Rotation-invariant numerical values, called features, of the electron density are extracted from spherical regions in an unknown map and compared with features extracted around regions in maps generated from a database of known structures. Those regions in the database that match best provide the local coordinates of atoms and these are accumulated to form a model of the unknown structure. Similarity between the regions in the database and an uninterpreted region is determined firstly by evaluating the numerical difference in feature values and secondly by calculating the electron-density correlation coefficient for those regions with similar feature values. TEXTAL has been successful at building protein structures for a wide range of test electron-density maps and can automatically model entire protein structures in a few hours on a workstation. Models built by TEXTAL from test electron-density maps of known protein structures were accurate to within 0.6-0.7 A root-mean-square deviation, assuming prior knowledge of C(alpha) positions. The system represents a new approach to protein structure determination and has the potential to greatly reduce the time required to interpret electron-density maps in order to build accurate protein models.

Algorithms↗

Conservation of cys-cys trp structural triads and their geometry in the protein domains of immunoglobulin superfamily members.

In almost all members of the immunoglobulin superfamily (IgSF) for which an experimental structure has been determined, a triad (C-CW) consisting of two cysteine residues that form a disulfide bond and a neighboring tryptophan can be found in the core of the protein fold. We analyzed the geometry of these C-CW triads among a database of 60 Fab crystal structures and found it to be remarkably conserved. We identified C-CW triads of a similar configuration in other members of the IgSF such as T cell receptor (TCR), major histocompatibility complex antigens (MHC), cell surface antigens CD4 and CD8, and cell-adhesion molecules. We used this C-CW pattern to search a database of non-IgSF proteins, and identified several proteins that contain a disulfide bridge associated with a tryptophan in a similar configuration. Examination of the distances and orientations between triads found in adjacent domains in Fab fragments and TCR also reveal a high degree of conservation, which reflects the invariance of the inter-chain domain packing. This high degree of conservation of the geometry of the C-CW triad in IgSF structures suggests that the Trp may contribute significantly to the stability of the disulfide bond. Knowledge of these geometric parameters may prove useful in the construction and validation of theoretical models of Ig, TCR, and other IgSF members.

Animals↗

TEXTAL: a pattern recognition system for interpreting electron density maps.

X-ray crystallography is the most widely used method for determining the three-dimensional structures of proteins and other macromolecules. One of the most difficult steps in crystallography is interpreting the electron density map to build the final model. This is often done manually by crystallographers and is very time-consuming and error-prone. In this paper, we introduce a new automated system called TEXTAL for interpreting electron density maps using pattern recognition. Given a map to be modeled, TEXTAL divides the map into small regions and then finds regions with a similar pattern of density in a database of maps for proteins whose structures have already been solved. When a match is found, the coordinates of atoms in the region are inferred by analogy. The key to making the database lookup efficient is to extract numeric features that represent the patterns in each region and to compare feature values using a weighted Euclidean distance metric. It is crucial that the features be rotation-invariant, since regions with similar patterns of density can be oriented in any arbitrary way. This pattern-recognition approach can take advantage of data accumulated in large crystallographic databases to effectively learn the association between electron density and molecular structure by example.

Algorithms↗

The context-dependence of amino acid properties.

One of the current limitations of using sequence alignments to identify proteins with similar structures is that some proteins with similar structures do not have significant sequence similarity by identity. One way to address this "hidden-homology" problem is to match amino acids based on their chemical and physical properties. However, the amino acid properties overlap, creating orthogonal dimensions of similarity, the relative strengths of which are ambiguous. It has been observed that the role an amino acid plays (and hence the property that is important) at a site in a protein depends on its secondary and tertiary environment. To approximate and take advantage of this dependence on context for improving the sensitivity of alignments of proteins whose structures are unknown, we propose a surrogate definition of context based on the pattern of hydropathy in a small window of contiguous neighbors surrounding each amino acid. We present the results of an experiment in which a search-based program iteratively tests and selects various properties in independent contexts, and incrementally increases the ability of sequence alignments to detect relationships among distantly-related proteins. The method is shown to perform better than using the MDM78 substitution table for partial match scores.

Amino Acid Sequence↗

Constructive induction and protein tertiary structure prediction.

To date, the only methods that have been used successfully to predict protein structures have been based on identifying homologous proteins whose structures are known. However, such methods are limited by the fact that some proteins have similar structure but no significant sequence homology. We consider two ways of applying machine learning to facilitate protein structure prediction. We argue that a straightforward approach will not be able to improve the accuracy of classification achieved by clustering by alignment scores alone. In contrast, we present a novel constructive induction approach that learns better representations of amino acid sequences in terms of physical and chemical properties. Our learning method combines knowledge and search to shift the representation of sequences so that semantic similarity is more easily recognized by syntactic matching. Our approach promises not only to find new structural relationships among protein sequences, but also expands our understanding of the roles knowledge can play in learning via experience in this challenging domain.

Artificial Intelligence↗

Polymorphism at the self-incompatibility locus in Solanaceae predates speciation.

Sequences of 11 alleles of the gametophytic self-incompatibility locus (S locus) from three species of the Solanaceae family have recently been determined. Pairwise comparisons of these alleles reveal two unexpected observations: (i) amino acid sequence similarity can be as low as 40% within species and (ii) some interspecific similarities are higher than intraspecific similarities. The gene genealogy clearly illustrates this unusual pattern of relationships. The data suggest that some of the polymorphism at the S locus existed prior to the divergence of these species and has been maintained to the present. In support of this hypothesis, the number of shared polymorphic sites was found to exceed the number found in simulations with independent accumulation of mutations. Strictly neutral evolution is exceedingly unlikely to maintain the polymorphism for such a long time. The allele multiplicity and extreme age of the alleles is consistent with Wright's classic one-locus population genetic model of gametophytic self-incompatibility. Similarities between the plant S locus and the mammalian major histocompatibility complex are discussed.

Alleles↗