Search PubMed⌕ Search

Biomedical subjects

Arthur M Lesk

Publications and source records attributed to Arthur M Lesk.

12 recordsLinked to original sources

Contact patterns between helices and strands of sheet define protein folding patterns.

Comparing and classifying protein folding patterns allows organizing the known structures and enumerating possible protein structural patterns including those not yet observed. We capture the essence of protein folding patterns in a concise tableau representation based on the order and contact patterns of secondary structures: helices and strands of sheet. The tableaux are intelligible to both humans and computers. They provide a database, derived from the Protein Data Bank, mineable in studies of protein architecture. Using this database, we have: (i) determined statistical properties of secondary structure contacts in an unbiased set of protein domains from ASTRAL, (ii) observed that in 98% of cases, the tableau is a faithful representation of the folding pattern as classified in SCOP, (iii) demonstrated that to a large extent the local structure of proteins indicates their complete folding topology, and (iv) studied the use of the representation for fold identification.

Databases, Protein↗

MUSTANG: a multiple structural alignment algorithm.

Multiple structural alignment is a fundamental problem in structural genomics. In this article, we define a reliable and robust algorithm, MUSTANG (MUltiple STructural AligNment AlGorithm), for the alignment of multiple protein structures. Given a set of protein structures, the program constructs a multiple alignment using the spatial information of the C(alpha) atoms in the set. Broadly based on the progressive pairwise heuristic, this algorithm gains accuracy through novel and effective refinement phases. MUSTANG reports the multiple sequence alignment and the corresponding superposition of structures. Alignments generated by MUSTANG are compared with several handcurated alignments in the literature as well as with the benchmark alignments of 1033 alignment families from the HOMSTRAD database. The performance of MUSTANG was compared with DALI at a pairwise level, and with other multiple structural alignment tools such as POSA, CE-MC, MALECON, and MultiProt. MUSTANG performs comparably to popular pairwise and multiple structural alignment tools for closely related proteins, and performs more reliably than other multiple structural alignment methods on hard data sets containing distantly related proteins or proteins that show conformational changes.

Algorithms↗

What determines the spectrum of protein native state structures?

We present a brief summary of the key factors underlying protein structure, as developed in the investigations of Pauling, Ramachandran, and Rose. We then outline a simplified physical model of proteins that focuses on geometry and symmetry. Although this model superficially appears unrelated to the detailed chemical descriptions commonly applied to proteins, we show that it captures the essential elements of the chemistry and provides a unified framework for understanding the common characteristics of folded proteins. We suggest that the spectrum of protein native state structures is determined by geometry and symmetry and the role of the sequence is to choose its native state structure from this predetermined menu.

Humans↗

Computational study of the fibril organization of polyglutamine repeats reveals a common motif identified in beta-helices.

The formation of fibril aggregates by long polyglutamine sequences is assumed to play a major role in neurodegenerative diseases such as Huntington. Here, we model peptides rich in glutamine, through a series of molecular dynamics simulations. Starting from a rigid nanotube-like conformation, we have obtained a new conformational template that shares structural features of a tubular helix and of a beta-helix conformational organization. Our new model can be described as a super-helical arrangement of flat beta-sheet segments linked by planar turns or bends. Interestingly, our comprehensive analysis of the Protein Data Bank reveals that this is a common motif in beta-helices (termed beta-bend), although it has not been identified so far. The motif is based on the alternation of beta-sheet and helical conformation as the protein sequence is followed from the N to the C termini (beta-alpha(R)-beta-polyPro-beta). We further identify this motif in the ssNMR structure of the protofibril of the amyloidogenic peptide Abeta(1-40). The recurrence of the beta-bend suggests a general mode of connecting long parallel beta-sheet segments that would allow the growth of partially ordered fibril structures. The design allows the peptide backbone to change direction with a minimal loss of main chain hydrogen bonds. The identification of a coherent organization beyond that of the beta-sheet segments in different folds rich in parallel beta-sheets suggests a higher degree of ordered structure in protein fibrils, in agreement with their low solubility and dense molecular packing.

Amino Acid Motifs↗

Molecular forces in antibody maturation.

Analysis of x-ray crystal structures has clarified the nature of antibody-antigen interactions, and the conformational basis of specificity and affinity, but does not provide a clear picture of the dynamics of antigen recognition. In particular, we know that primary antibodies can bind a wider variety of ligands than their secondary counterparts--which are tuned for high specificity and affinity. Crystal structures show that in the absence of antigen the secondary antibody adopts a structure preformed for binding, but that the primary antibody does not. Our calculations show that the unligated state of the primary antibody has a well-defined structure, fluctuating no more widely than that of the secondary antibody, and undergoes a discrete structural rearrangement in response to ligand.

Antibodies↗

Structural divergence and distant relationships in proteins: evolution of the globins.

The globin family has long been known from studies of approximately 150-residue proteins such as vertebrate myoglobins and haemoglobins. Recently, this family has been enriched by the investigation of the sequences and structures of truncated globins, which have the same basic topology but are approximately 30 residues shorter and exhibit functions other than the familiar one of binding diatomic ligands. The divergence of protein sequences, structures and functions reveals Nature's exploration of the potential inherent in a folding pattern, that is, the topology of the native structure. The observation of what remains constant and what varies during the evolution of a protein family reveals essential features of structure and function. Study of proteins with a wide range of divergence can therefore sharpen our understanding of how different amino acid sequences can determine similar three-dimensional structures. Globins have provided, and continue to provide, interesting material for such studies.

Amino Acid Sequence↗

Functional insights from the distribution and role of homopeptide repeat-containing proteins.

Expansion of "low complex" repeats of amino acids such as glutamine (Poly-Q) is associated with protein misfolding and the development of degenerative diseases such as Huntington's disease. The mechanism by which such regions promote misfolding remains controversial, the function of many repeat-containing proteins (RCPs) remains obscure, and the role (if any) of repeat regions remains to be determined. Here, a Web-accessible database of RCPs is presented. The distribution and evolution of RCPs that contain homopeptide repeats tracts are considered, and the existence of functional patterns investigated. Generally, it is found that while polyamino acid repeats are extremely rare in prokaryotes, several eukaryote putative homologs of prokaryote RCP-involved in important housekeeping processes-retain the repetitive region, suggesting an ancient origin for certain repeats. Within eukarya, the most common uninterrupted amino acid repeats are glutamine, asparagines, and alanine. Interestingly, while poly-Q repeats are found in vertebrates and nonvertebrates, poly-N repeats are only common in more primitive nonvertebrate organisms, such as insects and nematodes. We have assigned function to eukaryote RCPs using Online Mendelian Inheritance in Man (OMIM), the Human Reference Protein Database (HRPD), FlyBase, and Wormpep. Prokaryote RCPs were annotated using BLASTp searches and Gene Ontology. These data reveal that the majority of RCPs are involved in processes that require the assembly of large, multiprotein complexes, such as transcription and signaling.

Amino Acid Sequence↗

The high resolution crystal structure of a native thermostable serpin reveals the complex mechanism underpinning the stressed to relaxed transition.

Serpins fold into a native metastable state and utilize a complex conformational change to inhibit target proteases. An undesirable result of this conformational flexibility is that most inhibitory serpins are heat sensitive, forming inactive polymers at elevated temperatures. However, the prokaryote serpin, thermopin, from Thermobifida fusca is able to function in a heated environment. We have determined the 1.8 A x-ray crystal structure of thermopin in the native, inhibitory conformation. A structural comparison with the previously determined 1.5 A structure of cleaved thermopin provides detailed insight into the complex mechanism of conformational change in serpins. Flexibility in the shutter region and electrostatic interactions at the top of the A beta-sheet (the breach) involving the C-terminal tail, a unique structural feature of thermopin, are postulated to be important for controlling inhibitory activity and triggering conformational change, respectively, in the native state. Here we have discussed the structural basis of how this serpin reconciles the thermodynamic instability necessary for function with the stability required to withstand elevated temperatures.

Binding Sites↗

Hydrophobicity--getting into hot water.

In his famous 1959 review, Walter Kauzmann clarified important features of the thermodynamic stabilities of proteins. The hydrophobic effect is recognised as an important contributor to the stability of proteins and an important determinant of their structural patterns. As generally understood, it depends on the unusual properties of cold water and its interactions with nonpolar solutes. Here we comment on the relationship between this paradigm and the stabilities and structures of proteins from thermophilic organisms.

Archaeal Proteins↗

Prediction of protein function from protein sequence and structure.

The sequence of a genome contains the plans of the possible life of an organism, but implementation of genetic information depends on the functions of the proteins and nucleic acids that it encodes. Many individual proteins of known sequence and structure present challenges to the understanding of their function. In particular, a number of genes responsible for diseases have been identified but their specific functions are unknown. Whole-genome sequencing projects are a major source of proteins of unknown function. Annotation of a genome involves assignment of functions to gene products, in most cases on the basis of amino-acid sequence alone. 3D structure can aid the assignment of function, motivating the challenge of structural genomics projects to make structural information available for novel uncharacterized proteins. Structure-based identification of homologues often succeeds where sequence-alone-based methods fail, because in many cases evolution retains the folding pattern long after sequence similarity becomes undetectable. Nevertheless, prediction of protein function from sequence and structure is a difficult problem, because homologous proteins often have different functions. Many methods of function prediction rely on identifying similarity in sequence and/or structure between a protein of unknown function and one or more well-understood proteins. Alternative methods include inferring conservation patterns in members of a functionally uncharacterized family for which many sequences and structures are known. However, these inferences are tenuous. Such methods provide reasonable guesses at function, but are far from foolproof. It is therefore fortunate that the development of whole-organism approaches and comparative genomics permits other approaches to function prediction when the data are available. These include the use of protein-protein interaction patterns, and correlations between occurrences of related proteins in different organisms, as indicators of functional properties. Even if it is possible to ascribe a particular function to a gene product, the protein may have multiple functions. A fundamental problem is that function is in many cases an ill-defined concept. In this article we review the state of the art in function prediction and describe some of the underlying difficulties and successes.

Amino Acid Sequence↗

Evolution of amino acid frequencies in proteins over deep time: inferred order of introduction of amino acids into the genetic code.

To understand more fully how amino acid composition of proteins has changed over the course of evolution, a method has been developed for estimating the composition of proteins in an ancestral genome. Estimates are based upon the composition of conserved residues in descendant sequences and empirical knowledge of the relative probability of conservation of various amino acids. Simulations are used to model and correct for errors in the estimates. The method was used to infer the amino acid composition of a large protein set in the Last Universal Ancestor (LUA) of all extant species. Relative to the modern protein set, LUA proteins were found to be generally richer in those amino acids that are believed to have been most abundant in the prebiotic environment and poorer in those amino acids that are believed to have been unavailable or scarce. It is proposed that the inferred amino acid composition of proteins in the LUA probably reflects historical events in the establishment of the genetic code.

Amino Acid Sequence↗

Serpins in prokaryotes.

Members of the serpin (serine proteinase inhibitor) superfamily have been identified in higher multicellular eukaryotes (plants and animals) and viruses but not in bacteria, archaea, or fungi. Thus, the ancestral serpin and the origin of the serpin inhibitory mechanism remain obscure. In this study we characterize 12 serpin-like sequences in the genomes of prokaryotic organisms, extending this protein family to all major branches of life. Notably, these organisms live in dramatically different environments and some are evolutionarily distantly related. A sequence-based analysis suggests that all 12 serpins are inhibitory. Despite considerable sequence divergence between the proteins, in four of the 12 sequences the region of the serpin that determines proteinase specificity is highly conserved, indicating that these inhibitors are likely to share a common target. Inhibitory serpins are typically prone to polymerization upon heating; thus, the existence of serpins in the moderate thermophilic bacterium Thermobifida fusca, the thermophilic bacterium Thermoanaerobacter tengcongensis, and the hyperthermophilic archaeon Pyrobaculum aerophilum is of particular interest. Using molecular modeling, we predict the means by which heat stability in the latter protein may be achieved without compromising inhibitory activity.

Amino Acid Sequence↗