Search PubMed⌕ Search

Biomedical subjects

William R Atchley

Publications and source records attributed to William R Atchley.

5 recordsLinked to original sources

Molecular architecture of the DNA-binding region and its relationship to classification of basic helix-loop-helix proteins.

Multivariate statistical analyses are used to explore the molecular architecture of the DNA-binding and dimerization regions of basic helix-loop-helix (bHLH) proteins. Alphabetic amino acid data are transformed to biologically meaningful quantitative values using a set of 5 multivariate "indices." These multivariate indices summarize variation in a large suite of amino acid physiochemical attributes and reflect variability in polarity-accessibility-hydrophobicity, propensity for secondary structure, molecular size, codon composition, and electrostatic charge. Using these index score data, discriminant analyses describe the multidimensional aspects of physiochemical variation and clarify the structural basis of the prevailing evolutionary classification of bHLH proteins. A small number of amino acids from both the binding dimerization domains, when considered simultaneously, accurately distinguish the 5 known DNA-binding groups. The relevant sites often have well-documented structural and functional characteristics.

Animals↗

Networks of coevolving sites in structural and functional domains of serpin proteins.

Amino acids do not occur randomly in proteins; rather, their occurrence at any given site is strongly influenced by the amino acid composition at other sites, the structural and functional aspects of the region of the protein in which they occur, and the evolutionary history of the protein. The goal of our research study is to identify networks of coevolving sites within the serpin proteins (serine protease inhibitors) and classify them as being caused by structural-functional constraints or by evolutionary history. To address this, a matrix of pairwise normalized mutual information (NMI) values was computed among amino acid sites for the serpin proteins. The NMI matrix was partitioned into orthogonal patterns of amino acid variability by factor analysis. Each common factor pattern was interpreted as having phylogenetic and/or structural-functional explanations. In addition, we used a bootstrap factor analysis technique to limit the effects of phylogenetic history on our factor patterns. Our results show an extensive network of correlations among amino acid sites in key functional regions (reactive center loop, shutter, and breach). Additionally, we have discovered long-range coevolution for packed amino acids within the serpin protein core. Lastly, we have discovered a group of serpin sites which coevolve in the hydrophobic core region (s5B and s4B) and appear to represent sites important for formation of the "native" instead of the "latent" serpin structure. This research provides a better understanding on how protein structure evolves; in particular, it elucidates the selective forces creating coevolution among protein sites.

Amino Acids↗

Solving the protein sequence metric problem.

Biological sequences are composed of long strings of alphabetic letters rather than arrays of numerical values. Lack of a natural underlying metric for comparing such alphabetic data significantly inhibits sophisticated statistical analyses of sequences, modeling structural and functional aspects of proteins, and related problems. Herein, we use multivariate statistical analyses on almost 500 amino acid attributes to produce a small set of highly interpretable numeric patterns of amino acid variability. These high-dimensional attribute data are summarized by five multidimensional patterns of attribute covariation that reflect polarity, secondary structure, molecular volume, codon diversity, and electrostatic charge. Numerical scores for each amino acid then transform amino acid sequences for statistical analyses. Relationships between transformed data and amino acid substitution matrices show significant associations for polarity and codon diversity scores. Transformed alphabetic data are used in analysis of variance and discriminant analysis to study DNA binding in the basic helix-loop-helix proteins. The transformed scores offer a general solution for analyzing a wide variety of sequence analysis problems.

Amino Acid Sequence↗

Sequence signatures and the probabilistic identification of proteins in the Myc-Max-Mad network.

Accurate identification of specific groups of proteins by their amino acid sequence is an important goal in genome research. Here we combine information theory with fuzzy logic search procedures to identify sequence signatures or predictive motifs for members of the Myc-Max-Mad transcription factor network. Myc is a well known oncoprotein, and this family is involved in cell proliferation, apoptosis, and differentiation. We describe a small set of amino acid sites from the N-terminal portion of the basic helix-loop-helix (bHLH) domain that provide very accurate sequence signatures for the Myc-Max-Mad transcription factor network and three of its member proteins. A predictive motif involving 28 contiguous bHLH sequence elements found 337 network proteins in the GenBank NR database with no mismatches or misidentifications. This motif also identifies at least one previously unknown fungal protein with strong affinity to the Myc-Max-Mad network. Another motif found 96% of known Myc protein sequences with only a single mismatch, including sequences from genomes previously not thought to contain Myc proteins. The predictive motif for Myc is very similar to the ancestral sequence for the Myc group estimated from phylogenetic analyses. Based on available crystal structure studies, this motif is discussed in terms of its functional consequences. Our results provide insight into evolutionary diversification of DNA binding and dimerization in a well characterized family of regulatory proteins and provide a method of identifying signature motifs in protein families.

Amino Acid Sequence↗

Phylogenetic analysis of plant basic helix-loop-helix proteins.

The basic helix-loop-helix (bHLH) family of proteins is a group of functionally diverse transcription factors found in both plants and animals. These proteins evolved early in eukaryotic cells before the split of animals and plants, but appear to function in 'plant-specific' or 'animal-specific' processes. In animals bHLH proteins are involved in regulation of a wide variety of essential developmental processes. On the contrary, bHLH proteins have not been extensively studied in plants. Those that have been characterized function in anthocyanin biosynthesis, phytochrome signaling, globulin expression, fruit dehiscence, carpel and epidermal development. We have identified 118 different bHLH genes in the completely sequenced Arabidopsis thaliana genome and 131 bHLH genes in the rice genome. Here we report a phylogenetic analysis of these genes, including 46 genes from other plant species and a classification of these proteins into 15 distinct plant clades. Results imply a polyphyletic origin for the plant bHLH proteins related only by their bHLH DNA binding motif. We suggest that plant bHLH proteins are under weaker selective constraints than their animal counterparts and that lineage specific expansions and subfunctionalization have fashioned regulatory proteins for plant specific functions.

Amino Acid Sequence↗