Search PubMedSearch

Biomedical subjects

W J Bruno

Publications and source records attributed to W J Bruno.

8 recordsLinked to original sources

Major structural determinants of transmembrane proteins identified by principal component analysis.

We identify amino acid characteristics important in determining the secondary structures of transmembrane proteins, and compare them with characteristics important for cytoplasmic proteins. Using information derived from multiple sequence alignments, we perform a principal component analysis (PCA) to identify the directions in the 20-dimensional amino acid frequency space that comprise the most variance within each protein secondary structure. These vectors represent the important position-specific properties of the amino acids for coils, turns, beta sheets, and alpha helices. As expected, the most important axis for most of the datasets was hydrophobicity. Additional axes, distinct from hydrophobicity, are surprising, especially in the case of transmembrane alpha helices, where the effects of aromaticity and beta-branching are the next two most significant characteristics. The axis representing beta-branching also has equal importance in cytoplasmic and transmembrane helices, a finding that contrasts with some experimental results in membrane-like environments. In a further analysis, we examine trends for some of the PCA axes over averaged transmembrane alpha helices, and find interesting results for aromaticity.

Amino Acids

Evolutionary distances for protein-coding sequences: modeling site-specific residue frequencies.

Estimation of evolutionary distances from coding sequences must take into account protein-level selection to avoid relative underestimation of longer evolutionary distances. Current modeling of selection via site-to-site rate heterogeneity generally neglects another aspect of selection, namely position-specific amino acid frequencies. These frequencies determine the maximum dissimilarity expected for highly diverged but functionally and structurally conserved sequences, and hence are crucial for estimating long distances. We introduce a codon-level model of coding sequence evolution in which position-specific amino acid frequencies are free parameters. In our implementation, these are estimated from an alignment using methods described previously. We use simulations to demonstrate the importance and feasibility of modeling such behavior; our model produces linear distance estimates over a wide range of distances, while several alternative models underestimate long distances relative to short distances. Site-to-site differences in rates, as well as synonymous/nonsynonymous and first/second/third-codon-position differences, arise as a natural consequence of the site-to-site differences in amino acid frequencies.

Amino Acids

Estimation of reversible substitution matrices from multiple pairs of sequences.

We present a method for estimating the most general reversible substitution matrix corresponding to a given collection of pairwise aligned DNA sequences. This matrix can then be used to calculate evolutionary distances between pairs of sequences in the collection. If only two sequences are considered, our method is equivalent to that of Lanave et al. (1984). The main novelty of our approach is in combining data from different sequence pairs. We describe a weighting method for pairs of taxa related by a known tree that results in uniform weights for all branches. Our method for estimating the rate matrix results in fast execution times, even on large data sets, and does not require knowledge of the phylogenetic relationships among sequences. In a test case on a primate pseudogene, the matrix we arrived at resembles one obtained using maximum likelihood, and the resulting distance measure is shown to have better linearity than is obtained in a less general model.

Algorithms

Modeling residue usage in aligned protein sequences via maximum likelihood.

A computational method is presented for characterizing residue usage, i.e., site-specific residue frequencies, in aligned protein sequences. The method obtains frequency estimates that maximize the likelihood of the sequences in a simple model for sequence evolution, given a tree or a set of candidate trees computed by other methods. These maximum-likelihood frequencies constitute a profile of the sequences, and thus the method offers a rigorous alternative to sequence weighting for constructing such a profile. The ability of this method to discard misleading phylogenetic effects allows the biochemical propensities of different positions in a sequence to be more clearly observed and interpreted.

Amino Acids

Single-base sequencing and similarity comparisons.

A "single-base sequence" is a DNA sequence in which the identities and locations of bases of only one type have been determined. We present experimental procedures for single-base sequencing and describe the effective use of existing software (FASTA) in similarity comparisons of single-base sequences. We determined the theoretical and experimental minimum sequence lengths required for identification of a sequence within a large dataset and optimized the FASTA parameters for use in single-base similarity comparisons. Single-base sequences have been used to identify cDNAs occurring in a database. Single-base sequencing could be used to reduce the redundancy of "shot-gun sequencing."

Databases, Factual

Efficient pooling designs for library screening.

We describe efficient methods for screening clone libraries, based on pooling schemes that we call "random k-sets designs." In these designs, the pools in which any clone occurs are equally likely to be any possible selection of k from the v pools. The values of k and v can be chosen to optimize desirable properties. Random k-sets designs have substantial advantages over alternative pooling schemes: they are efficient, flexible, and easy to specify, require fewer pools, and have error-correcting and error-detecting capabilities. In addition, screening can often be achieved in only one pass, thus facilitating automation. For design comparison, we assume a binomial distribution for the number of "positive" clones, with parameters n, the number of clones, and c, the coverage. We propose the expected number of resolved positive clones--clones that are definitely positive based upon the pool assays--as a criterion for the efficiency of a pooling design. We determine the value of k that is optimal, with respect to this criterion, as a function of v, n, and c. We also describe superior k-sets designs called k-sets packing designs. As an illustration, we discuss a robotically implemented design for a 2.5-fold-coverage, human chromosome 16 YAC library of n = 1298 clones. We also estimate the probability that each clone is positive, given the pool-assay data and a model for experimental errors.

Binomial Distribution

Vibrationally enhanced tunneling as a mechanism for enzymatic hydrogen transfer.

We present a theory of enzymatic hydrogen transfer in which hydrogen tunneling is mediated by thermal fluctuations of the enzyme's active site. These fluctuations greatly increase the tunneling rate by shortening the distance the hydrogen must tunnel. The average tunneling distance is shown to decrease when heavier isotopes are substituted for the hydrogen or when the temperature is increased, leading to kinetic isotope effects (KIEs)--defined as the factor by which the reaction slows down when isotopically substituted substrates are used--that need be no larger than KIEs for nontunneling mechanisms. Within this theory we derive a simple KIE expression for vibrationally enhanced ground state tunneling that is able to fit the data for the bovine serum amine oxidase (BSAO) system, correctly predicting the large temperature dependence of the KIEs. Because the KIEs in this theory can resemble those for nontunneling dynamics, distinguishing the two possibilities requires careful measurements over a range of temperatures, as has been done for BSAO.

Alcohol Dehydrogenase