Search PubMed⌕ Search

Biomedical subjects

W J Wilbur

Publications and source records attributed to W J Wilbur.

At least 19 recordsLinked to original sources

Hidden Markov models and optimized sequence alignments.

We present a formulation of the Needleman-Wunsch type algorithm for sequence alignment in which the mutation matrix is allowed to vary under the control of a hidden Markov process. The fully trainable model is applied to two problems in bioinformatics: the recognition of related gene/protein names and the alignment and scoring of homologous proteins.

Animals↗

Amino acid residue environments and predictions of residue type.

The determination of a protein's structure from the knowledge of its linear chain is one of the important problems that remains as a bottleneck in interpreting the rapidly increasing repository of genetic sequence data. One approach to this problem that has shown promise and given a measure of success is threading. In this approach contact energies between different amino acids are first determined by statistical methods applied to known structures. These contact energies are then applied to a sequence whose structure is to be determined by threading it through various known structures and determining the total threading energy for each candidate structure. That structure that yields the lowest total energy is then considered the leading candidate among all the structures tested. Additional information is often needed in order to support the results of threading studies, as it is well known in the field that the contact potentials used are not sufficiently sensitive to allow definitive conclusions. Here, we investigate the hypothesis that the environment of an amino acid residue realized as all those residues not local to it on the chain but sufficiently close spatially can supply information predictive of the type of that residue that is not adequately reflected in the individual contact energies. We present evidence that confirms this hypothesis and suggests a high order cooperativity between the residues that surround a given residue and how they interact with it. We suggest a possible application to threading.

Algorithms↗

Automatic MeSH term assignment and quality assessment.

For computational purposes documents or other objects are most often represented by a collection of individual attributes that may be strings or numbers. Such attributes are often called features and success in solving a given problem can depend critically on the nature of the features selected to represent documents. Feature selection has received considerable attention in the machine learning literature. In the area of document retrieval we refer to feature selection as indexing. Indexing has not traditionally been evaluated by the same methods used in machine learning feature selection. Here we show how indexing quality may be evaluated in a machine learning setting and apply this methodology to results of the Indexing Initiative at the National Library of Medicine.

Abstracting and Indexing↗

A theory of information with special application to search problems.

Classical information theory concerns itself with communication through a noisy channel and how much one can infer about the channel input from a knowledge of the channel output. Because the channel is noisy the input and output are only related statistically and the rate of information transmission is a statistical concept with little meaning for the individual symbol used in transmission. Here we develop a more intuitive notion of information that is concerned with asking the right questions--that is, with finding those questions whose answer conveys the most information. We call this confirmatory information. In the first part of the paper we develop the general theory, show how it relates to classical information theory, and how in the special case of search problems it allows us to quantify the efficacy of information transmission regarding individual events. That is, confirmatory information measures how well a search for items having certain observable properties retrieves items having some unobserved property of interest. Thus confirmatory information facilitates a useful analysis of search problems and contrasts with classical information theory, which quantifies the efficiency of information transmission but is indifferent to the nature of the particular information being transmitted. The last part of the paper presents several examples where confirmatory information is used to quantify protein structural properties in a search setting.

Information Theory↗

Genes, themes and microarrays: using information retrieval for large-scale gene analysis.

The immense volume of data resulting from DNA microarray experiments, accompanied by an increase in the number of publications discussing gene-related discoveries, presents a major data analysis challenge. Current methods for genome-wide analysis of expression data typically rely on cluster analysis of gene expression patterns. Clustering indeed reveals potentially meaningful relationships among genes, but can not explain the underlying biological mechanisms. In an attempt to address this problem, we have developed a new approach for utilizing the literature in order to establish functional relationships among genes on a genome-wide scale. Our method is based on revealing coherent themes within the literature, using a similarity-based search in document space. Content-based relationships among abstracts are then translated into functional connections among genes. We describe preliminary experiments applying our algorithm to a database of documents discussing yeast genes. A comparison of the produced results with well-established yeast gene functions demonstrates the effectiveness of our approach.

Animals↗

The NLM Indexing Initiative.

The objective of NLM's Indexing Initiative (IND) is to investigate methods whereby automated indexing methods partially or completely substitute for current indexing practices. The project will be considered a success if methods can be designed and implemented that result in retrieval performance that is equal to or better than the retrieval performance of systems based principally on humanly assigned index terms. We describe the current state of the project and discuss our plans for the future.

Abstracting and Indexing↗

Boosting naïve Bayesian learning on a large subset of MEDLINE.

We are concerned with the rating of new documents that appear in a large database (MEDLINE) and are candidates for inclusion in a small specialty database (REBASE). The requirement is to rank the new documents as nearly in order of decreasing potential to be added to the smaller database as possible, so as to improve the coverage of the smaller database without increasing the effort of those who manage this specialty database. To perform this ranking task we have considered several machine learning approaches based on the naï ve Bayesian algorithm. We find that adaptive boosting outperforms naï ve Bayes, but that a new form of boosting which we term staged Bayesian retrieval outperforms adaptive boosting. Staged Bayesian retrieval involves two stages of Bayesian retrieval and we further find that if the second stage is replaced by a support vector machine we again obtain a significant improvement over the strictly Bayesian approach.

Algorithms↗

Analysis of biomedical text for chemical names: a comparison of three methods.

At the National Library of Medicine (NLM), a variety of biomedical vocabularies are found in data pertinent to its mission. In addition to standard medical terminology, there are specialized vocabularies including that of chemical nomenclature. Normal language tools including the lexically based ones used by the Unified Medical Language System (UMLS) to manipulate and normalize text do not work well on chemical nomenclature. In order to improve NLM's capabilities in chemical text processing, two approaches to the problem of recognizing chemical nomenclature were explored. The first approach was a lexical one and consisted of analyzing text for the presence of a fixed set of chemical segments. The approach was extended with general chemical patterns and also with terms from NLM's indexing vocabulary, MeSH, and the NLM SPECIALIST lexicon. The second approach applied Bayesian classification to n-grams of text via two different methods. The single lexical method and two statistical methods were tested against data from the 1999 UMLS Metathesaurus. One of the statistical methods had an overall classification accuracy of 97%.

Algorithms↗

A free energy analysis by unfolding applied to 125-mers on a cubic lattice.

BACKGROUND: A common approach to the protein folding problem involves computer simulation of folding using lattice models of amino acid sequences. Key factors for good performance in such models are the correct choice of the temperature and the average interaction energy between residues. In order to push the lattice approach to its limit it is important to have a method to adjust these parameters for optimal folding that is not limited by our ability to successfully simulate folding in a reasonable time. RESULTS: In this study, we adopt a simple cubic-lattice model and present a method for calculating the free energy of a chain as a function of the number of native contacts. This does not require that we are able to fold the sequence by simulation and it provides a method of estimating the folding transition temperature. For a given set of parameters, the free energy analysis also allows an estimate of foldability. By applying the method to sequences with 27 and 125 residues, we show that optimal folding occurs near the folding transition temperature and at either zero or small negative average interaction energy. We find ourselves able to fold only 125-mers that have significant short-range native contacts. CONCLUSIONS: A free energy analysis during unfolding is a useful tool for the study of foldability and should be applicable to a variety of folding models. In this way we are able to fold some 125-mer designed sequences and our results confirm the finding that short-range contacts contribute to foldability.

Markov Chains↗

The statistics of unique native states for random peptides.

Given a probability distribution from which the energy spectrum of a random peptide is to be sampled, we derive a general expression for the probability that such a peptide will fold to a unique native state and for the probability distribution of the native energy. This latter result allows us to localize the energy of folding based on model parameters and is one advantage of our formulation. Evidence from both the lattice theory of proteins and protein threading experiments suggest that the energy spectrum for the compact states of a peptide chain is Gaussian in form. For this reason we have derived from the more general framework the specific formulas that apply in the Gaussian case, where one requires only the number of states and the variance of the Gaussian distribution in order to apply the theory. This simplicity allows us to perform calculations that we compare with calculations previously made by others based on statistical thermodynamics. We find qualitative agreement, but a significant correction to prior estimates of folding probability derived from the Gaussian assumption is necessary.

Mathematical Computing↗

An analysis of statistical term strength and its use in the indexing and retrieval of molecular biology texts.

The biological literature presents a difficult challenge to information processing in its complexity, diversity, and in its sheer volume. Much of the diversity resides in its technical terminology, which has also become voluminous. In an effort to deal more effectively with this large vocabulary and improve information processing, a method of focus has been developed which allows one to classify terms based on a measure of their importance in describing the content of the documents in which they occur. The measurement is called the strength of a term and is a measure of how strongly the term's occurrences correlate with the subjects of documents in the database. If term occurrences are random then there will be no correlation and the strength will be zero, but if for any subject, the term is either always present or never present its strength will be one. We give here a new, information theoretical interpretation of term strength, review some of its uses in focusing the processing of documents for information retrieval and describe new results obtained in document categorization.

Abstracting and Indexing↗

Modelling neutral and selective evolution of protein folding.

We examine a model evolutionary space consisting of genotypes mapped to their corresponding phenotypes. This mapping is derived from a lattice model for proteins which, despite its highly idealized nature, has been shown to share general properties with real proteins. Large evolutionary networks are observed, with genotypes corresponding to non-lethal phenotypes linked by unit mutational steps. Neutral mutations are necessary for traversing the evolutionary networks, and even one neutral mutation in a genotype can change the phenotypes attainable by a unit mutational step.

Biological Evolution↗

A biophysical model of cochlear processing: intensity dependence of pure tone responses.

A mathematical model of cochlear processing is developed to account for the nonlinear dependence of frequency selectivity on intensity in inner hair cell and auditory nerve fiber responses. The model describes the transformation from acoustic stimulus to intracellular hair cell potentials in the cochlea. It incorporates a linear formulation of basilar membrane mechanics and subtectorial fluid-cilia displacement coupling, and a simplified description of the inner hair cell nonlinear transduction process. The analysis at this stage is restricted to low-frequency single tones. The computed responses to single tone inputs exhibit the experimentally observed nonlinear effects of increasing intensity such as the increase in the bandwidth of frequency selectivity and the downward shift of the best frequency. In the model, the first effect is primarily due to the saturating effect of the hair cell nonlinearity. The second results from the combined effects of both the nonlinearity and of the inner hair cell low-pass transfer function. In contrast to these shifts along the frequency axis, the model does not exhibit intensity dependent shifts of the spatial location along the cochlea of the peak response for a given single tone. The observed shifts therefore do not contradict an intensity invariant tonotopic code.

Acoustic Stimulation↗

On the PAM matrix model of protein evolution.

The internal consistency of the PAM matrix model of protein evolution is here investigated. The 1 PAM matrix has been constructed from amino acid replacements observed in closely related sequences. Such replacements are of two types, those that do not require an intermediate amino acid replacement and those that do. The second type of replacement must generally be produced by a repetition of the first. This allows data on the first type to be used in predicting data on the second type so that some elements of the 1 PAM matrix may be used to predict others. A discrepancy of more than two orders of magnitude is found between the predictions and the data when this is carried out. This is partly accounted for by an error in constructing the matrix. However, it also seems necessary that the basic model be modified. Several possibilities are considered. One of these is to incorporate a site-dependent spectrum of mutabilities associated with each amino acid.

Amino Acid Sequence↗

On the statistical significance of nucleic acid similarities.

When evaluating sequence similarities among nucleic acids by the usual methods, statistical significance is often found when the biological significance of the similarity is dubious. We demonstrate that the known statistical properties of nucleic acid sequences strongly affect the statistical distribution of similarity values when calculated by standard procedures. We propose a series of models which account for some of these known statistical properties. The utility of the method is demonstrated in evaluating high relative similarity scores in four specific cases in which there is little biological context by which to judge the similarities. In two of the cases we identify the statistical properties which are responsible for the apparent similarity. In the other two cases the statistical significance of the similarity persists even when the known statistical properties of sequences are modelled. For one of these cases biological significance is likely while the other case remains an enigma.

Base Sequence↗

A theoretical basis for large coefficient of variation and bimodality in neuronal interspike interval distributions.

We consider the classic Stein (1965) model for stochastic neuronal firing under random synaptic input. Our treatment includes the additional effect of synaptic reversal potentials. We develop and employ two numerical methods (in addition to Monte Carlo simulations) to study the relation of the various parameters of the model to the shape of the theoretical interspike interval distribution. Contrary to the results of Tuckwell (1979) we are unable to account, on the basis of substantial synaptic inhibition and with parameter settings in the known physiologic range, for experimental interspike interval distributions which exhibit large coefficients of variation or bimodality. We therefore introduce a time varying threshold into the model, which readily allows for such distributions and which has physiological justification.

Action Potentials↗

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence↗

Rapid similarity searches of nucleic acid and protein data banks.

With the development of large data banks of protein and nucleic acid sequences, the need for efficient methods of searching such banks for sequences similar to a given sequence has become evident. We present an algorithm for the global comparison of sequences based on matching k-tuples of sequence elements for a fixed k. The method results in substantial reduction in the time required to search a data bank when compared with prior techniques of similarity analysis, with minimal loss in sensitivity. The algorithm has also been adapted, in a separate implementation, to produce rigorous sequence alignments. Currently, using the DEC KL-10 system, we can compare all sequences in the entire Protein Data Bank of the National Biomedical Research Foundation with a 350-residue query sequence in less than 3 min and carry out a similar analysis with a 500-base query sequence against all eukaryotic sequences in the Los Alamos Nucleic Acid Data Base in less than 2 min.

Amino Acid Sequence↗