A note on Sattath and Tversky's, Saitou and Nei's, and Studier and Keppler's algorithms for inferring phylogenies from evolutionary distances.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to O Gascuel.
Explore the source record for details and available documents.
Inductive learning, also called 'learning from examples', is a subfield of artificial intelligence. Inductive learning methods are able to deal with 'structural descriptions'. These portray objects as composite structures consisting of various components. The use of structural descriptions to represent biological objects is appealing. For instance, they have been used by Rawlings et al [1] for symbolically and comprehensively representing the folding of proteins. This paper shows how inductive learning techniques may be used for extracting information from biological objects. We briefly describe some general techniques for describing objects in a structural way and for learning from these descriptions. We present details of a program that we developed, PLAGE, and show the application of this program for a study on signal peptides, which was done in collaboration with A Danchin [2,3]. Finally, we survey some other approaches and applications of inductive learning to molecular biology.
A method is presented for predicting the secondary structure of globular proteins from their amino acid sequence. It is based on a rigorous statistical exploitation of the well-known biological fact that the amino acid compositions of each secondary structure are different. We also propose an evaluation process that allows us to estimate the capacity of a method to predict the secondary structure of a new protein which does not have any homologous proteins whose structure is already known. This evaluation process shows that our method has a prediction accuracy of 58.7% over three states for the 62 proteins of the Kabsch and Sander (1983a) data bank. This result is better than that obtained by the most widely used methods--Lim (1974), Chou and Fasman (1978) and Garnier et al. (1978)--and also than that obtained by a recent method based on local homologies (Levin et al., 1986). Our prediction method is very simple and may be implemented on any microcomputer and even on programmable pocket calculators. A simple Pascal implementation of the method prediction algorithm is given. The interpretation of our results in terms of protein folding and directions for further work are discussed.
Investigation of possible variations between prokaryotic and eukaryotic signal sequences of exported proteins has revealed unexpected differences. Apart from the known similarities (presence of a core hydrophobic sequence preceded by a positively charged amino terminus and followed by a flexible structure), we have found that the core is much more rigid in eukaryotic signals than in their prokaryotic counterparts, and that at both ends the constraints are much more stringent in bacteria than in human cells. The differences have been summarized as a set of 17 criteria describing noteworthy features discriminating between the two classes of signal peptides. The program we used permitted each class of sequences to be learned; Escherichia coli sequences were well learned (i.e., they could be recognized by the programs as having common features), whereas human sequences were found to exhibit a much wider variation. Thus it was possible to propose a consensus in the case of the bacterial peptides, but none (or a much looser one) in the case of the human sequences. Two sequences were exceptional among the E. coli signal peptides, those of lipoprotein and plasmid-borne beta-lactamase, suggesting that they have special origins or destinations. Finally, the differences found strongly suggest that the mode of secretion is rather different in the two types of organisms, in spite of the common features of the signal sequences.
In order for the computer to learn about objects, the user must first provide a good description language for these objects. In this paper we present a new description language which is structural. Structural descriptions portray objects as composite structures consisting of various components. Structural descriptions can be contrasted with attribute descriptions, which specify only global properties of an object. Attribute descriptions can be expressed using propositional logic. Structural descriptions, however, must be expressed in first-order logic. We present how this language works, why it is suitable for biochemical objects and how one can discriminate on structural descriptions. Finally, we present an application on learning about tRNA which has yielded very good results brings evident of feasibility.
The information collected in national and international libraries on nucleotide and protein sequences cannot be directly treated for proper handling by existing software. Therefore we evaluated the feasibility of constructing a data base for Escherichia coli using the data present in the banks. The knowhow thus acquired was applied to Bacillus subtilis. Specific examples of the general procedure are given.