Search PubMed⌕ Search

Biomedical subjects

J Sallantin

Publications and source records attributed to J Sallantin.

13 recordsLinked to original sources

Improving the efficiency of a user-driven learning system with reconfigurable hardware. Application to DNA splicing.

This paper describes a new approach to problem solving by splitting up problem component parts between software and hardware. Our main idea arises from the combination of two previously published works. The first one proposed a conceptual environment of concept modelling in which the machine and the human expert interact. The second one reported an algorithm based on reconfigurable hardware system which outperforms any kind of previously published genetic data base scanning hardware or algorithms. Here we show how efficient the interaction between the machine and the expert is when the concept modelling is based on reconfigurable hardware system. Their cooperation is thus achieved with an real time interaction speed. The designed system has been partially applied to the recognition of primate splice junctions sites in genetic sequences.

Algorithms↗

A computer aided system for systematic production and revision of sequence patterns.

We used two complementary fields, object-oriented databases and machine learning, to produce and revise a set of protein sequence patterns. In a first stage, we show that object-oriented query languages are well suited for the production of patterns as well as for the interpretation of the biological function of new (uncharacterized) sequences. In a second stage, a classification is built from the set of sequences according to the pattern matches. This classification may be criticized by a specific analysis method, which yields back to revise sequences and patterns. In our application, we have used concept lattices as a classification method and sequence multiple alignment for criticism.

Algorithms↗

High speed pattern matching in genetic data base with reconfigurable hardware.

Homology detection in large data bases is probably the most time consuming operation in molecular genetic computing systems. Moreover, the progresses made all around the world concerning the mapping and sequencing of the genome of Homo Sapiens and other species have increased the size of data bases exponentially. Therefore even the best workstation would not be able to reach the scanning speed required. In order to answer this need we propose an algorithm, A2R2, and its implementation on a massively parallel system. Basically, two kinds of algorithms are used to search in molecular genetic data bases. The first kind is based on dynamic programming and the second on word processing, A2R2 belongs to the second kind. The structure of the motif (pattern) searched by A2R2 can support those from FAST, BLAST and FLASH algorithms. After a short presentation of the reconfigurable hardware concept and technology used in our massively parallel accelerator we present the A2R2 implementation. This parallel implementation outperforms any kind of previously published genetic data base scanning hardware or algorithms. We report up to 25 million nucleotides per scanning seconds as our best results.

Algorithms↗

Learning and alignment methods applied to protein structure prediction.

Learning techniques are able to extract structural knowledge specific to a selected set of proteins. We describe two algorithms that optimize scores expressing the propensity of a polypeptide sequence to adopt a local fold. The first algorithm generates secondary structure prediction rules based on a dictionary of geometrical patterns frequently found in the learning database. The second algorithm leads to scores that indicate the fit between an amino acid and a given local structural environment. Dynamic programming is then used to align structural information profiles by modifying the local mutation cost with the above learned functions. The main features of the system are exemplified on the structural prediction of the N-terminal domain of the CD4 antigen. Then the usefulness of additional 3-D information in the alignment is benchmarked on eight pairs of weakly homologous proteins.

Algorithms↗

Improved alignment of weakly homologous protein sequences using structural information.

Protein sequence alignments can be improved when at least one of the proteins to be aligned has a known 3-D structure. In this work, geometrical constraints extracted from the target fold are evaluated in independent units that deal with complementary structural features. This information is used to set up mutation tables specific to the locally observed structural environments. The resulting partial evaluations are then combined linearly into a global function which is optimized by dynamic programming. Eventually, a score based on tertiary interactions can be used as a selection criterion to discriminate among a set of suboptimal alignments. The relevance of the scores given by each unit is tested on a representative set of protein families. Finally, a method for combining the different scores is described and its efficiency is evaluated on a few pairs of weakly homologous proteins.

Algorithms↗

A modular learning environment for protein modeling.

We propose in this paper a modular learning environment for protein modeling. In this system, the protein modeling problem is tackled in two successive phases. First, partial structural informations are determined via numerical learning techniques. Then, in the second phase, the multiple available informations are combined in pattern matching searches via dynamic programming. It is shown on real problems that various protein structure predictions can be improved in this way, such as secondary structure prediction, alignment of weakly homologous protein sequences or protein model evaluations.

Amino Acid Sequence↗

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms↗

PCDRA: PC interactive molecular representation and modeling system.

PCDRA was designed to provide the average biologist with a user-friendly molecular display on a low-cost personal computer. The package is menu driven and is built so that a biologist, with little or no computing knowledge, finds it easy to use. The system gives a color representation with depth cueing of a protein whose atomic coordinates are stored as a PDB file. Moreover, the system presents several features similar to HYDRA and therefore is a good introduction to molecular graphics, especially for beginners in protein modeling.

Computer Graphics↗

Localization of the initiation of translation in messenger RNAs of prokaryotes by learning techniques.

Learning processes are applied to the recognition of protein coding regions in prokaryotes. Non-contradictory, statistical and logical rules are deduced from a set of known examples of coding sequences. These rules enable to build characteristic patterns on the m-RNA upstream of the initiating codon. These rules are applied with success to recognize more than 180 coding sequences and to detect and/or eliminate hypothetical reading frames or unknown genes.

Bacteria↗

Search for promoter sites of prokaryotic DNA using learning techniques.

Using learning techniques previously described in this journal, we have built an expert system able to point to the start DNA point of a sequence and therefore to recognize a promoter. However, to build this system, we have focused on the TATA box and its environment. We have used this expert system to look for new promoters and also to construct new promoters. The results obtained are discussed.

Base Sequence↗

Computer search of calcium binding sites in a gene data bank: use of learning techniques to build an expert system.

Using a learning set of 28 sequences able to bind calcium (each sequence is 12 residues long), we have built two filters by learning on this set. The first filter uses a pattern-matching technique and the second one takes into account the environment of amino-acids. These two filters have been used to find new calcium-binding proteins in a data bank. The results are discussed.

Amino Acid Sequence↗

[Localization of initiating codons in RNA prokaryotes messengers by learning technics].

Learning processes are applied to the recognition of protein coding regions in prokaryotes. Non-contradictory, statistical rules are deduced from a set of known examples of coding regions. These rules allow us to build characteristic patterns on the m-RNA upstream the initiating codon. These rules are applied to recognize more than 180 coding sequences.

Cell Physiological Phenomena↗

[The use of pattern recognition for analysing the possible relation between molecular structure of polycyclic hydrocarbons and their carcinogenicity].

Pattern recognition has been used for investigating the part played by the molecular structure of polycyclic hydrocarbons in their carcinogenic action. A series of molecules has been considered, whose physicochemical properties are known to be closely related to the number and location of their aromatic rings. From a pattern recognition program and two learning subsets (carcinogenic and non carcinogenic molecules), it has been shown that (i) the shape of the molecule is correlated with its carcinogenic power; (ii) an index of carcinogenicity can be estimated for any molecule in the considered set.

Carcinogens↗