Search PubMed⌕ Search

Biomedical subjects

J C Aude

Publications and source records attributed to J C Aude.

5 recordsLinked to original sources

An incremental algorithm for Z-value computations.

The Z-value (Comput. Chem. 23 (1999) 333) is an extension of the Z-score that is classically used to compare sets of biological sequences. The Z-value has been successfully used to handle complete genome studies as well as analyze large sets of proteins. The Z-value computation is based on a Monte Carlo approach to estimate the statistical significance of a Smith & Waterman alignment score. Comet et al. (Comput. Chem. 23 (1999) 333) have shown that, in contrast to the alignment score, the Z-value largely reduces the bias due to the lengths and compositions of the sequences. They also described an estimator of the deviation of Z-values, that we extend in this paper in order to optimize Z-values computation. The incremental algorithm described here provides two characteristics which are usually incompatible: (i) it improves the accuracy of Z-values calculation; (ii) it reduces the time complexity (this algorithm has been named incremental because it iteratively adds random sequences to the Monte-Carlo process when needed). Results are presented, originating from the all-by-all comparison of the proteins from Saccharomyces cerevisiae and Escherichia coli.

Algorithms↗

Massive sequence comparisons as a help in annotating genomic sequences.

An all-by-all comparison of all the publicly available protein sequences from plants has been performed, followed by a clusterization process. Within each of the 1064 resulting clusters-containing sequences that are orthologous as well as paralogous-the sequences have been submitted to a pyramidal classification and their domains delineated by an automated procedure à la. This process provides a means for easily checking for any apparent inconsistency in a cluster, for example, whether one sequence is shorter or longer than the others, one domain is missing, etc. In such cases, the alignment of the DNA sequence of the gene with that of a close homologous protein often reveals (in 10% of the clusters) probable sequencing errors (leading to frameshifts) or probable wrong intron/exon predictions. The composition of the clusters, their pyramidal classifications, and domain decomposition, as well as our comments when appropriate, are available from http://chlora.infobiogen.fr:1234/PHYTOPROT.

Amino Acid Sequence↗

Applications of the pyramidal clustering method to biological objects.

In conventional hierarchical clustering methods, any object can belong to only one class or cluster. We present here an application of the pyramidal classification method to biological objects, which illustrates the intuitively appealing idea that some objects may belong simultaneously to two classes. In a first step, we performed an all-by-all comparison of all the open reading frames in the genomes from S. cerevisiae, M. jannaschii, E. coli, H. influenzae and Synechocystis. In a second step, a series of connex classes was built, each connex class containing all those sequences that were linked by a Z-value (obtained after 100 sequence shufflings) greater than a given threshold. Finally, each connex class was submitted to a pyramidal classification. Three examples of such classifications are given, concerning two sets of multi-domains protein sequences and a family of aminoacyl-tRNA synthetases. They make it clear that the linear order among the classified objects that results from the pyramidal classification is useful in deciphering the multiple relationships that can exist between the objects under study. A program for calculating and displaying a pyramidal classification from a dissimilarity matrix is available from http:/(/)genome.genetique.uvsq.fr/Pyramids. The pyramidal classifications of the connex classes from the five organisms (intra- and inter-genomic comparisons) are available from http:/(/)www.gene-it.com under the family item.

Algorithms↗

Significance of Z-value statistics of Smith-Waterman scores for protein alignments.

The Z-value is an attempt to estimate the statistical significance of a Smith-Waterman dynamic alignment score (SW-score) through the use of a Monte-Carlo process. It partly reduces the bias induced by the composition and length of the sequences. This paper is not a theoretical study on the distribution of SW-scores and Z-values. Rather, it presents a statistical analysis of Z-values on large datasets of protein sequences, leading to a law of probability that the experimental Z-values follow. First, we determine the relationships between the computed Z-value, an estimation of its variance and the number of randomizations in the Monte-Carlo process. Then, we illustrate that Z-values are less correlated to sequence lengths than SW-scores. Then we show that pairwise alignments, performed on 'quasi-real' sequences (i.e., randomly shuffled sequences of the same length and amino acid composition as the real ones) lead to Z-value distributions that statistically fit the extreme value distribution, more precisely the Gumbel distribution (global EVD, Extreme Value Distribution). However, for real protein sequences, we observe an over-representation of high Z-values. We determine first a cutoff value which separates these overestimated Z-values from those which follow the global EVD. We then show that the interesting part of the tail of distribution of Z-values can be approximated by another EVD (i.e., an EVD which differs from the global EVD) or by a Pareto law. This has been confirmed for all proteins analysed so far, whether extracted from individual genomes, or from the ensemble of five complete microbial genomes comprising altogether 16956 protein sequences.

Computing Methodologies↗

Evolution of genes, evolution of species: the case of aminoacyl-tRNA synthetases.

All of the aminoacyl-tRNA synthetase (aaRS) sequences currently available in the data banks have been subjected to a systematic analysis aimed at finding gene duplications, genetic recombinations, and horizontal transfers. Evidence is provided for the occurrence (or probable occurrence) of such phenomena within this class of enzymes. In particular, it is suggested that the monomeric PheRS from the yeast mitochondrion is a chimera of the alpha and beta chains of the standard tetrameric protein. In addition, it is proposed that the dimeric and tetrameric forms of GlyRS are the result of a double and independent acquisition of the same specificity within two different subclasses of aaRS. The phylogenetic reconstructions of the evolutionary histories of the genes encoding aaRS are shown to be extremely diverse. While large segments of the population are consistent with the broad grouping into the three Woesian domains, some phylogenetic reconstructions do not place the Archae and the Eucarya as sister groups but, rather, show a gram-negative bacteria/eukaryote clustering. In addition, many individual genes pose difficulties that preclude any simple evolutionary scheme. Thus, aaRS's are clearly a paradigm of F. Jacob's "odd jobs of evolution" but, on the whole, do not call into question the evolutionary scenario originally proposed by Woese and subsequently refined by others.

Amino Acid Sequence↗