Search PubMed⌕ Search

Biomedical subjects

M Gouy

Publications and source records attributed to M Gouy.

At least 55 records · Page 3Linked to original sources

Phylogenetic analysis based on rRNA sequences supports the archaebacterial rather than the eocyte tree.

How many primary lineages of life exist and what are their evolutionary relationships? These are fundamental but highly controversial issues. Woese and co-workers propose that archaebacteria, eubacteria and eukaryotes are the three primary lines of descent and their relationships can be represented by Fig. 1a (the 'archaebacterial tree') if one neglects the root of the tree. In contrast, Lake claims that archaebacteria are paraphyletic, and he groups eocytes (extremely thermophilic, sulphur-dependent bacteria) with eukaryotes, and halobacteria with eubacteria (the 'eocyte tree', Fig. 1b). Lake's view has gained considerable support as a result of an analysis of small subunit ribosomal RNA sequence data by a new approach, the evolutionary parsimony method. Here we report that analysis of small subunit data by the neighbour-joining and maximum parasimony methods favours the archaebacterial tree and that computer simulations using either the archaebacterial or the eocyte tree as a model tree show that the probability of recovering the model tree is very high (greater than 90 per cent) for both the neighbour-joining and maximum parsimony methods but is relatively low for the evolutionary parsimony method. Moreover, analysis of large subunit rRNA sequences by all three methods strongly favours the archaebacterial tree.

Archaea↗

Date of the monocot-dicot divergence estimated from chloroplast DNA sequence data.

The divergence between monocots and dicots represents a major event in higher plant evolution, yet the date of its occurrence remains unknown because of the scarcity of relevant fossils. We have estimated this date by reconstructing phylogenetic trees from chloroplast DNA sequences, using two independent approaches: the rate of synonymous nucleotide substitution was calibrated from the divergence of maize, wheat, and rice, whereas the rate of nonsynonymous substitution was calibrated from the divergence of angiosperms and bryophytes. Both methods lead to an estimate of the monocot-dicot divergence at 200 million years (Myr) ago (with an uncertainty of about 40 Myr). This estimate is also supported by analyses of the nuclear genes encoding large and small subunit ribosomal RNAs. These results imply that the angiosperm lineage emerged in Jurassic-Triassic time, which considerably predates its appearance in the fossil record (approximately 120 Myr ago). We estimate the divergence between cycads and angiosperms to be approximately 340 Myr, which can be taken as an upper bound for the age of angiosperms.

Animals↗

Molecular phylogeny of the kingdoms Animalia, Plantae, and Fungi.

The branching order of the kingdoms Animalia, Plantae, and Fungi has been a controversial issue. Using the transformed distance method and the maximum parsimony method, we investigated this problem by comparing the sequences of several kinds of macromolecules in organisms spanning all three kingdoms. The analysis was based on the large-subunit and small-subunit ribosomal RNAs, 10 isoacceptor transfer RNA families, and six highly conserved proteins. All three sets of sequences support the same phylogenetic tree: plants and animals are sibling kingdoms that have diverged more recently than the fungi. The ribosomal RNA and protein data sets are large enough so that in both cases the inferred phylogeny is statistically significant. The present report appears to be the first to provide statistically conclusive molecular evidence for the phylogeny of the three kingdoms. The determination of this phylogeny will help us to understand the evolution of various molecular, cellular, and developmental characters shared by any two of the three kingdoms. Noting that the large-subunit rRNA sequences have evolved at similar rates in the three kingdoms, we estimated the ratio of the time since the animal-plant split to the time since the fungal divergence to be 0.90.

Animal Population Groups↗

RPA190, the gene coding for the largest subunit of yeast RNA polymerase A.

Yeast RNA polymerases are being extensively studied at the gene level. The entire gene encoding the largest subunit of RNA polymerase A, A190, was isolated and characterized in detail. Southern hybridization and gene disruption experiments showed that the RPA190 gene is unique in the haploid yeast genome and essential for cell viability. Nuclease S1 mapping was used to identify mRNA 5' and 3' termini. RPA190 encodes a polypeptide chain of 186,270 daltons in a large uninterrupted reading frame. A dot matrix comparison of the deduced amino acid sequence of subunit A190 with Escherichia coli beta' and cognate subunits B220 and C160 from yeast RNA polymerases B and C showed a conserved pattern of homology regions (I-VI). A potential DNA-binding site (zinc-binding motif) is conserved in the N-terminal region I. Remarkably, the A190 subunit does not harbor the heptapeptide repeated sequence present in the B220 subunit. The sequence of the A190 subunit diverges from B220 and C160 by the presence of two hydrophilic domains inserted between homology regions I and II, and V and VI. From their codon usage and third base pyrimidine bias, RNA polymerase genes RPA190, RPB220, RPC160, and RPC40 fall among yeast genes expressed at an average level. The RPA190 5'-flanking region contains features present in other polymerase genes that might function in regulation.

Amino Acid Sequence↗

Interfacing similarity search software with the sequence retrieval system ACNUC.

A method of interfacing sequence similarity search software with the fast sequence retrieval system ACNUC is described. The method is written in FORTRAN 77 and is straightforward to implement because no text-processing code is required--a minimum of 12 extra lines of FORTRAN provided the interface for most applications. The method is also efficient, since sequences are located by simple indexing techniques, with no linear searches of large database files necessary.

Algorithms↗

Codon contexts in enterobacterial and coliphage genes.

This investigation of the codon context of enterobacteria, plasmid, and phage protein genes was based on a search for correlations between the presence of one base type at codon position III and the presence of another base type at some other position in adjacent codons. Enterobacterial genes were compared with eukaryotic sequences for codon context effects. In enterobacterial genes, base usage at codon position III is correlated with the third position of the upstream adjacent codon and with all three positions of the downstream codon. Plasmid genes are free of context biases. Phage genes are heterogeneous: MS2 codons have no biased context, whereas lambda genes partly follow the trends of the host bacterium, and T7 genes have biased codon contexts that differ from those of the host. It has been reported that two successive third-codon positions tend to be occupied by two purines or two pyrimidines in Escherichia coli genes of low expression level. Here, the extent to which highly expressed protein genes can modulate base usage at two successive codon positions III, given the constraints on codon usage and protein sequence that act on them, was quantified. This demonstrates that the above-mentioned favored patterns are not a characteristic of weakly expressed genes but occur in all genes in which codon context can vary appreciably. The correlation between successive third-codon positions is a distinct feature of enterobacteria and of some phages, one that may result from adaptation of gene structure to translational efficiency. Conversely, codon context in yeast and human genes is biased--but for reasons unrelated to translation.

Animals↗

Evolution of the primate beta-globin gene region. High rate of variation in CpG dinucleotides and in short repeated sequences between man and chimpanzee.

A 5500 base-pair fragment including the beta-globin gene downstream from codon 122 and about 4000 base-pairs of its 5' flanking sequence was cloned from chimpanzee DNA and thoroughly sequenced before being compared with the corresponding human sequence: 88 point differences (83 substitutions and 5 deletions or insertions of 1 base-pair) were detected as well as seven more important deletion/insertion events. These changes occur preferentially in two kinds of structure. First, 40% of the CpG dinucleotides present in either human or chimpanzee sequences are affected by nucleotide variations. This corresponds to a divergence level considerably higher than that expected. Second, most short repeated sequences found in the 5' extragenic sequence are involved in mutational events (amplification or contraction of the number of basic motifs as well as point substitutions or deletions/insertions of 1 base-pair). Considering the very low level of nucleotide sequence divergence between these two closely related species, our data provide direct evidence for CpG and tandem array instability.

Animals↗

System analysis and nucleic acid sequence banks.

The mass of published nucleic acid sequence data has required the design of several computerized data bases. We show that this activity is related to the methodology of System Analysis and that data bases are a means of modeling biological knowledge. As an example, the ACNUC data base we have created is presented.

Base Sequence↗

Non-parametric statistics for nucleic acid sequence study.

The use of non-parametric statistics for nucleic acid sequence studies is illustrated by some examples. This method is highly flexible and allows design of specific tests for detecting sequence structure. Tests devoted to local repetitivity, codon nearest neighbors, and dinucleotide avoidance are discussed in detail. An appendix indicates all computations required to use these tests.

Animals↗

[Prediction of secondary structures of nucleic acids: algorithmic and physical aspects].

Prediction of secondary structures in nucleic acids requires both an adequate physical model and powerful calculation algorithms. In our approach, we cut the molecules in sections of which the contributions to the global energy are context-dependent but roughly additive. The structure of minimum energy is obtained by a tree search under constraints of binary incompatibilities. Our algorithm of the "incompatibility islets" is shown to be more powerful than the "bit parallel forward checking" algorithm, well known in Artificial Intelligence. Recurrent algorithms, proposed by other authors are even more rapid, but often miss the correct structures, for they demand a strict additivity of the energetic contributions, physically unjustified. New strategies, required to deal with molecules of more than 200 nucleotides are discussed. Our physical model has been improved by considering the special case of internal loops beginning with a G-A opposition. A bonus of 1.5 kcal. is attributed to such a feature, at each side of an internal loop. To illustrate our programs, we give the computed schemes for the 3' termini of the small subunit ribosomal RNA.

Base Sequence↗

ACNUC--a portable retrieval system for nucleic acid sequence databases: logical and physical designs and usage.

ACNUC is a database structure and retrieval software for use with either the GenBank or EMBL nucleic acid sequence data collections. The nucleotide and textual data furnished by both collections are each restructured into a database that allows sequence retrieval on a multi-criterion basis. The main selection criteria are: species (or higher order taxon), keyword, reference, journal, author, and organelle; all logical combinations of these criteria can be used. Direct access to sequence regions that code for a specific product (protein, tRNA or rRNA) is provided. A versatile extraction procedure copies selected sequences, or fragments of them, from the database to user files suitable to be analysed by user-supplied application programs. A detailed help mechanism is provided to aid the user at any time during the retrieval session. All software has been written in FORTRAN 77 which guarantees a high degree of transportability to minicomputers or mainframes.

Base Sequence↗

ACNUC: a nucleic acid sequence data base and analysis system.

Structured as a data base and associated with data analysis tools, ACNUC allows both on-line access to a central computer and local exploitation of published nucleotide sequences. Its data retrieval capabilities seem to be presently the most powerful available.

Base Sequence↗

An energy model that predicts the correct folding of both the tRNA and the 5S RNA molecules.

A new set of energy values to predict the secondary structures in RNA molecules has been derived through a multiple-step refinement procedure. It achieves more than 80% success in predicting the cloverleaf pattern in tRNA (200 sequences tested) and more than 60% success in predicting the consensus folding of 5S RNA (100 sequences). Improvements in our initial program for predicting secondary structures, based on the principle of the "incompatibility islets" made possible the work on 5S RNA. The program was speeded up by introducing a dynamic grouping of the islets into three disjoint blocks. The novel features in the energy model include i) an evaluation of the contribution of odd pairs according to their position within a segment ii) a penalty for internal loops related to their dissymmetry iii) a bonus for bulge loops when the two terminal paired bases at the junction point are both pyrimidines.

Computers↗

Codon usage in bacteria: correlation with gene expressivity.

The nucleic acid sequence bank now contains over 600 protein coding genes of which 107 are from prokaryotic organisms. Codon frequencies in each new prokaryotic gene are given. Analysis of genetic code usage in the 83 sequenced genes of the Escherichia coli genome (chromosome, transposons and plasmids) is presented, taking into account new data on gene expressivity and regulation as well as iso-tRNA specificity and cellular concentration. The codon composition of each gene is summarized using two indexes: one is based on the differential usage of iso-tRNA species during gene translation, the other on choice between Cytosine and Uracil for third base. A strong relationship between codon composition and mRNA expressivity is confirmed, even for genes transcribed in the same operon. The influence of codon use of peptide elongation rate and protein yield is discussed. Finally, the evolutionary aspect of codon selection in mRNA sequences is studied.

Amino Acid Sequence↗

Codon catalog usage is a genome strategy modulated for gene expressivity.

The nucleic acid sequence bank now contains 161 mRNAs, 43 new genes are added. One sequence, that of B. mori fibroin, is dropped due to uncertainty on the starting point for translation. Frequencies of all codons are given for each gene added and for each genome type in the total bank. A new series of correspondence analyses on codon use is presented, substantiating the genome hypothesis. Internal regulation of mRNA expression by different third base choices between quartet and duet codons is proposed for bacterial genes.

Amino Acid Sequence↗

Codon frequencies in 119 individual genes confirm consistent choices of degenerate bases according to genome type.

The poor printing of our previous Figure 2 (1) is corrected. Codon usage in mRNA sequences just published is also given. A new correspondence analysis is done, based on simultaneous comparison in all mRNA of use of the 61 codons. This analysis reinforces our claim that most genes in a genome, or genome type, have the same coding strategy; that is, they show similar choices among synonymous codons, or among degenerate bases (2). Like analysis on frequency variation in the amino acids coded reveals an entirely different pattern.

Animals↗

Codon catalog usage and the genome hypothesis.

Frequencies for each of the 61 amino acid codons have been determined in every published mRNA sequence of 50 or more codons. The frequencies are shown for each kind of genome and for each individual gene. A surprising consistency of choices exists among genes of the same or similar genomes. Thus each genome, or kind of genome, appears to possess a "system" for choosing between codons. Frameshift genes, however, have widely different choice strategies from normal genes. Our work indicates that the main factors distinguishing between mRNA sequences relate to choices among degenerate bases. These systematic third base choices can therefore be used to establish a new kind of genetic distance, which reflects differences in coding strategy. The choice patterns we find seem compatible with the idea that the genome and not the individual gene is the unit of selection. Each gene in a genome tends to conform to its species' usage of the codon catalog; this is our genome hypothesis.

Animals↗