Search PubMed⌕ Search

Biomedical subjects

P Liò

Publications and source records attributed to P Liò.

17 recordsLinked to original sources

Analysing gene function after duplication.

After gene duplication, mutations cause the gene copies to diverge. The classical model predicts that these mutations will generally lead to the loss of function of one gene copy; rarely, new functions will be created and both duplicate genes are conserved. In contrast, under the subfunctionalization model both duplicates are preserved due to the partition of different functions between the duplicates. A recent study provides support for the subfunctionalization model, identifying several expressed gene duplicates common to humans and mice that contain regions conserved in one duplicate but variable in the other (and vice versa). We discuss both the methodology used in this study and also how gene phylogeny may lead to additional evidence for the importance of subfunctionalization in the evolution of new genes.

Animals↗

Molecular phylogenetics: state-of-the-art methods for looking into the past.

As the amount of molecular sequence data in the public domain grows, so does the range of biological topics that it influences through evolutionary considerations. In recent years, a number of developments have enabled molecular phylogenetic methodology to keep pace. Likelihood-based inferential techniques, although controversial in the past, lie at the heart of these new methods and are producing the promised advances in the understanding of sequence evolution. They allow both a wide variety of phylogenetic inferences from sequence data and robust statistical assessment of all results. It cannot remain acceptable to use outdated data analysis techniques when superior alternatives exist. Here, we discuss the most important and exciting methods currently available to the molecular phylogeneticist.

Animals↗

Molecular evolution of nitrogen fixation: the evolutionary history of the nifD, nifK, nifE, and nifN genes.

The pairs of nitrogen fixation genes nifDK and nifEN encode for the alpha and beta subunits of nitrogenase and for the two subunits of the NifNE protein complex, involved in the biosynthesis of the FeMo cofactor, respectively. Comparative analysis of the amino acid sequences of the four NifD, NifK, NifE, and NifN in several archaeal and bacterial diazotrophs showed extensive sequence similarity between them, suggesting that their encoding genes constitute a novel paralogous gene family. We propose a two-step model to reconstruct the possible evolutionary history of the four genes. Accordingly, an ancestor gene gave rise, by an in-tandem paralogous duplication event followed by divergence, to an ancestral bicistronic operon; the latter, in turn, underwent a paralogous operon duplication event followed by evolutionary divergence leading to the ancestors of the present-day nifDK and nifEN operons. Both these paralogous duplication events very likely predated the appearance of the last universal common ancestor. The possible role of the ancestral gene and operon in nitrogen fixation is also discussed.

Dinitrogenase Reductase↗

Finding pathogenicity islands and gene transfer events in genome data.

MOTIVATION: There is a growing literature on wavelet theory and wavelet methods showing improvements on more classical techniques, especially in the contexts of smoothing and extraction of fundamental components of signals. G+C patterns occur at different lengths (scales) and, for this reason, G+C plots are usually difficult to interpret. Current methods for genome analysis choose a window size and compute a chi(2) statistics of the average value for each window with respect to the whole genome. RESULTS: Firstly, wavelets are used to smooth G+C profiles to locate characteristic patterns in genome sequences. The method we use is based on performing a chi(2) statistics on the wavelet coefficients of a profile; thus we do not need to choose a fixed window size, in that the smoothing occurs at a set of different scales. Secondly, a wavelet scalogram is used as a measure for sequence profile comparison; this tool is very general and can be applied to other sequence profiles commonly used in genome analysis. We show applications to the analysis of Deinococcus radiodurans chromosome I, of two strains of Helicobacter pylori (26695, J99) and two of Neisseria meningitidis (serogroup B strain MC58 and serogroup A strain Z2491). We report a list of loci that have different G+C content with respect to the nearby regions; the analysis of N. meningitidis serogroup B shows two new large regions with low G+C content that are putative pathogenicity islands. AVAILABILITY: Software and numerical results (profiles, scalograms, high and low frequency components) for all the genome sequences analyzed are available upon request from the authors.

Base Composition↗

Molecular nature of RAPD markers from Haemophilus influenzae Rd genome.

Despite the widespread application of the random amplified polymorphic DNA (RAPD) technique, there is no experimental evidence of the molecular mechanism of random amplification starting from a complex template. To investigate this mechanism, we cloned and sequenced 23 selected RAPD bands amplified from Haemophilus influenzae Rd genomic DNA using eight decamer primers different in GC content and/or nucleotide sequence. As the whole genome sequence of H. influenzae Rd has been reported, the exact nucleotide sequence of each primer-template annealing site was identified. Results showed that, on an average, a homology of eight base pairs was involved in priming events and that the number of nonhomologous base pairings declined exponentially from the 5' end of the primer to its 3' end. The interaction between the primer and the template DNA was stabilized by the formation of secondary structures, and a perfect match of the 3' terminal region of the primer was not necessary for successful amplification. The complexity of the annealing process suggested that, in the studied reaction conditions, many primer-template annealing sites were extended in the first cycles and that differences in the efficiency of priming and replication processes led to amplification of RAPD fragments. Moreover, the distribution of the amplified regions on the H. influenzae chromosome was analyzed.

Base Sequence↗

Using protein structural information in evolutionary inference: transmembrane proteins.

We present a model of amino acid sequence evolution based on a hidden Markov model that extends to transmembrane proteins previous methods that incorporate protein structural information into phylogenetics. Our model aims to give a better understanding of processes of molecular evolution and to extract structural information from multiple alignments of transmembrane sequences and use such information to improve phylogenetic analyses. This should be of value in phylogenetic studies of transmembrane proteins: for example, mitochondrial proteins have acquired a special importance in phylogenetics and are mostly transmembrane proteins. The improvement in fit to example data sets of our new model relative to less complex models of amino acid sequence evolution is statistically tested. To further illustrate the potential utility of our method, phylogeny estimation is performed on primate CCR5 receptor sequences, sequences of l and m subunits of the light reaction center in purple bacteria, guinea pig sequences with respect to lagomorph and rodent sequences of calcitonin receptor and K-substance receptor, and cetacean sequences of cytochrome b.

Animals↗

PASSML: combining evolutionary inference and protein secondary structure prediction.

MOTIVATION: Evolutionary models of amino acid sequences can be adapted to incorporate structure information; protein structure biologists can use phylogenetic relationships among species to improve prediction accuracy. Results : A computer program called PASSML ('Phylogeny and Secondary Structure using Maximum Likelihood') has been developed to implement an evolutionary model that combines protein secondary structure and amino acid replacement. The model is related to that of Dayhoff and co-workers, but we distinguish eight categories of structural environment: alpha helix, beta sheet, turn and coil, each further classified according to solvent accessibility, i.e. buried or exposed. The model of sequence evolution for each of the eight categories is a Markov process with discrete states in continuous time, and the organization of structure along protein sequences is described by a hidden Markov model. This paper describes the PASSML software and illustrates how it allows both the reconstruction of phylogenies and prediction of secondary structure from aligned amino acid sequences. AVAILABILITY: PASSML 'ANSI C' source code and the example data sets described here are available at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html and 'downstream' Web pages. CONTACT: P.Lio@gen.cam.ac.uk

Adenylate Kinase↗

Models of molecular evolution and phylogeny.

Phylogenetic reconstruction is a fast-growing field that is enriched by different statistical approaches and by findings and applications in a broad range of biological areas. Fundamental to these are the mathematical models used to describe the patterns of DNA base substitution and amino acid replacement. These may become some of the basic models for comparative genome research. We discuss these models, including the analysis of observed DNA base and amino acid mutation patterns, the concept of site heterogeneity, and the incorporation of structural biology data, all of which have become particularly important in recent years. We also describe the use of such models in phylogenetic reconstruction and statistical methods for the comparison of different models.

Amino Acid Substitution↗

Paralogous histidine biosynthetic genes: evolutionary analysis of the Saccharomyces cerevisiae HIS6 and HIS7 genes.

The HIS6 gene from Saccharomyces cerevisiae strain YNN282 is able to complement both the S. cerevisiae his6 and the Escherichia coli hisA mutations. The cloning and the nucleotide sequence indicated that this gene encodes a putative phosphoribosyl-5-amino-1-phosphoribosyl-4-imidazolecarboxiamide isomerase (5' Pro-FAR isomerase, EC 5.3.1.16) of 261 amino acids, with a molecular weight of 29,554. The HIS6 gene product shares a significant degree of sequence similarity with the prokaryotic HisA proteins and HisF proteins, and with the C-terminal domain of the S. cerevisiae HIS7 protein (homologous to HisF), indicating that the yeast HIS6 and HIS7 genes are paralogous. Moreover, the HIS6 gene is organized into two homologous modules half the size of the entire gene, typical of all the known prokaryotic hisA and hisF genes. The structure of the yeast HIS6 gene supports the two-step evolutionary model suggested by Fani et al. (J. Mol. Evol. 1994; 38: 489-495) to explain the present-day hisA and hisF genes. According to this idea, the hisF gene originated from the duplication of an ancestral hisA gene which, in turn, was the result of an earlier gene elongation event involving an ancestral module half the size of the extant gene. Results reported in this paper also suggest that these two successive paralogous gene duplications took probably place in the early steps of molecular evolution of the histidine pathway, well before the diversification of the three domains, and that this pathway was one of the metabolic activities of the last common ancestor. The molecular evolution of the yeast HIS6 and HIS7 genes is also discussed.

Aldose-Ketose Isomerases↗

Comparison of parametric and nonparametric methods to map oligogenes by linkage.

A sample of 95 sib pairs affected with insulin-dependent diabetes and typed with their normal parents for 28 markers on chromosome 6 has been analyzed by several methods. When appropriate parameters are efficiently estimated, a parametric model is equivalent to the beta model, which is superior to nonparametric alternatives both in single point tests (as found previously) and in multipoint tests. Theory is given for meta-analysis combined with allelic association, and problems that may be associated with errors of map location and/or marker typing are identified. Reducing by multipoint analysis the number of association tests in a dense map can give a 3-fold reduction in the critical lod, and therefore in the cost of positional cloning.

Adult↗

A physiological and molecular analysis of the genus Nicotiana.

An analysis of the evolution of the genus Nicotiana was carried out with physiological and molecular tools. The capacity of explants from seedlings of several species of Nicotiana to differentiate roots or shoots or to habituate was used to ascertain whether the in vitro behavior of species has a nonrandom distribution in the genus. The results obtained allowed us to identify two groups of species, one root-forming prone composed of Paniculatae (subgenus Rustica) and the other composed of Alatae, Repandae, and Noctiflorae (subgenus Petunioides), with a major tendency toward the production of shoots. Habituation capacity was characteristic of species randomly distributed throughout the phylogenetic tree. These data suggest fixation throughout the evolution of coadapted gene complexes (hormone-related genes) involved in the control of developmental processes. RAPDs, on the other hand, used as molecular markers for the clustering of related species, seem entirely coherent both with classical morphological and karyological studies and with in vitro physiological methods, supporting an early subdivision of the whole genus into two diverging developmental patterns.

Adaptation, Physiological↗

Analysis of genomic patchiness of Haemophilus influenzae and Saccharomyces cerevisiae chromosomes.

We have analysed some aspects of the primary structure of the chromosome of the prokaryote Haemophilus influenzae and of the eukaryote Saccharomyces cerevisiae that share the same G + C content. In particular, we have investigated genomic patchiness over the gene size level (10 Kb) and that patchiness due to long homogenous tracts. Long polypurine and polypyrmidine tracts that are largely over-represented in S. cerevisiae chromosomes and under-represented in H. influenzae, are responsible for a large fraction of long correlation signals. Generating mechanisms of long homogenous tracts are DNA replication slippage and duplication events that appear to be linked processes driving chromosome primary structure evolution.

Base Sequence↗

High statistics block entropy measures of DNA sequences.

We have used an improved block-entropy measure in order to gain some further insights into the short-range correlations present in whole chromosomes of S. cerevisiae, viruses and organelles and very large genomic regions of E. coli. Although DNA sequences are largely inhomogeneous and word frequencies are unevenly distributed, the comparison of entire chromosomes and large genomic regions show a "bulk" composition homogeneity. This property suggests that biases in selection, directional mutational pressure and recombination processes act in homogenizing the base composition of the DNA molecules within a genome but their mode of action, relative impact and direction may vary in different organisms. The most interesting results appear to be the differences between the SW (C,G/A,T) and RY (A,G/C,T) two-letter alphabet entropies. Deviations from randomness in E. coli and S. cerevisiae sequences particularly concern SW dinucleotide frequencies and RY tetranucleotide frequencies.

Base Sequence↗

Selection, mutations and codon usage in a bacterial model.

We present a statistical model of bacterial evolution based on the coupling between codon usage and tRNA abundance. Such a model interprets this aspect of the evolutionary process as a balance between the codon homogenization effect due to mutation process and the improvement of the translation phase due to natural selection. We develop a thermodynamical description of the asymptotic state of the model. The analysis of naturally occurring sequences shows that the effect of natural selection on codon bias affects genes whose products are largely required at maximal growth rate conditions or undergo rapid transient increases.

Bacteria↗

Molecular evolution of the histidine biosynthetic pathway.

The available sequences of genes encoding the enzymes associated with histidine biosynthesis suggest that this is an ancient metabolic pathway that was assembled prior to the diversification of the Bacteria, Archaea, and Eucarya. Paralogous duplications, gene elongation, and fusion events involving different his genes have played a major role in shaping this biosynthetic route. Evidence that the hisA and the hisF genes and their homologous are the result of two successive duplication events that apparently took place before the separation of the three cellular lineages is extended. These two successive gene duplication events as well as the homology between the hisH genes and the sequences encoding the TrpG-type amidotransferases support the idea that during the early stages of metabolic evolution at least parts of the histidine biosynthetic pathway were mediated by enzymes of broader substrate specificities. Maximum likelihood trees calculated for the available sequences of genes encoding these enzymes have been obtained. Their topologies support the possibility of an evolutionary proximity of archaebacteria with low GC Gram-positive bacteria. This observation is consistent with those detected by other workers using the sequences of heat-shock proteins (HSP70), glutamine synthetases, glutamate dehydrogenases, and carbamoylphosphate synthetases.

Aldose-Ketose Isomerases↗

The evolution of the histidine biosynthetic genes in prokaryotes: a common ancestor for the hisA and hisF genes.

The hisA and hisF genes belong to the histidine operon that has been extensively studied in the enterobacteria Escherichia coli and Salmonella typhimurium where the hisA gene codes for the phosphoribosyl-5-amino-1-phosphoribosyl-4-imidazolecarboxamide isomerase (EC 5.3.1.16) catalyzing the fourth step of the histidine biosynthetic pathway, and the hisF gene codes for a cyclase catalyzing the sixth reaction. Comparative analysis of nucleotide and predicted amino acid sequence of hisA and hisF genes in different microorganisms showed extensive sequence homology (43% considering similar amino acids), suggesting that the two genes arose from an ancestral gene by duplication and subsequent evolutionary divergence. A more detailed analysis, including mutual information, revealed an internal duplication both in hisA and hisF genes in each of the considered microorganisms. We propose that the hisA and hisF have originated from the duplication of a smaller ancestral gene corresponding to half the size of the actual genes followed by rapid evolutionary divergence. The involvement of gene elongation, gene duplication, and gene fusion in the evolution of the histidine biosynthetic genes is also discussed.

Aldose-Ketose Isomerases↗