Search PubMed⌕ Search

Biomedical subjects

Masatoshi Nei

Publications and source records attributed to Masatoshi Nei.

34 records · Page 2Linked to original sources

Type I MADS-box genes have experienced faster birth-and-death evolution than type II MADS-box genes in angiosperms.

Plant MADS-box genes form a large gene family for transcription factors and are involved in various aspects of developmental processes, including flower development. They are known to be subject to birth-and-death evolution, but the detailed features of this mode of evolution remain unclear. To have a deeper insight into the evolutionary pattern of this gene family, we enumerated all available functional and nonfunctional (pseudogene) MADS-box genes from the Arabidopsis and rice genomes. Plant MADS-box genes can be classified into types I and II genes on the basis of phylogenetic analysis. Conducting extensive homology search and phylogenetic analysis, we found 64 presumed functional and 37 nonfunctional type I genes and 43 presumed functional and 4 nonfunctional type II genes in Arabidopsis. We also found 24 presumed functional and 6 nonfunctional type I genes and 47 presumed functional and 1 nonfunctional type II genes in rice. Our phylogenetic analysis indicated there were at least about four to eight type I genes and approximately 15-20 type II genes in the most recent common ancestor of Arabidopsis and rice. It has also been suggested that type I genes have experienced a higher rate of birth-and-death evolution than type II genes in angiosperms. Furthermore, the higher rate of birth-and-death evolution in type I genes appeared partly due to a higher frequency of segmental gene duplication and weaker purifying selection in type I than in type II genes.

Arabidopsis↗

MEGA3: Integrated software for Molecular Evolutionary Genetics Analysis and sequence alignment.

With its theoretical basis firmly established in molecular evolutionary and population genetics, the comparative DNA and protein sequence analysis plays a central role in reconstructing the evolutionary histories of species and multigene families, estimating rates of molecular evolution, and inferring the nature and extent of selective forces shaping the evolution of genes and genomes. The scope of these investigations has now expanded greatly owing to the development of high-throughput sequencing techniques and novel statistical and computational methods. These methods require easy-to-use computer programs. One such effort has been to produce Molecular Evolutionary Genetics Analysis (MEGA) software, with its focus on facilitating the exploration and analysis of the DNA and protein sequence variation from an evolutionary perspective. Currently in its third major release, MEGA3 contains facilities for automatic and manual sequence alignment, web-based mining of databases, inference of the phylogenetic trees, estimation of evolutionary distances and testing evolutionary hypotheses. This paper provides an overview of the statistical methods, computational tools, and visual exploration modules for data input and the results obtainable in MEGA.

Databases, Genetic↗

Concerted and nonconcerted evolution of the Hsp70 gene superfamily in two sibling species of nematodes.

We have identified the Hsp70 gene superfamily of the nematode Caenorhabditis briggsae and investigated the evolution of these genes in comparison with Hsp70 genes from C. elegans, Drosophila, and yeast. The Hsp70 genes are classified into three monophyletic groups according to their subcellular localization, namely, cytoplasm (CYT), endoplasmic reticulum (ER), and mitochondria (MT). The Hsp110 genes can be classified into the polyphyletic CYT group and the monophyletic ER group. The different Hsp70 and Hsp110 groups appeared to evolve following the model of divergent evolution. This model can also explain the evolution of the ER and MT genes. On the other hand, the CYT genes are divided into heat-inducible and constitutively expressed genes. The constitutively expressed genes have evolved more or less following the birth-and-death process, and the rates of gene birth and gene death are different between the two nematode species. By contrast, some heat-inducible genes show an intraspecies phylogenetic clustering. This suggests that they are subject to sequence homogenization resulting from gene conversion-like events. In addition, the heat-inducible genes show high levels of sequence conservation in both intra-species and inter-species comparisons, and in most cases, amino acid sequence similarity is higher than nucleotide sequence similarity. This indicates that purifying selection also plays an important role in maintaining high sequence similarity among paralogous Hsp70 genes. Therefore, we suggest that the CYT heat-inducible genes have been subjected to a combination of purifying selection, birth-and-death process, and gene conversion-like events.

Animals↗

Evolution of olfactory receptor genes in the human genome.

Olfactory receptor (OR) genes form the largest known multigene family in the human genome. To obtain some insight into their evolutionary history, we have identified the complete set of OR genes and their chromosomal locations from the latest human genome sequences. We detected 388 potentially functional genes that have intact ORFs and 414 apparent pseudogenes. The number and the fraction (48%) of functional genes are considerably larger than the ones previously reported. The human OR genes can clearly be divided into class I and class II genes, as was previously noted. Our phylogenetic analysis has shown that the class II OR genes can further be classified into 19 phylogenetic clades supported by high bootstrap values. We have also found that there are many tandem arrays of OR genes that are phylogenetically closely related. These genes appear to have been generated by tandem gene duplication. However, the relationships between genomic clusters and phylogenetic clades are very complicated. There are a substantial number of cases in which the genes in the same phylogenetic clade are located on different chromosomal regions. In addition, OR genes belonging to distantly related phylogenetic clades are sometimes located very closely in a chromosomal region and form a tight genomic cluster. These observations can be explained by the assumption that several chromosomal rearrangements have occurred at the regions of OR gene clusters and the OR genes contained in different genomic clusters are shuffled.

Biological Evolution↗

Antiquity and evolution of the MADS-box gene family controlling flower development in plants.

MADS-box genes in plants control various aspects of development and reproductive processes including flower formation. To obtain some insight into the roles of these genes in morphological evolution, we investigated the origin and diversification of floral MADS-box genes by conducting molecular evolutionary genetics analyses. Our results suggest that the most recent common ancestor of today's floral MADS-box genes evolved roughly 650 MYA, much earlier than the Cambrian explosion. They also suggest that the functional classes T (SVP), B (and Bs), C, F (AGL20 or TM3), A, and G (AGL6) of floral MADS-box genes diverged sequentially in this order from the class E gene lineage. The divergence between the class G and E genes apparently occurred around the time of the angiosperm/gymnosperm split. Furthermore, the ancestors of three classes of genes (class T genes, class B/Bs genes, and the common ancestor of the other classes of genes) might have existed at the time of the Cambrian explosion. We also conducted a phylogenetic analysis of MADS-domain sequences from various species of plants and animals and presented a hypothetical scenario of the evolution of MADS-box genes in plants and animals, taking into account paleontological information. Our study supports the idea that there are two main evolutionary lineages (type I and type II) of MADS-box genes in plants and animals.

Animals↗

Birth-and-death evolution in primate MHC class I genes: divergence time estimates.

The major histocompatibility complex (MHC) is a multigene family that mediates the host immune response by helping T lymphocytes to recognize and respond to foreign antigens. The high degree of polymorphism and a quick turnover of the genetic loci make the evolution of MHC genes an intriguing subject of study. To understand the evolutionary pattern of this multigene family, we studied the phylogeny and divergence times of six functional MHC class I loci from primate species. On the phylogenetic trees, locus F occupies the most basal position among these loci. Our results suggest that the F locus diverged from the other MHC class I loci about 46-66 MYA. The major diversification of the other class I loci was estimated to have occurred at about 35-49 MYA, which is before the time of separation of Old World-New World monkeys. The gene duplication leading to the classical C locus in great apes appears to have occurred about 21-28 MYA. At approximately the same time the duplication of the B locus occurred in macaques. The oldest allelic lineages of A, B, and C loci in humans seem to have appeared at least 14-19, 10-15, and 13-17 MYA, respectively. Our phylogenetic analysis supports the hypothesis that the nonclassical locus F has diverged from the rest of class I loci very early in primate evolution. The overall phylogenetic pattern observed among class I genes is consistent with the model of birth-and-death evolution.

Animals↗

Reanalysis of Murphy et al.'s data gives various mammalian phylogenies and suggests overcredibility of Bayesian trees.

Murphy and colleagues reported that the mammalian phylogeny was resolved by Bayesian phylogenetics. However, the DNA sequences they used had many alignment gaps and undetermined nucleotide sites. We therefore reanalyzed their data by minimizing unshared nucleotide sites and retaining as many species as possible (13 species). In constructing phylogenetic trees, we used the Bayesian, maximum likelihood (ML), maximum parsimony (MP), and neighbor-joining (NJ) methods with different substitution models. These trees were constructed by using both protein and DNA sequences. The results showed that the posterior probabilities for Bayesian trees were generally much higher than the bootstrap values for ML, MP, and NJ trees. Two different Bayesian topologies for the same set of species were sometimes supported by high posterior probabilities, implying that two different topologies can be judged to be correct by Bayesian phylogenetics. This suggests that the posterior probability in Bayesian analysis can be excessively high as an indication of statistical confidence and therefore Murphy et al.'s tree, which largely depends on Bayesian posterior probability, may not be correct.

Animals↗

Estimation of divergence times for major lineages of primate species.

Although the phylogenetic relationships of major lineages of primate species are relatively well established, the times of divergence of these lineages as estimated by molecular data are still controversial. This controversy has been generated in part because different authors have used different types of molecular data, different statistical methods, and different calibration points. We have therefore examined the effects of these factors on the estimates of divergence times and reached the following conclusions: (1) It is advisable to concatenate many gene sequences and use a multigene gamma distance for estimating divergence times rather than using the individual gene approach. (2) When sequence data from many nuclear genes are available, protein sequences appear to give more robust estimates than DNA sequences. (3) Nuclear proteins are generally more suitable than mitochondrial proteins for time estimation. (4) It is important first to construct a phylogenetic tree for a group of species using some outgroups and then estimate the branch lengths. (5) It appears to be better to use a few reliable calibration points rather than many unreliable ones. Considering all these factors and using two calibration points, we estimated that the human lineage diverged from the chimpanzee, gorilla, orangutan, Old World monkey, and New World monkey lineages approximately 6 MYA (with a range of 5-7), 7 MYA (range, 6-8), 13 MYA (range, 12-15), 23 MYA (range, 21-25), and 33 MYA (range 32-36).

Animals↗

Overcredibility of molecular phylogenies obtained by Bayesian phylogenetics.

Bayesian phylogenetics has recently been proposed as a powerful method for inferring molecular phylogenies, and it has been reported that the mammalian and some plant phylogenies were resolved by using this method. The statistical confidence of interior branches as judged by posterior probabilities in Bayesian analysis is generally higher than that as judged by bootstrap probabilities in maximum likelihood analysis, and this difference has been interpreted as an indication that bootstrap support may be too conservative. However, it is possible that the posterior probabilities are too high or too liberal instead. Here, we show by computer simulation that posterior probabilities in Bayesian analysis can be excessively liberal when concatenated gene sequences are used, whereas bootstrap probabilities in neighbor-joining and maximum likelihood analyses are generally slightly conservative. These results indicate that bootstrap probabilities are more suitable for assessing the reliability of phylogenetic trees than posterior probabilities and that the mammalian and plant phylogenies may not have been fully resolved.

Amino Acid Substitution↗

Acceleration of genomic evolution caused by enhanced mutation rate in endocellular symbionts.

Endosymbionts, which are widely observed in nature, have undergone reductive genome evolution because of their long-term intracellular lifestyle. Here we compared the complete genome sequences of two different endosymbionts, Buchnera and a protist mitochondrion, with their close relatives to study the evolutionary rates of functional genes in endosymbionts. The results indicate that the rate of amino acid substitution is two times higher in symbionts than in their relatives. This rate increase was observed uniformly among different functional classes of genes, although strong purifying selection may have counterbalanced the rate increase in a few cases. Our data suggest that, contrary to current views, neither the Muller's ratchet effect nor the slightly deleterious mutation theory sufficiently accounts for the elevated evolutionary rate. Rather, the elevated evolutionary rate appears to be mainly due to enhanced mutation rate, although the possibility of relaxation of purifying selection cannot be ruled out.

Buchnera↗

Molecular evolution of the nontandemly repeated genes of the histone 3 multigene family.

In some species, histone gene clusters consist of tandem arrays of each type of histone gene, whereas in other species the genes may be clustered but not arranged in tandem. In certain species, however, histone genes are found scattered across several different chromosomes. This study examines the evolution of histone 3 (H3) genes that are not arranged in large clusters of tandem repeats. Although H3 amino acid sequences are highly conserved both within and between species, we found that the nucleotide sequence divergence at synonymous sites is high, indicating that purifying selection is the major force for maintaining H3 amino acid sequence homogeneity over long-term evolution. In cases where synonymous-site divergence was low, recent gene duplication appeared to be a better explanation than gene conversion. These results, and other observations on gene inactivation, organization, and phylogeny, indicated that these H3 genes evolve according to a birth-and-death process under strong purifying selection. Thus, we found little evidence to support previous claims that all H3 proteins, regardless of their genome organization, undergo concerted evolution. Further analyses of the structure of H3 proteins revealed that the histones of higher eukaryotes might have evolved from a replication-independent-like H3 gene.

Amino Acid Sequence↗

Simulation study of the reliability and robustness of the statistical methods for detecting positive selection at single amino acid sites.

Inferring positive selection at single amino acid sites is of biological and medical importance. Parsimony-based and likelihood-based methods have been developed for this purpose, but the reliabilities of these methods are not well understood. Because the evolutionary models assumed in these methods are only rough approximations to reality, it is desirable that the methods are not very sensitive to violation of the assumptions made. In this study we show by computer simulation that the likelihood-based method is sensitive to violation of the assumptions and produces many false-positive results under certain conditions, whereas the parsimony-based method tends to be conservative. These observations, together with those from previous studies, suggest that the positively selected sites inferred by the parsimony-based method are more reliable than those inferred by the likelihood-based method.

Amino Acid Sequence↗

Adaptive evolution of variable region genes encoding an unusual type of immunoglobulin in camelids.

A typical immunoglobulin (Ig) molecule is composed of four polypeptide chains: two identical heavy (H) chains and two identical light (L) chains. This tetrameric structure is conserved in almost all jawed vertebrate species. However, it has been discovered that camels and llamas (family: Camelidae) possess a type of dimeric Ig that consists of two H chains only. These H chains do not associate with L chains, and they do not have the first constant region (CH1), which is present in the conventional Ig. In spite of these changes, the dimeric Ig maintains the normal immune function. To understand the evolution of the dimeric Ig, we studied the phylogenetic relationships of the variable region (V(H)H) genes of the dimeric Ig from Camelidae and those (V(H)) of the conventional Ig from mammals. The results showed that the V(H)H genes form a monophyletic cluster within one of the mammalian V(H) groups, group C. We examined the type of selective force in complementarity-determining regions (CDRs) and framework regions (FRs) by comparing the rate of synonymous (dS) and nonsynonymous (dN) substitutions. We found that the results obtained from V(H)H genes were similar to those from V(H) genes in that CDRs showed an excess of dN over dS (indicating positive selection), whereas the reverse was true for FRs (purifying selection). However, when the extent of positive selection or purifying selection was investigated at each codon site, three major differences between V(H)H and V(H) genes were found. That is, very different types of selective force were observed between V(H)H and V(H) genes (1) at the sites that contact the L chain in the conventional Ig, (2) at the sites that interact with the CH1 region in the conventional Ig, and (3) in the H1 loop. Our findings suggest that adaptive evolution has occurred in the functionally important sites of the V(H)H genes to maintain the normal immune function in the dimeric Ig.

Amino Acid Substitution↗

Origin and evolution of influenza virus hemagglutinin genes.

Influenza A, B, and C viruses are the etiological agents of influenza. Hemagglutinin (HA) is the major envelope glycoprotein of influenza A and B viruses, and hemagglutinin-esterase (HE) in influenza C viruses is a protein homologous to HA. Because influenza A virus pandemics in humans appear to occur when new subtypes of HA genes are introduced from aquatic birds that are known to be the natural reservoir of the viruses, an understanding of the origin and evolution of HA genes is of particular importance. We therefore conducted a phylogenetic analysis of HA and HE genes and showed that the influenza A and B virus HA genes diverged much earlier than the divergence between different subtypes of influenza A virus HA genes. The rate of amino acid substitution for A virus HAs from duck, a natural reservoir, was estimated to be 3.19 x 10(-4) per site per year, which was slower than that for human and swine A virus HAs but similar to that for influenza B and C virus HAs (HEs). Using this substitution rate from the duck, we estimated that the divergences between different subtypes of A virus HA genes occurred from several thousand to several hundred years ago. In particular, the earliest divergence time was estimated to be about 2,000 years ago. Also, the A virus HA gene diverged from the B virus HA gene about 4,000 years ago and from the C virus HE gene about 8,000 years ago. These time estimates are much earlier than the previous ones.

Amino Acid Sequence↗

Purifying selection and birth-and-death evolution in the histone H4 gene family.

Histones are small basic proteins encoded by a multigene family and are responsible for the nucleosomal organization of chromatin in eukaryotes. Because of the high degree of protein sequence conservation, it is generally believed that histone genes are subject to concerted evolution. However, purifying selection can also generate a high degree of sequence homogeneity. In this study, we examined the long-term evolution of histone H4 genes to determine whether concerted evolution or purifying selection was the major factor for maintaining sequence homogeneity. We analyzed the proportion (p(S)) of synonymous nucleotide differences between the H4 genes from 59 species of fungi, plants, animals, and protists and found that p(S) is generally very high and often close to the saturation level (p(S) ranging from 0.3 to 0.6) even though protein sequences are virtually identical for all H4 genes. A small proportion of genes showed a low level of p(S) values, but this appeared to be caused by recent gene duplication. Our findings suggest that the members of this gene family evolve according to the birth-and-death model of evolution under strong purifying selection. Using histone-like genes in archaebacteria as outgroups, we also showed that H1, H2A, H2B, H3, and H4 histone genes in eukaryotes form separate clusters and that these classes of genes diverged nearly at the same time, before the eukaryotic kingdoms diverged.

Animals↗

The Wilhelmine E. Key 2001 Invitational Lecture. Estimation of divergence times for a few mammalian and several primate species.

Statistical methods for estimating divergence times by using multiprotein gamma distances are discussed. When a large number of proteins are used, even a small degree of deviation from the molecular clock hypothesis can be detected. In this case, one may use the stem-lineage method for estimating divergence times. However, the estimates obtained by this method are often similar to those obtained by the linearized tree method. Application of these methods to a dataset of 104 proteins from several vertebrate species indicated that the divergence times between humans and mice and between mice and rats are about 96 and 33 million years (MY) ago, respectively. These estimates were obtained by assuming that birds and mammals diverged 310 MY ago. Similarly application of the methods to the protein sequence data from primate species indicated that the human lineage separated from the chimpanzee, gorilla, Old World monkeys, and New World monkeys about 6.0, 7.0, 23.0, and 33.0 MY ago, respectively. In this case the use of two calibration points, that is, the divergence time (13 MY ago) between humans and orangutans and between primates and artiodactyls (90 MY ago) gave essentially the same estimates.

Animals↗