Search PubMed⌕ Search

Biomedical subjects

Masatoshi Nei

Publications and source records attributed to Masatoshi Nei.

At least 19 recordsLinked to original sources

Rapid expansion of killer cell immunoglobulin-like receptor genes in primates and their coevolution with MHC Class I genes.

The gene family of killer cell immunoglobulin-like receptors (KIRs) in primates provides the first line of defense against virus infection and tumor transformation. Interacting with MHC class I molecules, KIRs can regulate the cytotoxic activity of natural killer (NK) cells and distinguish the tumor and virus infected cells from normal body cells. Phylogenetic analysis and comparison of domain structures identified three major groups of KIR genes (group I, II, and III genes). These groups of KIR genes, generated by a series of gene duplications, have acquired different MHC-binding specificity. Inference of ancestral KIR sequences suggested that the functional divergence of group I genes from group II genes occurred by positive selection at the MHC-binding sites after duplication. Our evolutionary study has shown that group I genes diverged from group II genes about 17 million years ago (Mya) apparently after separation of hominoids from Old World (OW) monkeys. Around the same time, gene duplication generating the class I MHC-C locus appears to have occurred. These findings suggest that KIR and MHC class I genes have coevolved as an interacting system. The KIR gene family has experienced a rapid expansion in primate species. The rate of expansion of this gene family seems to be one of the highest among all hominoid gene families. The KIR gene family is also subject to birth-and-death evolution.

Amino Acid Sequence↗

Eighty percent of proteins are different between humans and chimpanzees.

The chimpanzee is our closest living relative. The morphological differences between the two species are so large that there is no problem in distinguishing between them. However, the nucleotide difference between the two species is surprisingly small. The early genome comparison by DNA hybridization techniques suggested a nucleotide difference of 1-2%. Recently, direct nucleotide sequencing confirmed this estimate. These findings generated the common belief that the human is extremely close to the chimpanzee at the genetic level. However, if one looks at proteins, which are mainly responsible for phenotypic differences, the picture is quite different, and about 80% of proteins are different between the two species. Still, the number of proteins responsible for the phenotypic differences may be smaller since not all genes are directly responsible for phenotypic characters.

Animals↗

Origin and evolution of the Ig-like domains present in mammalian leukocyte receptors: insights from chicken, frog, and fish homologues.

In mammals many natural killer (NK) cell receptors, encoded by the leukocyte receptor complex (LRC), regulate the cytotoxic activity of NK cells and provide protection against virus-infected and tumor cells. To investigate the origin of the Ig-like domains encoded by the LRC genes, a subset of C2-type Ig-like domain sequences was compiled from mammals, birds, amphibians, and fish. Phylogenetic analysis of these sequences generated seven monophyletic groups in mammals (MI, MII, and FcI, FcIIa, FcIIb, FcIII, FcIV), two in chicken (CI, CII), four in frog (FI-FIV), and five in zebrafish (ZI-ZV). The analysis of the major groups supported the following order of divergence: ZI [or a common ancestor of ZI and F (a cluster composed of the FcIII and FIII groups)], F, CII (or a common ancestor of CII and MII), MII, and MI-CI. The relationships of the remaining groups were unclear, since the phylogenetic positions of these groups were not supported by high bootstrap values. Two main conclusions can be drawn from this analysis. First, the two groups of mammalian LRC sequences must diverged before the separation of the avian and mammalian lineages. Second, the mammalian LRC sequences are most closely related to the Fc receptor sequences and these two groups diverged before the separation of birds and mammals.

Animals↗

A simple method for predicting the functional differentiation of duplicate genes and its application to MIKC-type MADS-box genes.

A simple statistical method for predicting the functional differentiation of duplicate genes was developed. This method is based on the premise that the extent of functional differentiation between duplicate genes is reflected in the difference in evolutionary rate because the functional change of genes is often caused by relaxation or intensification of functional constraints. With this idea in mind, we developed a window analysis of protein sequences to identify the protein regions in which the significant rate difference exists. We applied this method to MIKC-type MADS-box proteins that control flower development in plants. We examined 23 pairs of sequences of floral MADS-box proteins from petunia and found that the rate differences for 14 pairs are significant. The significant rate differences were observed mostly in the K domain, which is important for dimerization between MADS-box proteins. These results indicate that our statistical method may be useful for predicting protein regions that are likely to be functionally differentiated. These regions may be chosen for further experimental studies.

Data Interpretation, Statistical↗

Comparative evolutionary analysis of olfactory receptor gene clusters between humans and mice.

Olfactory receptor (OR) genes form the largest multigene family in mammalian genomes. Humans have approximately 800 OR genes, but >50% of them are pseudogenes. By contrast, mice have approximately 1400 OR genes and pseudogenes are approximately 25%. To understand the evolutionary processes that shaped the difference of OR gene families between humans and mice, we studied the genomic locations of all human and mouse OR genes and conducted a detailed phylogenetic analysis using functional genes and pseudogenes. We identified 40 phylogenetic clades with high bootstrap supports, most of which contain both human and mouse genes. Interestingly, a particular clade contains approximately 100 pseudogenes in humans, whereas the numbers of pseudogenes are <20 for most of the mouse clades. We also found that the organization of OR genomic clusters is well conserved between humans and mice in many chromosomal locations. Despite the difference in the numbers of genes, the numbers of large genomic clusters are nearly the same for humans and mice. These observations suggest that the greater OR gene repertoire in mice has been generated mainly by tandem gene duplication within each genomic cluster.

Animals↗

Evolutionary changes of the number of olfactory receptor genes in the human and mouse lineages.

The numbers of functional olfactory receptor (OR) genes are quite variable among mammalian species. Previously we have reported that humans have 388 functional OR genes and 414 pseudogenes, while mice have 1037 functional genes and 354 pseudogenes. These observations suggest either that humans lost many functional OR genes after the human-mouse divergence (HMD) or that mice gained many functional genes. To distinguish between these two hypotheses, we devised a new method of inferring the number of functional OR genes in the most recent common ancestor (MRCA) of humans and mice. An application of this method suggested that the MRCA had approximately 750 functional OR genes and that mice acquired approximately 350 new OR genes after the HMD whereas approximately 430 OR genes in the MRCA have become pseudogenes or eliminated in the human lineage. Therefore, the two evolutionary hypotheses mentioned above are not mutually exclusive and both are nearly equally responsible for the difference in the number of OR genes between humans and mice.

Animals↗

Genomic organization and evolutionary analysis of Ly49 genes encoding the rodent natural killer cell receptors: rapid evolution by repeated gene duplication.

Ly49 genes regulate the cytotoxic activity of natural killer (NK) cells in rodents and provide important protection against virus-infected or tumor cells. About 15 Ly49 genes have been identified in mice, but only a few genes have been reported to date in rats. Here we studied all Ly49 genes in the entire rat genome sequence and identified 17 putative functional and 16 putative non-functional genes together with their genomic locations in a 1.8-Mb region of chromosome 4. Phylogenetic analysis of these genes indicated that the Ly49 gene family expanded rapidly in recent years, and this expansion was mediated by both tandem and genomic block duplication. The joint phylogenetic analysis of mouse and rat genes suggested that the most recent common ancestor of the two species had at least several Ly49 genes, but that the majority of current duplicate genes were generated after divergence of the two species. In both species Ly49 genes are apparently subject to birth-and-death evolution, but the birth and death rates of Ly49 genes are higher in rats than in mice. The rate of gene expansion in the Ly49 gene family in rats is one of the highest among all mammalian multigene families so far studied. The biochemical function of Ly49 genes is essentially the same as that of KIR genes in primates, but the molecular structures of the two groups of NK cell receptors are very different. A hypothesis was presented to explain the origin of the differential use of Ly49 and KIR genes in rodents and primates.

Animals↗

Prospects for inferring very large phylogenies by using the neighbor-joining method.

Current efforts to reconstruct the tree of life and histories of multigene families demand the inference of phylogenies consisting of thousands of gene sequences. However, for such large data sets even a moderate exploration of the tree space needed to identify the optimal tree is virtually impossible. For these cases the neighbor-joining (NJ) method is frequently used because of its demonstrated accuracy for smaller data sets and its computational speed. As data sets grow, however, the fraction of the tree space examined by the NJ algorithm becomes minuscule. Here, we report the results of our computer simulation for examining the accuracy of NJ trees for inferring very large phylogenies. First we present a likelihood method for the simultaneous estimation of all pairwise distances by using biologically realistic models of nucleotide substitution. Use of this method corrects up to 60% of NJ tree errors. Our simulation results show that the accuracy of NJ trees decline only by approximately 5% when the number of sequences used increases from 32 to 4,096 (128 times) even in the presence of extensive variation in the evolutionary rate among lineages or significant biases in the nucleotide composition and transition/transversion ratio. Our results encourage the use of complex models of nucleotide substitution for estimating evolutionary distances and hint at bright prospects for the application of the NJ and related methods in inferring large phylogenies.

Computer Simulation↗

False-positive selection identified by ML-based methods: examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus.

Sexually induced gene 1 (Sig1) in the centric diatom Thalassiosira weissflogii is considered to encode a gamete recognition protein. Sorhannus (2003) analyzed nucleotide sequences of Sig1 using parsimony analysis and the maximum-likelihood (ML)-based Bayesian method for inferring positive selection at single amino acid sites and reported that positively selected sites were detected by the latter method but not by the former. He then concluded that for this type of study, the ML-based method is more reliable than parsimony analysis. Here we show that his results apparently represent false-positive cases of the ML-based method and that there is no solid evidence that this gene contains positively selected sites. We further demonstrate that in the tax gene of human T-cell lymphotropic virus type I (HTLV-I), all codon sites, including invariable sites, can be inferred as positively selected sites by the ML-based method. These observations indicate that the ML-based method may produce many false-positive sites. One of the main reasons for the occurrence of false positives is that in the ML-based method, codon sites are grouped into several categories, with different nonsynonymous/synonymous rate ratios (omegas), on a purely statistical basis, and positive selection is inferred indirectly by examining whether the average omega for each category is greater than 1. In parsimony analysis, however, the evolutionary change of nucleotides at each codon site is examined. For this reason, parsimony-based methods rarely produce false positives and are safer than ML-based methods for detecting positive selection at individual codon sites, although a large number of sequences are necessary.

Bayes Theorem↗

Type I MADS-box genes have experienced faster birth-and-death evolution than type II MADS-box genes in angiosperms.

Plant MADS-box genes form a large gene family for transcription factors and are involved in various aspects of developmental processes, including flower development. They are known to be subject to birth-and-death evolution, but the detailed features of this mode of evolution remain unclear. To have a deeper insight into the evolutionary pattern of this gene family, we enumerated all available functional and nonfunctional (pseudogene) MADS-box genes from the Arabidopsis and rice genomes. Plant MADS-box genes can be classified into types I and II genes on the basis of phylogenetic analysis. Conducting extensive homology search and phylogenetic analysis, we found 64 presumed functional and 37 nonfunctional type I genes and 43 presumed functional and 4 nonfunctional type II genes in Arabidopsis. We also found 24 presumed functional and 6 nonfunctional type I genes and 47 presumed functional and 1 nonfunctional type II genes in rice. Our phylogenetic analysis indicated there were at least about four to eight type I genes and approximately 15-20 type II genes in the most recent common ancestor of Arabidopsis and rice. It has also been suggested that type I genes have experienced a higher rate of birth-and-death evolution than type II genes in angiosperms. Furthermore, the higher rate of birth-and-death evolution in type I genes appeared partly due to a higher frequency of segmental gene duplication and weaker purifying selection in type I than in type II genes.

Arabidopsis↗

MEGA3: Integrated software for Molecular Evolutionary Genetics Analysis and sequence alignment.

With its theoretical basis firmly established in molecular evolutionary and population genetics, the comparative DNA and protein sequence analysis plays a central role in reconstructing the evolutionary histories of species and multigene families, estimating rates of molecular evolution, and inferring the nature and extent of selective forces shaping the evolution of genes and genomes. The scope of these investigations has now expanded greatly owing to the development of high-throughput sequencing techniques and novel statistical and computational methods. These methods require easy-to-use computer programs. One such effort has been to produce Molecular Evolutionary Genetics Analysis (MEGA) software, with its focus on facilitating the exploration and analysis of the DNA and protein sequence variation from an evolutionary perspective. Currently in its third major release, MEGA3 contains facilities for automatic and manual sequence alignment, web-based mining of databases, inference of the phylogenetic trees, estimation of evolutionary distances and testing evolutionary hypotheses. This paper provides an overview of the statistical methods, computational tools, and visual exploration modules for data input and the results obtainable in MEGA.

Databases, Genetic↗

Concerted and nonconcerted evolution of the Hsp70 gene superfamily in two sibling species of nematodes.

We have identified the Hsp70 gene superfamily of the nematode Caenorhabditis briggsae and investigated the evolution of these genes in comparison with Hsp70 genes from C. elegans, Drosophila, and yeast. The Hsp70 genes are classified into three monophyletic groups according to their subcellular localization, namely, cytoplasm (CYT), endoplasmic reticulum (ER), and mitochondria (MT). The Hsp110 genes can be classified into the polyphyletic CYT group and the monophyletic ER group. The different Hsp70 and Hsp110 groups appeared to evolve following the model of divergent evolution. This model can also explain the evolution of the ER and MT genes. On the other hand, the CYT genes are divided into heat-inducible and constitutively expressed genes. The constitutively expressed genes have evolved more or less following the birth-and-death process, and the rates of gene birth and gene death are different between the two nematode species. By contrast, some heat-inducible genes show an intraspecies phylogenetic clustering. This suggests that they are subject to sequence homogenization resulting from gene conversion-like events. In addition, the heat-inducible genes show high levels of sequence conservation in both intra-species and inter-species comparisons, and in most cases, amino acid sequence similarity is higher than nucleotide sequence similarity. This indicates that purifying selection also plays an important role in maintaining high sequence similarity among paralogous Hsp70 genes. Therefore, we suggest that the CYT heat-inducible genes have been subjected to a combination of purifying selection, birth-and-death process, and gene conversion-like events.

Animals↗

Evolution of olfactory receptor genes in the human genome.

Olfactory receptor (OR) genes form the largest known multigene family in the human genome. To obtain some insight into their evolutionary history, we have identified the complete set of OR genes and their chromosomal locations from the latest human genome sequences. We detected 388 potentially functional genes that have intact ORFs and 414 apparent pseudogenes. The number and the fraction (48%) of functional genes are considerably larger than the ones previously reported. The human OR genes can clearly be divided into class I and class II genes, as was previously noted. Our phylogenetic analysis has shown that the class II OR genes can further be classified into 19 phylogenetic clades supported by high bootstrap values. We have also found that there are many tandem arrays of OR genes that are phylogenetically closely related. These genes appear to have been generated by tandem gene duplication. However, the relationships between genomic clusters and phylogenetic clades are very complicated. There are a substantial number of cases in which the genes in the same phylogenetic clade are located on different chromosomal regions. In addition, OR genes belonging to distantly related phylogenetic clades are sometimes located very closely in a chromosomal region and form a tight genomic cluster. These observations can be explained by the assumption that several chromosomal rearrangements have occurred at the regions of OR gene clusters and the OR genes contained in different genomic clusters are shuffled.

Biological Evolution↗

Antiquity and evolution of the MADS-box gene family controlling flower development in plants.

MADS-box genes in plants control various aspects of development and reproductive processes including flower formation. To obtain some insight into the roles of these genes in morphological evolution, we investigated the origin and diversification of floral MADS-box genes by conducting molecular evolutionary genetics analyses. Our results suggest that the most recent common ancestor of today's floral MADS-box genes evolved roughly 650 MYA, much earlier than the Cambrian explosion. They also suggest that the functional classes T (SVP), B (and Bs), C, F (AGL20 or TM3), A, and G (AGL6) of floral MADS-box genes diverged sequentially in this order from the class E gene lineage. The divergence between the class G and E genes apparently occurred around the time of the angiosperm/gymnosperm split. Furthermore, the ancestors of three classes of genes (class T genes, class B/Bs genes, and the common ancestor of the other classes of genes) might have existed at the time of the Cambrian explosion. We also conducted a phylogenetic analysis of MADS-domain sequences from various species of plants and animals and presented a hypothetical scenario of the evolution of MADS-box genes in plants and animals, taking into account paleontological information. Our study supports the idea that there are two main evolutionary lineages (type I and type II) of MADS-box genes in plants and animals.

Animals↗

Birth-and-death evolution in primate MHC class I genes: divergence time estimates.

The major histocompatibility complex (MHC) is a multigene family that mediates the host immune response by helping T lymphocytes to recognize and respond to foreign antigens. The high degree of polymorphism and a quick turnover of the genetic loci make the evolution of MHC genes an intriguing subject of study. To understand the evolutionary pattern of this multigene family, we studied the phylogeny and divergence times of six functional MHC class I loci from primate species. On the phylogenetic trees, locus F occupies the most basal position among these loci. Our results suggest that the F locus diverged from the other MHC class I loci about 46-66 MYA. The major diversification of the other class I loci was estimated to have occurred at about 35-49 MYA, which is before the time of separation of Old World-New World monkeys. The gene duplication leading to the classical C locus in great apes appears to have occurred about 21-28 MYA. At approximately the same time the duplication of the B locus occurred in macaques. The oldest allelic lineages of A, B, and C loci in humans seem to have appeared at least 14-19, 10-15, and 13-17 MYA, respectively. Our phylogenetic analysis supports the hypothesis that the nonclassical locus F has diverged from the rest of class I loci very early in primate evolution. The overall phylogenetic pattern observed among class I genes is consistent with the model of birth-and-death evolution.

Animals↗

Reanalysis of Murphy et al.'s data gives various mammalian phylogenies and suggests overcredibility of Bayesian trees.

Murphy and colleagues reported that the mammalian phylogeny was resolved by Bayesian phylogenetics. However, the DNA sequences they used had many alignment gaps and undetermined nucleotide sites. We therefore reanalyzed their data by minimizing unshared nucleotide sites and retaining as many species as possible (13 species). In constructing phylogenetic trees, we used the Bayesian, maximum likelihood (ML), maximum parsimony (MP), and neighbor-joining (NJ) methods with different substitution models. These trees were constructed by using both protein and DNA sequences. The results showed that the posterior probabilities for Bayesian trees were generally much higher than the bootstrap values for ML, MP, and NJ trees. Two different Bayesian topologies for the same set of species were sometimes supported by high posterior probabilities, implying that two different topologies can be judged to be correct by Bayesian phylogenetics. This suggests that the posterior probability in Bayesian analysis can be excessively high as an indication of statistical confidence and therefore Murphy et al.'s tree, which largely depends on Bayesian posterior probability, may not be correct.

Animals↗

Estimation of divergence times for major lineages of primate species.

Although the phylogenetic relationships of major lineages of primate species are relatively well established, the times of divergence of these lineages as estimated by molecular data are still controversial. This controversy has been generated in part because different authors have used different types of molecular data, different statistical methods, and different calibration points. We have therefore examined the effects of these factors on the estimates of divergence times and reached the following conclusions: (1) It is advisable to concatenate many gene sequences and use a multigene gamma distance for estimating divergence times rather than using the individual gene approach. (2) When sequence data from many nuclear genes are available, protein sequences appear to give more robust estimates than DNA sequences. (3) Nuclear proteins are generally more suitable than mitochondrial proteins for time estimation. (4) It is important first to construct a phylogenetic tree for a group of species using some outgroups and then estimate the branch lengths. (5) It appears to be better to use a few reliable calibration points rather than many unreliable ones. Considering all these factors and using two calibration points, we estimated that the human lineage diverged from the chimpanzee, gorilla, orangutan, Old World monkey, and New World monkey lineages approximately 6 MYA (with a range of 5-7), 7 MYA (range, 6-8), 13 MYA (range, 12-15), 23 MYA (range, 21-25), and 33 MYA (range 32-36).

Animals↗

Overcredibility of molecular phylogenies obtained by Bayesian phylogenetics.

Bayesian phylogenetics has recently been proposed as a powerful method for inferring molecular phylogenies, and it has been reported that the mammalian and some plant phylogenies were resolved by using this method. The statistical confidence of interior branches as judged by posterior probabilities in Bayesian analysis is generally higher than that as judged by bootstrap probabilities in maximum likelihood analysis, and this difference has been interpreted as an indication that bootstrap support may be too conservative. However, it is possible that the posterior probabilities are too high or too liberal instead. Here, we show by computer simulation that posterior probabilities in Bayesian analysis can be excessively liberal when concatenated gene sequences are used, whereas bootstrap probabilities in neighbor-joining and maximum likelihood analyses are generally slightly conservative. These results indicate that bootstrap probabilities are more suitable for assessing the reliability of phylogenetic trees than posterior probabilities and that the mammalian and plant phylogenies may not have been fully resolved.

Amino Acid Substitution↗