Search PubMed⌕ Search

Biomedical subjects

W H Li

Publications and source records attributed to W H Li.

At least 19 recordsLinked to original sources

Evolutionary analyses of the human genome.

The completion of the human genome will greatly accelerate the development of a new branch of science--evolutionary genomics. We can now directly address important questions about the evolutionary history of human genes and their regulatory sequences. Computational analyses of the human genome will reveal the number of genes and repetitive elements, the extent of gene duplication and compositional heterogeneity in the human genome, and the extent of domain shuffling and domain sharing among proteins. Here we present some first glimpses of these features.

Conserved Sequence↗

Genomic divergences between humans and other hominoids and the effective population size of the common ancestor of humans and chimpanzees.

To study the genomic divergences among hominoids and to estimate the effective population size of the common ancestor of humans and chimpanzees, we selected 53 autosomal intergenic nonrepetitive DNA segments from the human genome and sequenced them in a human, a chimpanzee, a gorilla, and an orangutan. The average sequence divergence was only 1.24% +/- 0.07% for the human-chimpanzee pair, 1.62% +/- 0.08% for the human-gorilla pair, and 1.63% +/- 0.08% for the chimpanzee-gorilla pair. These estimates, which were confirmed by additional data from GenBank, are substantially lower than previous ones, which included repetitive sequences and might have been based on less-accurate sequence data. The average sequence divergences between orangutans and humans, chimpanzees, and gorillas were 3.08% +/- 0.11%, 3.12% +/- 0.11%, and 3.09% +/- 0.11%, respectively, which also are substantially lower than previous estimates. The sequence divergences in other regions between hominoids were estimated from extensive data in GenBank and the literature, and Alus showed the highest divergence, followed in order by Y-linked noncoding regions, pseudogenes, autosomal intergenic regions, X-linked noncoding regions, synonymous sites, introns, and nonsynonymous sites. The neighbor-joining tree derived from the concatenated sequence of the 53 segments--24,234 bp in length--supports the Homo-Pan clade with a 100% bootstrap value. However, when each segment is analyzed separately, 22 of the 53 segments (approximately 42%) give a tree that is incongruent with the species tree, suggesting a large effective population size (N(e)) of the common ancestor of Homo and Pan. Indeed, a parsimony analysis of the 53 segments and 37 protein-coding genes leads to an estimate of N(e) = 52,000 to 96,000. As this estimate is 5 to 9 times larger than the long-term effective population size of humans (approximately 10,000) estimated from various genetic polymorphism data, the human lineage apparently had experienced a large reduction in effective population size after its separation from the chimpanzee lineage. Our analysis assumes a molecular clock, which is in fact supported by the sequence data used. Taking the orangutan speciation date as 12 to 16 million years ago, we obtain an estimate of 4.6 to 6.2 million years for the Homo-Pan divergence and an estimate of 6.2 to 8.4 million years for the gorilla speciation date, suggesting that the gorilla lineage branched off 1.6 to 2.2 million years earlier than did the human-chimpanzee divergence.

Alu Elements↗

Transposable elements are found in a large number of human protein-coding genes.

To study the genome-wide impact of transposable elements (TEs) on the evolution of protein-coding regions, we examined 13 799 human genes and found 533 (approximately 4%) cases of TEs within protein-coding regions. The majority of these TEs (approximately 89.5%) reside within 'introns' and were recruited into coding regions as novel exons. We found that TE integration often has an effect on gene function. In particular, there were two mouse genes whose coding regions consist largely of TEs, suggesting that TE insertion might create new genes. Thus, there is increasing evidence for an important role of TEs in gene evolution. Because many TEs are taxon-specific, their integration into coding regions could accelerate species divergence.

Animals↗

Human DNA sequence variation in a 6.6-kb region containing the melanocortin 1 receptor promoter.

An approximately 6.6-kb region located upstream from the melanocortin 1 receptor (MC1R) gene and containing its promoter was sequenced in 54 humans (18 Africans, 18 Asians, and 18 Europeans) and in one chimpanzee, gorilla, and orangutan. Seventy-six polymorphic sites were found among the human sequences and the average nucleotide diversity (pi) was 0.141%, one of the highest among all studies of nuclear sequence variation in humans. Opposite to the pattern observed in the MC1R coding region, in the present region pi is highest in Africans (0.136%) compared to Asians (0.116%) and Europeans (0.122%). The distributions of pi, theta, and Fu and Li's F-statistic are nonuniform along the sequence and among continents. The pattern of genetic variation is consistent with a population expansion in Africans. We also suggest a possible phase of population size reduction in non-Africans and purifying selection acting in the middle subregion and parts of the 5' subregion in Africans. We hypothesize diversifying selection acting on some sites in the 5' and 3' subregions or in the MC1R coding region in Asians and Europeans, though we cannot reject the possibility of relaxation of functional constraints in the MC1R gene in Asians and Europeans. The mutation rate in the sequenced region is 1.65 x 10(-9) per site per year. The age of the most recent common ancestor for this region is similar to that for the other long noncoding regions studied to date, providing evidence for ancient gene genealogies. Our population screening and phylogenetic footprinting suggest potentially important sites for the MC1R promoter function.

Animals↗

Bushbaby growth hormone is much more similar to nonprimate growth hormones than to rhesus monkey and human growth hormones.

Unlike other mammals, Old World primates have five growth hormone-like genes that are highly divergent at the amino acid level from the single growth hormone genes found in nonprimates. Additionally, there is a change in the interaction of growth hormone with its receptor in humans such that human growth hormone functions in nonprimates, whereas nonprimate growth hormone is ineffective in humans. A Southern blotting analysis of the genome of a prosimian, Galago senegalensis, revealed a single growth hormone locus. This single gene was PCR-amplified from genomic DNA and sequenced. It has a rate of nonsynonymous nucleotide substitution less than one fourth that of the human growth hormone gene, while the rates of synonymous substitution in the two species are less different. Human and rhesus monkey growth hormones exhibit variation at a number of amino acid residues that can affect receptor binding. The galago growth hormone is conservative at each of these sites, indicating that this growth hormone is functionally like nonprimate growth hormones. These observations indicate that the amplification and rapid divergence of primate growth hormones occurred after the separation of the higher primate lineage from the galago lineage.

Amino Acid Sequence↗

Isolation of Cladonema Pax-B genes and studies of the DNA-binding properties of cnidarian Pax paired domains.

Pax genes encode nuclear transcription factors that are involved in developmental control. They contain a conserved DNA-binding domain, the paired domain. The DNA-binding specificity of paired domains is directly related to the gene regulation function of Pax proteins. Pax genes were previously divided into five groups on the basis of a phylogenetic analysis of paired domains. In this study, two highly similar cnidarian Pax-B genes from Cladonema californicum, a jellyfish with eyes, were found and sequenced. In an effort to understand the function of the cnidarian Pax genes isolated in this and a previous study, we characterized the consensus DNA sequences bound by the cnidarian paired domains using a PCR-based method and electrophoretic mobility shift assays. The consensus DNA sequences obtained are very similar to those bound by mammalian Pax proteins. Comparison of known consensus sequences indicates that they are all partially palindromic, but this characteristic is most prominent in cnidarians, which suggests that the DNA sequences bound by the ancestral paired domain could have been palindromic. Also, cnidarian paired domains, like those of Pax-2/5/8, possess a broader binding specificity than other paired domains, which implies that the common ancestor of Pax-2/5/8 and Pax-4/6 paired domains could also have had a similar broad DNA-binding specificity. Thus far, a definitive Pax-6 gene has not been found in several cnidarian species examined, which is consistent with a later origin of the Pax-6 gene and raises two possibilities: the Pax genes of cnidarians are multifunctional and control two or more developmental pathways, including eye development, or they use a Pax-independent pathway for eye development. Whether this pathway does exist and is unique to cnidarians or it whether it represents a true master control under which Pax-6 was later included remains to be determined.

Animals↗

NJML+: an extension of the NJML method to handle protein sequence data and computer software implementation.

While the maximum-likelihood (ML) method of tree reconstruction is statistically rigorous, it is extremely time-consuming for reconstructing large trees. We previously developed a hybrid method (NJML) that combines the neighbor-joining (NJ) and ML methods and thus is much faster than the ML method and improves the performance of NJ. However, we considered only nucleotide sequence data, so NJML is not suitable for handling amino acid sequence data, which requires even more computer time. NJML+ is an implementation of a further improved method for practical data analyses (including protein sequence data). Our extensive simulations using nucleotide and amino acid sequences showed that NJML+ gave good results in tree reconstruction. Indeed, NJML+ showed substantial improvements over existing methods in terms of both computational times and efficiencies, especially for amino acid sequence data. We also developed a "user-friendly" interface for the NJML+ program, including a simple tree viewer.

Amino Acid Sequence↗

Global patterns of human DNA sequence variation in a 10-kb region on chromosome 1.

Human DNA variation is currently a subject of intense research because of its importance for studying human origins, evolution, and demographic history and for association studies of complex diseases. A approximately 10-kb region on chromosome 1, which contains only four small exons (each <155 bp), was sequenced for 61 humans (20 Africans, 20 Asians, and 21 Europeans) and for 1 chimpanzee, 1 gorilla, and 1 orangutan. We found 52 polymorphic sites among the 122 human sequences and 382 variant sites among the human, chimpanzee, gorilla, and orangutan sequences. For the introns sequenced (8,991 bp), the nucleotide diversity (pi) was 0.058% among all sequences, 0.076% among the African sequences, 0.047% among the Asian sequences, and 0.045% among the European sequences. A compilation of data revealed that autosomal regions have, on average, the highest pi value (0.091%), X-linked regions have a somewhat lower pi value (0.079%), and Y-linked regions have a very low pi value (0.008%). The lower polymorphism in the present region may be due to a lower mutation rate and/or selection in the gene containing these introns or in genes linked to this region. The present region and two other 10-kb noncoding regions all show a strong excess of low-frequency variants, indicating a relatively recent population expansion. This region has a low mutation rate, which was estimated to be 0.74 x 10 per nucleotide per year. An average estimate of approximately 12,600 for the long-term effective population size was obtained using various methods; the estimate was not far from the commonly used value of 10,000. Fu and Li's tests rejected the assumption of an equilibrium neutral Wright-Fisher population, largely owing to the high proportion of low-frequency variants. The age of the most recent common ancestor of the sequences in our sample was estimated to be more than 1 Myr. Allowing for some unrealistic assumptions in the model, this estimate would still suggest an age of more than 500,000 years, providing further evidence for a genetic history of humans much more ancient than the emergence of modern humans. The fact that many unique variants exist in Europe and Asia also suggests a fairly long genetic history outside of Africa and argues against a complete replacement of all indigenous populations in Europe and Asia by a small Africa stock. Moreover, the ancient genetic history of humans indicates no severe bottleneck during the evolution of humans in the last half million years; otherwise, much of the ancient genetic history would have been lost during a severe bottleneck. We suggest that both the "Out of Africa" and the multiregional models are too simple to explain the evolution of modern humans.

Africa↗

Episodic evolution of growth hormone in primates and emergence of the species specificity of human growth hormone receptor.

Growth hormone (GH) evolution is very conservative among mammals, except for primates and ruminant artiodactyls. In fact, most known mammalian GH sequences differ from the inferred ancestral mammalian sequence by only a few amino acids. In contrast, the human GH sequence differs from the inferred ancestral sequence by 59 amino acids. However, it is not known when this rapid evolution of GH occurred during primate evolution or whether it was due to positive selection. Also, human growth hormone receptor (GHR) displays species specificity; i.e., it can interact only with human (or rhesus monkey) GH, not with nonprimate GHS: The species specificity of human GHR is largely due to the Leu-->Arg change at position 43, and it has been hypothesized that this change must have been preceded by the His-->Asp change at position 171 of GH. Is this hypothesis true? And when did these changes occur? To address the above issues, we sequenced GH and GHR genes in prosimians and simians. Our data supported the above hypothesis and revealed that the species specificity of human GHR actually emerged in the common ancestor of Old World primates, but the transitional phase still persists in New World monkeys. Our data showed that the rapid evolution of primate GH occurred during a relatively short period (in the common ancestor of higher primates) and that the rate of change was especially high at functionally important sites, suggesting positive selection. However, the nonsynonymous rate/synonymous rate ratio at these sites was <1, so relaxation of purifying selection might have played a role in the rapid evolution of the GH gene in simians, possibly as a result of multiple gene duplications. Similar to GH, GHR displayed an accelerated rate of evolution in primates. Our data revealed proportionally more amino acid replacements at the functionally important sites in both GH and GHR in simians but, surprisingly, showed few coincidental replacements of amino acids forming the same intermolecular contacts between the two proteins.

Amino Acids↗

Densities, length proportions, and other distributional features of repetitive sequences in the human genome estimated from 430 megabases of genomic sequence.

The densities of repetitive elements in the human genome were calculated in each GC content class using non-overlapping windows of 50kb. The density of Alu is two to three times higher in GC-rich regions than in AT-rich regions, while the opposite is true for LINE1. In contrast, LINE2 and other elements, such as DNA transposons, are more uniformly distributed in the genome. The number of Alus in the human genome was estimated to be 1.4 million, higher than previous estimates. About 40% of the autosomes and approximately 51% of the X and Y chromosomes are occupied by repetitive elements. In total, the human genome is estimated to contain more than 4 million repetitive elements. The GC contents (%) of repetitive elements and their flanking regions were also calculated. The GC contents of almost all kinds of repeats are positively correlated with the window GC contents, suggesting that a repetitive sequence is subject to the same mutation pressure as its surrounding regions, so it tends to have the same GC content as its surrounding regions. This observation supports the regional mutation hypothesis. The only two exceptions are AluYa and AluYb8, the two youngest Alu subfamilies. The GC content of AluYb8 is negatively correlated with that of its surrounding regions, while AluYa shows no correlation, suggesting different insertion patterns for these two young Alu subfamilies. This suggestion was supported by the fact that the average genetic distance between members of AluYb8 in each GC window class is positively correlated with the GC content of the window, but no correlation was found for AluYa. AluYa is more frequent in Y chromosome than in other chromosomes; the same is true for LTR retroviruses. This pattern might be correlated with the evolutionary history of Y chromosome.

Alu Elements↗

Worldwide DNA sequence variation in a 10-kilobase noncoding region on human chromosome 22.

Human DNA sequence variation data are useful for studying the origin, evolution, and demographic history of modern humans and the mechanisms of maintenance of genetic variability in human populations, and for detecting linkage association of disease. Here, we report worldwide variation data from a approximately 10-kilobase noncoding autosomal region. We identified 75 variant sites in 64 humans (128 sequences) and 463 variant sites among the human, chimpanzee, and orangutan sequences. Statistical tests suggested that the region is selectively neutral. The average nucleotide diversity (pi) across the region was 0.088% among all of the human sequences obtained, 0.085% among African sequences, and 0.082% among non-African sequences, supporting the view of a low nucleotide diversity ( approximately 0.1%) in humans. The comparable pi value in non-Africans to that in Africans indicates no severe bottleneck during the evolution of modern non-Africans; however, the possibility of a mild bottleneck cannot be excluded because non-Africans showed considerably fewer variants than Africans. The present and two previous large data sets all show a strong excess of low frequency variants in comparison to that expected from an equilibrium population, indicating a relatively recent population expansion. The mutation rate was estimated to be 1.15 x 10(-9) per nucleotide per year. Estimates of the long-term effective population size N(e) by various statistical methods were similar to those in other studies. The age of the most recent common ancestor was estimated to be approximately 1.29 million years ago among all of the sequences obtained and approximately 634,000 years ago among the non-African sequences, providing the first evidence from a noncoding autosomal region for ancient human histories, even among non-Africans.

Animals↗

Cellular regulation of cytosolic group IV phospholipase A2 by phosphatidylinositol bisphosphate levels.

Cytosolic group IV phospholipase A2 (cPLA2) is a ubiquitously expressed enzyme with key roles in intracellular signaling. The current paradigm for activation of cPLA2 by stimuli proposes that both an increase in intracellular calcium and mitogen-activated protein kinase-mediated phosphorylation occur together to fully activate the enzyme. Calcium is currently thought to be needed for translocation of the cPLA2 to the membrane via a C2 domain, whereas the role of cPLA2 phosphorylation is less clearly defined. Herein, we report that brief exposure of P388D1 macrophages to UV radiation results in a rapid, cPLA2-mediated arachidonic acid mobilization, without increases in intracellular calcium. Thus, increased Ca2+ availability is a dispensable signal for cPLA2 activation, which suggests the existence of alternative mechanisms for the enzyme to efficiently interact with membranes. Our previous in vitro data suggested the importance of phosphatidylinositol 4,5-bisphosphate (PtdInsP2) in the association of cPLA2 to model membranes and hence in the regulation of cPLA2 activity. Experiments described herein show that PtdInsP2 also serves a similar role in vivo. Moreover, inhibition of PtdInsP2 formation during activation conditions leads to inhibition of the cPLA2-mediated arachidonic acid mobilization. These results suggest that cellular PtdInsP2 levels are involved in the regulation of group IV cPLA2 activation.

Animals↗

Molecular evolution of growth hormone and receptor in the guinea-pig, a mammal unresponsive to growth hormone.

Growth in the guinea-pig is completely unresponsive to endogenous or exogenous growth hormone, despite the fact that the guinea-pig produces normal to high levels of growth hormone and receptor. In primates and artiodactyls, growth hormone exhibits accelerated rates of evolution that appear to be correlated with changes in function. Surprisingly, both guinea-pig growth hormone and receptor exhibit slow rates of evolution similar to those seen in other mammals, implying that both proteins are as functionally conserved in the guinea-pig as in other mammals or that any loss or relaxation of functional constraint was very recent. However, the guinea-pig growth hormone and receptor both exhibit a single amino acid replacement at a site known to have functional significance. Nevertheless, it is unclear whether the aberrant nature of the guinea-pig growth hormone-growth hormone receptor axis is due to these replacements or whether it is due to a defect in post-receptor signalling.

Amino Acid Sequence↗

A rapid heuristic algorithm for finding minimum evolution trees.

The minimum sum of branch lengths (S), or the minimum evolution (ME) principle, has been shown to be a good optimization criterion in phylogenetic inference. Unfortunately, the number of topologies to be analyzed is computationally prohibitive when a large number of taxa are involved. Therefore, simplified, heuristic methods, such as the neighbor-joining (NJ) method, are usually employed instead. The NJ method analyzes only a small number of trees (compared with the size of the entire search space); so, the tree obtained may not be the ME tree (for which the S value is minimum over the entire search space). Different compromises between very restrictive and exhaustive search spaces have been proposed recently. In particular, the "stepwise algorithm" (SA) utilizes what is known in computer science as the "beam search," whereas the NJ method employs a "greedy search." SA is virtually guaranteed to find the ME trees while being much faster than exhaustive search algorithms. In this study we propose an even faster method for finding the ME tree. The new algorithm adjusts its search exhaustiveness (from greedy to complete) according to the statistical reliability of the tree node being reconstructed. It is also virtually guaranteed to find the ME tree. The performances and computational efficiencies of ME, SA, NJ, and our new method were compared in extensive simulation studies. The new algorithm was found to perform practically as well as the SA (and, therefore, ME) methods and slightly better than the NJ method. For searching for the globally optimal ME tree, the new algorithm is significantly faster than existing ones, thus making it relatively practical for obtaining all trees with an S value equal to or smaller than that of the NJ tree, even when a large number of taxa is involved.

Algorithms↗

Molecular systematics of pikas (genus Ochotona) inferred from mitochondrial DNA sequences.

The phylogenetic relationships among worldwide species of genus Ochotona were investigated by sequencing mitochondrial cytochrome b and ND4 genes. Parsimony and neighbor-joining analyses of the sequence data yielded congruent results that strongly indicated three major clusters: the shrub-steppe group, the northern group, and the mountain group. The subgeneric classification of Ochotona species needs to be revised because each of the two subgenera in the present classification contains species from the mountain group. To solve this taxonomic problem so that each taxon is monophyletic, i.e. , represents a natural clade, Ochotona could be divided into three subgenera, one for the shrub-steppe species, a second for the northern species, and a third for the mountain species. The inferred tree suggests that the differentiation of this genus in the Palearctic Region was closely related to the gradual uplifting of the Tibet (Qinghai-Xizang) Plateau, as hypothesized previously, and that vicariance might have played a major role in the differentiation of this genus on the Plateau. On the other hand, the North American species, O. princeps, is most likely a dispersal event, which might have happened during the Pliocene through the opening of the Bering Strait. The phylogenetic relationships within the shrub-steppe group are worth noting in that instead of a monophyletic shrub-dwelling group, shrub dwellers and steppe dwellers are intermingled with each other. Moreover, the sequence divergence within the sister taxa of one steppe dweller and one shrub dweller is very low. These findings support the hypothesis that pikas have entered the steppe environment several times and that morphological similarities within steppe dwellers were due to convergent evolution.

Animals↗

Selective constraints, amino acid composition, and the rate of protein evolution.

What are the major forces governing protein evolution? A common view is that proteins with strong structural and functional requirements evolve more slowly than proteins with weak constraints, because a stringent negative selection pressure limits the number of substitutions. In contrast, Graur claimed that the substitution rate of a protein is mainly determined by its amino acid composition and the changeabilities of amino acids. In this paper, however, we found that the relative changeabilities of amino acids in mammalian proteins are different for transmembranal and nontransmembranal segments, which have very distinct structural requirements. This indicates that the changeability of a given residue is influenced by the structural and functional context. We also reexamined the relationship between substitution rate and amino acid composition. Indeed, the two kinds of segments exhibit contrasting amino acid compositions: transmembranal regions are made up mainly of hydrophobic residues (a total frequency of approximately 60%) and are very poor in polar amino acids (<5%), whereas nontransmembranal segments have frequencies of 30% and 22%, respectively. Interestingly, we found that within a given integral membrane protein, nontransmembranal segments accumulate, on average, twice as many substitutions as transmembranal regions. However, regression analyses showed that the variability in amino acid frequencies among proteins cannot explain more than 30% of the variability in substitution rate for the transmembranal and nontransmembranal data sets. Furthermore, transmembranal and nontransmembranal segments evolving at the same rate in different proteins have different compositions, and the compositions of slowly evolving and rapidly evolving segments of the same type are similar. From these observations, we conclude that the rate of protein evolution is only weakly affected by amino acid composition but is mostly determined by the strength of functional requirements or selective constraints.

Amino Acid Substitution↗

NJML: a hybrid algorithm for the neighbor-joining and maximum-likelihood methods.

In the reconstruction of a large phylogenetic tree, the most difficult part is usually the problem of how to explore the topology space to find the optimal topology. We have developed a "divide-and-conquer" heuristic algorithm in which an initial neighbor-joining (NJ) tree is divided into subtrees at internal branches having bootstrap values higher than a threshold. The topology search is then conducted by using the maximum-likelihood method to reevaluate all branches with a bootstrap value lower than the threshold while keeping the other branches intact. Extensive simulation showed that our simple method, the neighbor-joining maximum-likelihood (NJML) method, is highly efficient in improving NJ trees. Furthermore, the performance of the NJML method is nearly equal to or better than existing time-consuming heuristic maximum-likelihood methods. Our method is suitable for reconstructing relatively large molecular phylogenetic trees (number of taxa >/= 16).

Algorithms↗

Assessment of compositional heterogeneity within and between eukaryotic genomes.

Using large amounts of long genomic sequences, we studied the compositional patterns of eukaryotic genomes. We developed a simple measure, the compositional heterogeneity (or variability) index, to compare the differences in compositional heterogeneity between long genomic sequences. The index measures the average difference in GC content between two adjacent windows normalized by the standard error expected under the assumption of random distribution of nucleotides in a window. We report the following findings: (1) The extent of the compositional heterogeneity in a genomic sequence strongly correlates with its GC content in all multicellular eukaryotes studied regardless of genome size. (2) The human genome appears to be highly compositionally heterogeneous both within and between individual chromosomes; the heterogeneity goes much beyond the predictions of the isochore model. (3) All genomes of multicellular eukaryotes examined in this study are compositionally heterogeneous, although they also contain compositionally uniform segments, or isochores. (4) The true uniqueness of the human (or mammalian) genome is the presence of very high GC regions, which exhibit unusually high compositional heterogeneity and contain few long homogeneous segments (isochores). In general, GC-poor isochores tend to be longer than GC-rich ones. These findings indicate that the genomes of multicellular organisms are much more heterogeneous in nucleotide composition than depicted by the isochore model and so lead to a looser definition of isochores.

Algorithms↗