Search PubMed⌕ Search

Biomedical subjects

R R Hudson

Publications and source records attributed to R R Hudson.

At least 37 records · Page 2Linked to original sources

The coalescent process and background selection.

Some statistical properties of gene trees are described for a model with background deleterious mutations. It is argued that the history of a small sample of genes under this model is a continuous time Markov chain that quickly reaches stationarity. This observation leads to simple expressions for the expected nucleotide diversity and suggests that the frequency spectrum in small samples should be approximately the same as under a strict neutral model. The results concerning expected nucleotide diversity are compared with observed variation on the third chromosome of Drosophila melanogaster.

Animals↗

The hitchhiking effect on the site frequency spectrum of DNA polymorphisms.

The level of DNA sequence variation is reduced in regions of the Drosophila melanogaster genome where the rate of crossing over per physical distance is also reduced. This observation has been interpreted as support for the simple model of genetic hitchhiking, in which directional selection on rare variants, e.g., newly arising advantageous mutants, sweeps linked neutral alleles to fixation, thus eliminating polymorphisms near the selected site. However, the frequency spectra of segregating sites of several loci from some populations exhibiting reduced levels of nucleotide diversity and reduced numbers of segregating sites did not appear different from what would be expected under a neutral equilibrium model. Specifically, a skew toward an excess of rare sites was not observed in these samples, as measured by Tajima's D. Because this skew was predicted by a simple hitchhiking model, yet it had never been expressed quantitatively and compared directly to DNA polymorphism data, this paper investigates the hitchhiking effect on the site frequency spectrum, as measured by Tajima's D and several other statistics, using a computer simulation model based on the coalescent process and recurrent hitchhiking events. The results presented here demonstrate that under the simple hitchhiking model (1) the expected value of Tajima's D is large and negative (indicating a skew toward rare variants), (2) that Tajima's test has reasonable power to detect a skew in the frequency spectrum for parameters comparable to those from actual data sets, and (3) that the Tajima's Ds observed in several data sets are very unlikely to have been the result of simple hitchhiking. Consequently, the simple hitchhiking model is not a sufficient explanation for the DNA polymorphism at those loci exhibiting a decreased number of segregating sites yet not exhibiting a skew in the frequency spectrum.

Animals↗

Deleterious background selection with recombination.

An analytic expression for the expected nucleotide diversity is obtained for a neutral locus in a region with deleterious mutation and recombination. Our analytic results are used to predict levels of variation for the entire third chromosome of Drosophila melanogaster. The predictions are consistent with the low levels of variation that have been observed at loci near the centromeres of the third chromosome of D. melanogaster. However, the low levels of variation observed near the tips of this chromosome are not predicted using currently available estimates of the deleterious mutation rate and of selection coefficients. If considerably smaller selection coefficients are assumed, the low observed levels of variation at the tips of the third chromosome are consistent with the background selection model.

Animals↗

How can the low levels of DNA sequence variation in regions of the drosophila genome with low recombination rates be explained?

Different regions of the Drosophila genome have very different rates of recombination. For example, near centromeres and near the tips of chromosomes, the rates of recombination are much lower than in other regions. Several surveys of polymorphisms in Drosophila have now documented that levels of DNA polymorphism are positively correlated with rates of recombination; i.e., regions with low rates of recombination tend to have low levels of DNA polymorphism within populations of Drosophila. Three hypotheses are reviewed that might account for these observations. The first hypothesis is that regions of low recombination have low neutral mutation rates. Under this hypothesis between-species divergences should also be low in regions of low recombination. In fact, regions of low recombination have diverged at the same rate as other regions of the genome. On this basis, this strictly neutral hypothesis is rejected. The second hypothesis is that the process of fixation of favorable mutations leads to the observed correlation between polymorphism and recombination. This occurs via genetic hitchhiking, in which linked regions of the genome are swept along with the selectively favored mutant as it increases in frequency and eventually fixes in the population. This hitchhiking model with fixation of favorable mutations is compatible with major features of the data. By assuming this model is correct, one can estimate the rate of fixation of favorable mutations. The third hypothesis is that selection against continually arising deleterious mutations results in reduced levels of polymorphism at linked loci. Analysis of this background selection model shows that it can produce some reduction in levels of polymorphism but cannot explain some extreme cases that have been observed. Thus, it appears that hitchhiking of favorable mutations and background selection against deleterious mutations must be considered together to correctly account for the patterns of polymorphism that are observed in Drosophila.

Animals↗

Evidence for positive selection in the superoxide dismutase (Sod) region of Drosophila melanogaster.

DNA sequence variation in a 1410-bp region including the Cu,Zn Sod locus was examined in 41 homozygous lines of Drosophila melanogaster. Fourteen lines were from Barcelona, Spain, 25 were from California populations and the other two were from laboratory stocks. Two common electromorphs, SODS and SODF, are segregating in the populations. Our sample of 41 lines included 19 SodS and 22 SodF alleles (henceforward referred to as Slow and Fast alleles). All 19 Slow alleles were identical in sequence. Of the 22 Fast alleles sequenced, nine were identical in sequence and are referred to as the Fast A haplotypes. The Slow allele sequence differed from the Fast A haplotype at a single nucleotide site, the site that accounts for the amino acid difference between SODS and SODF. There were nine other haplotypes among the remaining 13 Fast alleles sequenced. The overall level of nucleotide diversity (pi) in this sample is not greatly different than that found at other loci in D. melanogaster. It is concluded that the Slow/Fast polymorphism is a recently arisen polymorphism, not an old balanced polymorphism. The large group of nearly identical haplotypes suggests that a recent mutation, at the Sod locus or tightly linked to it, has increased rapidly in frequency to around 50%, both in California and Spain. The application of a new statistical test demonstrates that the occurrence of such large numbers of haplotypes with so little variation among them is very unlikely under the usual equilibrium neutral model. We suggest that the high frequency of some haplotypes is due to natural selection at the Sod locus or at a tightly linked locus.

Animals↗

Hierarchical analysis of linkage disequilibrium in Rhizobium populations: evidence for sex?

Many bacterial species exhibit strong linkage disequilibrium of their chromosomal genes, which apparently indicates restricted recombination between alleles at different loci. The extent to which restricted recombination reflects limited migration between geographically isolated populations versus infrequent mixis of genotypes within populations is more difficult to determine. We examined the genetic structure of Rhizobium leguminosarum biovar phaseoli populations associated with wild and cultivated beans (Phaseolus spp.) over several spatial scales, ranging from individual host plants to throughout the Western Hemisphere. We observed significant linkage disequilibrium at scales at least as small as a cultivated plot. However, the amount of disequilibrium was much greater among isolates collected throughout the Western Hemisphere than among isolates from one area of Mexico, even when disequilibrium was quantified using an index that scales for allelic diversity. This finding suggests that limited migration between populations contributes substantially to linkage disequilibrium in Rhizobium. We also compared the genetic structure for R. leguminosarum bv. phaseoli taken from a cultivated plot with that for Escherichia coli obtained from one human host in an earlier study. Even at this fine scale, linkage disequilibrium in E. coli was very near the theoretical maximum level, whereas it was much less extreme in the local population of Rhizobium. Thus, the genetic structure for R. leguminosarum bv. phaseoli does not exclude the possibility of frequent mixis within local populations.

Enzymes↗

Estimation of levels of gene flow from DNA sequence data.

We compare the utility of two methods for estimating the average levels of gene flow from DNA sequence data. One method is based on estimating FST from frequencies at polymorphic sites, treating each site as a separate locus. The other method is based on computing the minimum number of migration events consistent with the gene tree inferred from their sequences. We compared the performance of these two methods on data that were generated by a computer simulation program that assumed the infinite sites model of mutation and that assumed an island model of migration. We found that in general when there is no recombination, the cladistic method performed better than FST while the reverse was true for rates of recombination similar to those found in eukaryotic nuclear genes, although FST performed better for all recombination rates for very low levels of migration (Nm = 0.1).

Computer Simulation↗

A statistical test for detecting geographic subdivision.

A statistical test for detecting genetic differentiation of subpopulations is described that uses molecular variation in samples of DNA sequences from two or more localities. The statistical significance of the test is determined with Monte Carlo simulations. The power of the test to detect genetic differentiation in a selectively neutral Wright-Fisher island model depends on both sample size and the rates of migration, mutation, and recombination. It is found that the power of the test is substantial with samples of size 50, when 4Nm less than 10, where N is the subpopulation size and m is the fraction of migrants in each subpopulation each generation. More powerful tests are obtained with genes with recombination than with genes without recombination.

Alcohol Dehydrogenase↗

The coalescent process in models with selection, recombination and geographic subdivision.

A population genetic model with a single locus at which balancing selection acts and many linked loci at which neutral mutations can occur is analysed using the coalescent approach. The model incorporates geographic subdivision with migration, as well as mutation, recombination, and genetic drift of neutral variation. It is found that geographic subdivision can affect genetic variation even with high rates of migration, providing that selection is strong enough to maintain different allele frequencies at the selected locus. Published sequence data from the alcohol dehydrogenase locus of Drosophila melanogaster are found to fit the proposed model slightly better than a similar model without subdivision.

Alcohol Dehydrogenase↗

Inferring the evolutionary histories of the Adh and Adh-dup loci in Drosophila melanogaster from patterns of polymorphism and divergence.

The DNA sequences of 11 Drosophila melanogaster lines are compared across three contiguous regions, the Adh and Adh-dup loci and a noncoding 5' flanking region of Adh. Ninety-eight of approximately 4750 sites are segregating in the sample, 36 in the 5' flanking region, 38 in Adh and 24 in Adh-dup. Several methods are presented to test whether the patterns and levels of polymorphism are consistent with neutral molecular evolution. The analysis of within- and between-species polymorphism indicates that the region is evolving in a nonneutral and complex fashion. A graphical analysis of the data provides support for a hypothesized balanced polymorphism at or near position 1490, site of the amino acid replacement difference between Adhf and Adhs. The Adh-dup locus is less polymorphic than Adh and all 24 of its polymorphisms occur at low frequency--suggestive of a recent selective substitution in the Adh-dup region. Adhs alleles form two distinct evolutionary lineages that differ one from another at a total of nineteen sites in the Adh and Adh-dup loci. The polymorphisms are in complete linkage disequilibrium. A recombination experiment failed to find evidence for recombination suppression between the two allelic classes. Two hypotheses are presented to account for the widespread distribution of the two divergent lineages in natural populations. Natural selection appears to have played an important role in governing the overall patterns of nucleotide variation across the two-gene region.

Alcohol Dehydrogenase↗

Pairwise comparisons of mitochondrial DNA sequences in stable and exponentially growing populations.

We consider the distribution of pairwise sequence differences of mitochondrial DNA or of other nonrecombining portions of the genome in a population that has been of constant size and in a population that has been growing in size exponentially for a long time. We show that, in a population of constant size, the sample distribution of pairwise differences will typically deviate substantially from the geometric distribution expected, because the history of coalescent events in a single sample of genes imposes a substantial correlation on pairwise differences. Consequently, a goodness-of-fit test of observed pairwise differences to the geometric distribution, which assumes that each pairwise comparison is independent, is not a valid test of the hypothesis that the genes were sampled from a panmictic population of constant size. In an exponentially growing population in which the product of the current population size and the growth rate is substantially larger than one, our analytical and simulation results show that most coalescent events occur relatively early and in a restricted range of times. Hence, the "gene tree" will be nearly a "star phylogeny" and the distribution of pairwise differences will be nearly a Poisson distribution. In that case, it is possible to estimate r, the population growth rate, if the mutation rate, mu, and current population size, N0, are assumed known. The estimate of r is the solution to ri/mu = ln(N0r) - gamma, where i is the average pairwise difference and gamma approximately 0.577 is Euler's constant.

Animals↗

The rate of Cu,Zn superoxide dismutase evolution.

The rate of amino acid replacement in Cu,Zn SOD greatly departs from the expectations of the molecular clock. We examine 27 Cu,Zn SOD sequences available and conclude that: (1) the SOD enzymes from different mammal families differ from each other by roughly the same number of replacements, which is consistent with a simultaneous mammalian radiation; (2) over the most recent 60 million years (MY) the rate of SOD evolution is fairly high (15 aa/100 aa/100 MYR) and may be considered constant; (3) the rate of accumulation of amino acid replacements since the divergence of fungi, plants and animals to the present is inconstant along different branches of the evolutionary tree; moreover it steadily decreases with time, to the same extent in all lineages; (4) some comparisons exhibit divergences that are in any case greater than expected from a Poisson process on the assumption of a molecular clock; (5) plant chloroplast enzymes display fewer differences from each other than cytoplasmic ones; (6) bacteriocuprein (from Photobacterium leiognathi), fluke and human extracellular SOD are all three extremely remotely related to one another and to the SOD of other eukaryotes. The process of consistent decline of the rate of evolution of Cu, Zn SOD can be described by a number of mathematical functions. We explore simple models that assume constant rates and might be applicable to other proteins or genes that apparently evolve at disparate rates.

Amino Acid Sequence↗

A numerical method for calculating moments of coalescent times in finite populations with selection.

A numerical method is developed for solving a nonstandard singular system of second-order differential equations arising from a problem in population genetics concerning the coalescent process for a sample from a population undergoing selection. The nonstandard feature of the system is that there are terms in the equations that approach infinity as one approaches the boundary. The numerical recipe is patterned after the LU decomposition for tridiagonal matrices. Although there is no analytic proof that this method leads to the correct solution, various examples are presented that suggest that the method works. This method allows one to calculate the expected number of segregating sites in a random sample of n genes from a population whose evolution is described by a model which is not selectively neutral.

Genetics, Population↗

How often are polymorphic restriction sites due to a single mutation?

An approximate expression is obtained for the probability that a restriction site, which is polymorphic in a random sample, is a site at which two or more mutations have occurred in the descent to the sample from the most recent common ancestor of the sample. The analysis is based on the assumption that the population from which the sample is obtained is at equilibrium under a selectively neutral Wright-Fisher model. Monte Carlo simulations show that the approximation is quite accurate. For commonly observed levels of genetic variation in humans and in natural populations of Drosophila, it is found that multiple mutations would occur at 5 to 10 percent of polymorphic restriction sites assuming that six-cutter enzymes are used on samples of size 50 to 100. Simulations are also used to investigate the bias and mean square error of four estimators of 4Nu, where N is the population size and u is the neutral mutation rate per nucleotide site. Two of the estimators are biased by approximately 20 percent when levels of variation are similar to those which have been observed in natural populations of Drosophila.

Animals↗

The "hitchhiking effect" revisited.

The number of selectively neutral polymorphic sites in a random sample of genes can be affected by ancestral selectively favored substitutions at linked loci. The degree to which this happens depends on when in the history of the sample the selected substitutions happen, the strength of selection and the amount of crossing over between the sampled locus and the loci at which the selected substitutions occur. This phenomenon is commonly called hitchhiking. Using the coalescent process for a random sample of genes from a selectively neutral locus that is linked to a locus at which selection is taking place, a stochastic, finite population model is developed that describes the steady state effect of hitchhiking on the distribution of the number of selectively neutral polymorphic sites in a random sample. A prediction of the model is that, in regions of low crossing over, strongly selected substitutions in the history of the sample can substantially reduce the number of polymorphic sites in a random sample of genes from that expected under a neutral model.

Base Sequence↗

DNA polymorphism haplotypes of the human apolipoprotein APOA1-APOC3-APOA4 gene cluster.

The genes coding for apolipoproteins A1, C3, and A4 (APOA1, APOC3, APOA4) are closely linked and tandemly organized within a 15-kilobase (kb) DNA segment on the long arm of human chromosome 11. The nucleotide variability of a 61-kb DNA segment containing these genes and their flanking sequences was studied by restriction analysis of a sample of 18 unrelated Northern Europeans using seven different genomic DNA probes. Eleven restriction site polymorphisms located within this DNA segment were used for haplotype analysis of 129 Mediterranean and 67 American black chromosomes. Estimation of the extent of nonrandom association between these polymorphisms indicated considerable linkage disequilibrium within the APOA1-APOC3-APOA4 gene cluster. Several haplotypes arose by recombination, and the rate of recombination within this gene cluster was estimated to be at least 4 times greater than that expected based on uniform recombination. The polymorphism information content of each of these polymorphisms, taken individually, ranges between 0.053 and 0.375, while that of their haplotypes ranges between 0.858 and 0.862. Therefore, DNA polymorphism haplotypes in the APOA1-APOC3-APOA4 gene cluster constitute a highly informative genetic marker on the long arm of human chromosome 11.

Apolipoprotein A-I↗