Search PubMed⌕ Search

Biomedical subjects

M Slatkin

Publications and source records attributed to M Slatkin.

At least 37 records · Page 2Linked to original sources

The sampling distribution of disease-associated alleles.

A theory is developed that provides the sampling distribution of low frequency alleles at a single locus under the assumption that each allele is the result of a unique mutation. The numbers of copies of each allele is assumed to follow a linear birth-death process with sampling. If the population is of constant size, standard results from theory of birth-death processes show that the distribution of numbers of copies of each allele is logarithmic and that the joint distribution of numbers of copies of k alleles found in a sample of size n follows the Ewens sampling distribution. If the population from which the sample was obtained was increasing in size, if there are different selective classes of alleles, or if there are differences in penetrance among alleles, the Ewens distribution no longer applies. Likelihood functions for a given set of observations are obtained under different alternative hypotheses. These results are applied to published data from the BRCA1 locus (associated with early onset breast cancer) and the factor VIII locus (associated with hemophilia A) in humans. In both cases, the sampling distribution of alleles allows rejection of the null hypothesis, but relatively small deviations from the null model can account for the data. In particular, roughly the same population growth rate appears consistent with both data sets.

Alleles↗

Estimating the age of alleles by use of intraallelic variability.

A method is presented for estimating the age of an allele by use of its frequency and the extent of variation among different copies. The method uses the joint distribution of the number of copies in a population sample and the coalescence times of the intraallelic gene genealogy conditioned on the number of copies. The linear birth-death process is used to approximate the dynamics of a rare allele in a finite population. A maximum-likelihood estimate of the age of the allele is obtained by Monte Carlo integration over the coalescence times. The method is applied to two alleles at the cystic fibrosis (CFTR) locus, deltaF508 and G542X, for which intraallelic variability at three intronic microsatellite loci has been examined. Our results indicate that G542X is somewhat older than deltaF508. Although absolute estimates depend on the mutation rates at the microsatellite loci, our results support the hypothesis that deltaF508 arose < 500 generations (approximately 10,000 years) ago.

Alleles↗

Microsatellites: evolution and mutational processes.

Microsatellites (simple sequence repeats) are ubiquitous in eukaryotic genomes, and they are highly polymorphic. They are currently the primary tools for most genetic mapping and for studies comparing the differentiation of human and other mammalian populations. More and more inherited human diseases are now recognized as resulting from mutations in particular microsatellites, and such microsatellite mutations can serve as markers for some cancers. The majority of microsatellite mutational changes probably consist of insertion or deletion of one or a few repeat units through replication slippage, whereas larger (much rarer) changes are important in producing observed allele distributions. Comparisons of microsatellite allele frequencies between humans and chimpanzees suggest that there are constraints on the overall length of microsatellites. Sequence analyses of microsatellites in diverse human and non-human populations indicate that the structure of many repeats may not be as simple as previously believed, in that alleles differ in base composition as well as in repeat length. Single base changes that result in long uninterrupted repeats may lead to increased mutation rates, including the extreme trinucleotide repeat instability responsible for several inherited diseases.

Animals↗

Interaction of selection and recombination in the fixation of negative-epistatic genes.

We investigated the interaction of recombination and selection on the process of fixation of two linked loci with epistatic interactions in fitness. We consider both the probability of fixation of newly arising mutants (the static model) and the time to fixation under continued mutation (the dynamic model). Our results show that the fixation of a new advantageous combination is facilitated by higher fitness of the advantageous genotype and by weaker selection against the intermediate deleterious genotypes. Fixation occurs more rapidly when the recombination rates are small, except when selection against intermediate genotypes is weak and selection in favour of the double mutant is very strong. In these cases fixation is more rapid when the recombinant rate is large. Mutations of strong effects, deleterious when alone but beneficial when coupled, are fixed more easily than mutations of intermediate effects, at least for large recombination rates. Among the possible pathways the process of fixation might follow, independent substitutions lead to the fixation of the double mutant only when selection is weak. The relative importance of the other pathways depends on the interaction between recombination and selection. The coupled-gamete pathway (i.e. when the population waits until the double mutant appears and then drives it to fixation) is more important as selection intensity increases and the recombination rate is reduced. For all recombination rates, asymmetries in fitness of the intermediate genotypes increase the rate at which fixations occur. Finally, throughout the fixation process, the population will be monomorphic at least at one of the two loci for most of the time, which implies that there would be little opportunity to detect the presence of negative epistasis even if it were important for occasional evolutionary transitions.

Alleles↗

A correction to the exact test based on the Ewens sampling distribution.

The exact test for neutrality based on the Ewens sampling distribution described previously (Slatkin, 1994) is not correct. The problem is that the test as described is based on the probability of the ordered configuration of numbers of alleles, while it should be based on the probability of the unordered configuration. The correctly implemented exact test leads to results that are similar to those from the homozygosity test proposed by Watterson (1977) for relatively small sample sizes but can still differ substantially for larger sample sizes. Programs to perform the exact test are available from the author.

Computer Simulation↗

Testing for linkage disequilibrium in genotypic data using the Expectation-Maximization algorithm.

We generalize an approach suggested by Hill (Heredity, 33, 229-239, 1974) for testing for significant association among alleles at two loci when only genotype and not haplotype frequencies are available. The principle is to use the Expectation-Maximization (EM) algorithm to resolve double heterozygotes into haplotypes and then apply a likelihood ratio test in order to determine whether the resolutions of haplotypes are significantly nonrandom, which is equivalent to testing whether there is statistically significant linkage disequilibrium between loci. The EM algorithm in this case relies on the assumption that genotype frequencies at each locus are in Hardy-Weinberg proportions. This method can accommodate X-linked loci and samples from haplodiploid species. We use three methods for testing significance of the likelihood ratio: the empirical distribution in a large number of randomized data sets, the X2 approximation for the distribution of likelihood ratios, and the Z2 test. The performance of each method is evaluated by applying it to simulated data sets and comparing the tail probability with the tail probability from Fisher's exact test applied to the actual haplotype data. For realistic sample sizes (50-150 individuals) all three methods perform well with two or three alleles per locus, but only the empirical distribution is adequate when there are five to eight alleles per locus, as is typical of hypervariable loci such as microsatellites. The method is applied to a data set of 32 microsatellite loci in a Finnish population and the results confirm the theoretical predictions. We conclude that with highly polymorphic loci, the EM algorithm does lead to a useful test for linkage disequilibrium, but that it is necessary to find the empirical distribution of likelihood ratios in order to perform a test of significance correctly.

Algorithms↗

Gene genealogies within mutant allelic classes.

A coalescent theory of the gene genealogy within an allelic class that arises by a unique mutational event is developed and analyzed. To interpret this theory it was necessary to expand on existing theory for populations of varying size. Two features of the gene genealogy--the average pairwise distance and the total tree length--within the mutant class and within the nonmutant class are found. An index, I, is proposed that describes the extent to which a genealogy is similar to one from a population of constant size (for which I = 0) or to a star genealogy (for which I = 1). The value of I is positive in growing populations and is generally positive for the gene genealogy for the mutant class. The value of I is negative for a population decreasing in size and for the nonmutant class, if the mutant arose recently. The results are discussed in the context of the infinite sites model of mutation, which is appropriate for nucleotide sequence data, and the generalized stepwise mutation model, which is appropriate for microsatellite loci. The same genealogical methods are used to find the probability of at least one recombination event between the nucleotide that defines an allelic class and a marker at a nearby linked site.

Alleles↗

A fine-scale comparison of the human and chimpanzee genomes: linkage, linkage disequilibrium and sequence analysis.

We have performed a fine-scale comparative study of the human and chimpanzee genomes, using linkage, linkage disequilibrium and sequence analyses on microsatellite loci spanning a region of approximately 30 cM on human chromosome 4p. Our results extend the findings of previous studies that indicated virtually complete conservation between the human and chimpanzee genomes at the chromosomal and sub-chromosomal level and support the hypothesis, derived from previous analyses of mitochondrial DNA, that chimpanzee populations are more diverse than human ones. By sequencing several human and chimpanzee alleles of two microsatellites we showed that base substitutions that diminish the length of perfect repeats (but do not change allele sizes) are probably responsible for the low heterozygosity of these loci in chimpanzees; our results suggest that the evolutionary history of microsatellites should not be inferred from comparisons of mean allele lengths between populations or species.

Alleles↗

A measure of population subdivision based on microsatellite allele frequencies.

A new measure of the extent of population subdivision as inferred from allele frequencies at microsatellite loci is proposed and tested with computer simulations. This measure, called R(ST), is analogous to Wright's F(ST) in representing the proportion of variation between populations. It differs in taking explicit account of the mutation process at microsatellite loci, for which a generalized stepwise mutation model appears appropriate. Simulations of subdivided populations were carried out to test the performance of R(ST) and F(ST). It was found that, under the generalized stepwise mutation model, R(ST) provides relatively unbiased estimates of migration rates and times of population divergence while F(ST) tends to show too much population similarity, particularly when migration rates are low or divergence times are long [corrected].

Alleles↗

The distribution of linkage disequilibrium over anonymous genome regions.

Linkage disequilibrium (LD), the association of alleles at different loci, is a powerful tool for genetic mapping and for investigating, at the population level, processes such as recombination, selection, mutation and admixture. Little is known about the distribution of LD across mammalian genomes. Therefore, a survey was undertaken, using microsatellite loci, to evaluate the distribution of LD over several regions of human chromosome 4. Radiation hybrid (RH) and linkage maps provided information on physical and genetic distances between these loci. A Finnish population sample was genotyped using 32 microsatellite loci, and partial marker haplotypes were determined. Assessment of LD was performed, between all pairs of loci, using the Fisher's exact test. LD was detected between several loci that were separated by more than 1 Mb or 1 cM. Detection of LD was strongly associated with small physical distance; its relation to genetic distance was more equivocal. This result may support the hypothesis that linkage maps are relatively inaccurate over small distances. Our results suggest that LD is widely distributed in anonymous regions of the human genome and its use may allow more accurate measurement of small genetic distances than does standard linkage analysis. Alternative explanations are considered for comparisons in which LD is not detected between tightly linked loci.

Alleles↗

Hitchhiking and associative overdominance at a microsatellite locus.

The possible effects of a selected locus on a closely linked microsatellite locus are discussed and analyzed in terms of coalescent theory and models of the mutation process. Background selection caused by recurrent deleterious mutations will reduce the variance of allele size at a microsatellite locus. The occasional substitution of advantageous alleles (genetic hitchhiking) will also reduce the variance, but a high mutation rate at a microsatellite locus can restore the variance relatively rapidly. Overdominance at the selected locus will increase the variance at the microsatellite locus and create partitioning of the variation in allele size among gametes carrying one or the other of the overdominant alleles. These results suggest that neutral microsatellite loci can provide indicators of selective processes at closely linked loci.

Alleles↗

Microsatellite allele frequencies in humans and chimpanzees, with implications for constraints on allele size.

The distributions of allele sizes at eight simple-sequence repeat (SSR) or microsatellite loci in chimpanzees are found and compared with the distributions previously obtained from several human populations. At several loci, the differences in average allele size between chimpanzees and humans are sufficiently small that there might be a constraint on the evolution of average allele size. Furthermore, a model that allows for a bias in the mutation process shows that for some loci a weak bias can account for the observations. Several alleles at one of the loci (Mfd 59) were sequenced. Differences between alleles of different lengths were found to be more complex than previously assumed. An 8-base-pair deletion was present in the nonvariable region of the chimpanzee locus. This locus contains a previously unrecognized repeated region, which is imperfect in humans and perfect in chimpanzees. The apparently greater opportunity for mutation conferred by the two perfect repeat regions in chimpanzees is reflected in the higher variance in repeat number at Mfd 59 in chimpanzees than in humans. These data indicate that interspecific differences in allele length are not always attributable to simple changes in the number of repeats.

Alleles↗

The number of segregating sites in expanding human populations, with implications for estimates of demographic parameters.

The frequency distribution of pairwise differences between sequences of mtDNA has recently been used to estimate the size of human populations before and after a hypothetical episode of rapid population growth and the time at which the population grew. To test the internal consistency of this method, we used three different sets of human mtDNA data and the corresponding demographic parameters estimated from the distribution of pairwise differences to determine by simulation the expected number of segregating sites, S, and its empirical distribution. The results indicate that the observed values of S are significantly lower than expected in two of three cases under the assumption of the infinite-sites model. Further simulations in which mutations were allowed to occur more than once at the same site and in which there was variation in mutation rate among sites show that the expected number of segregating sites can be much lower than under the infinite-site assumption. Nevertheless, the observed value of S is still significantly different from the value expected under the expansion hypothesis in two of three cases.

Animals↗

Maximum-likelihood estimation of molecular haplotype frequencies in a diploid population.

Molecular techniques allow the survey of a large number of linked polymorphic loci in random samples from diploid populations. However, the gametic phase of haplotypes is usually unknown when diploid individuals are heterozygous at more than one locus. To overcome this difficulty, we implement an expectation-maximization (EM) algorithm leading to maximum-likelihood estimates of molecular haplotype frequencies under the assumption of Hardy-Weinberg proportions. The performance of the algorithm is evaluated for simulated data representing both DNA sequences and highly polymorphic loci with different levels of recombination. As expected, the EM algorithm is found to perform best for large samples, regardless of recombination rates among loci. To ensure finding the global maximum likelihood estimate, the EM algorithm should be started from several initial conditions. The present approach appears to be useful for the analysis of nuclear DNA sequences or highly variable loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci.

Algorithms↗

Mutational processes of simple-sequence repeat loci in human populations.

Mutational processes of simple sequence repeats (SSRs) in complex genomes are poorly understood. We examined these processes by introducing a two-phase mutation model. In this model, most mutations are single-step changes, but infrequent large jumps in repeat number also occur. We used computer simulations to determine expected values of statistics that reflect frequency distributions of allele size for the two-phase model and two alternatives, the one-step and geometric models. The theoretical expectations for each model were tested by comparison with observed values for 10 SSR loci genotyped in the Sardinian population, whose genetic and demographic histories have been previously reconstructed. The two-phase model provided the best fit to the data for most of these loci in this population. In the analysis we assumed that the loci were neutral and that this population had undergone rapid population growth. Recent observations made for unstable trinucleotide repeats support our suggestion that frequent small changes and rare large changes in repeat number represent two distinct classes of mutation at SSR loci. We genotyped the same 10 loci in Egyptian and sub-Saharan African samples to assess the utility of SSRs for studying the divergence of populations and found that estimates of interpopulation distances from SSRs were similar to those derived from analysis of mitochondrial DNA.

Gene Frequency↗

Segregation variance after hybridization of isolated populations.

We develop a model to predict the increase in genetic variance of a quantitative character in a hybrid population produced by crossing two previously isolated populations of the same species. The increase in variance in the F2 hybrids, the 'segregation variance', is caused by differences in the average allelic effects at each locus and by linkage disequilibrium among loci. We focus on the case in which the character is additively based and the average value of the character does not differ in the two populations. In that case the predicted segregation variance depends strongly on what is assumed about the genetic basis of the character. If the genetic variance of the character in each population is attributable to loci with numerous alleles of small effect that are in moderate frequency, as in Lande's (1975) model, the segregation variance should increase linearly with time since the populations were isolated, at a rate determined by the inverse of the effective population size. If the genetic variance is attributable to loci with alleles in very low frequency, as in Turelli's (1984) house-of-cards model or in Barton's (1990) model of pleiotropic, deleterious alleles, then the segregation variance in the hybrid population increases at a much lower rate.

Genetic Variation↗

An exact test for neutrality based on the Ewens sampling distribution.

Using the Ewens sampling distribution of selectively neutral alleles in a finite population, it is possible to develop an exact test of neutrality by finding the probability of each configuration with the same sample size and observed number of allelic classes. The exact test provides the probability of obtaining a configuration with the same or smaller probability as the observed configuration under the null hypothesis. The results from the exact test may be quite different from those from the Ewens-Watterson test based on the homozygosity in the sample. The advantages and disadvantages of using an exact test in this and other population genetic contexts are discussed.

Alleles↗