Search PubMed⌕ Search

Biomedical subjects

Michael G B Blum

Publications and source records attributed to Michael G B Blum.

6 recordsLinked to original sources

Sampling properties of homozygosity-based statistics for linkage disequilibrium.

Homozygosity-based statistics such as Ohta's identity-in-state (IIS) excess offer the potential to measure linkage disequilibrium for multiallelic loci in small samples. However, previous observations have suggested that for independent loci, in small samples these statistics might produce values that more frequently lie on one side rather than on the other side of zero. Here we investigate the sampling properties of the IIS excess. We find that for any pair of independent polymorphic loci, as sample size n approaches infinity, the sampling distribution of the IIS excess approaches a normal distribution. For large samples, the IIS excess tends towards symmetry around zero, and the probabilities of positive and of negative IIS excess both approach 1/2. Surprisingly, however, we also find that for sufficiently large n, independent loci can be chosen so that the probability of a sample having positive IIS excess is arbitrarily close to either 0 or 1. The results are applied to interpretation of data from human populations, and we conclude that before employing homozygosity-based statistics to measure LD in a particular sample, especially for loci with either very small or very large homozygosities, it is useful to verify that loci with the observed homozygosity values are not likely to produce a large bias in IIS excess in samples of the given size.

Algorithms↗

Matrilineal fertility inheritance detected in hunter-gatherer populations using the imbalance of gene genealogies.

Fertility inheritance, a phenomenon in which an individual's number of offspring is positively correlated with his or her number of siblings, is a cultural process that can have a strong impact on genetic diversity. Until now, fertility inheritance has been detected primarily using genealogical databases. In this study, we develop a new method to infer fertility inheritance from genetic data in human populations. The method is based on the reconstruction of the gene genealogy of a sample of sequences from a given population and on the computation of the degree of imbalance in this genealogy. We show indeed that this level of imbalance increases with the level of fertility inheritance, and that other phenomena such as hidden population structure are unlikely to generate a signal of imbalance in the genealogy that would be confounded with fertility inheritance. By applying our method to mtDNA samples from 37 human populations, we show that matrilineal fertility inheritance is more frequent in hunter-gatherer populations than in food-producer populations. One possible explanation for this result is that in hunter-gatherer populations, individuals belonging to large kin networks may benefit from stronger social support and may be more likely to have a large number of offspring.

Allelic Imbalance↗

Low levels of genetic divergence across geographically and linguistically diverse populations from India.

Ongoing modernization in India has elevated the prevalence of many complex genetic diseases associated with a western lifestyle and diet to near-epidemic proportions. However, although India comprises more than one sixth of the world's human population, it has largely been omitted from genomic surveys that provide the backdrop for association studies of genetic disease. Here, by genotyping India-born individuals sampled in the United States, we carry out an extensive study of Indian genetic variation. We analyze 1,200 genome-wide polymorphisms in 432 individuals from 15 Indian populations. We find that populations from India, and populations from South Asia more generally, constitute one of the major human subgroups with increased similarity of genetic ancestry. However, only a relatively small amount of genetic differentiation exists among the Indian populations. Although caution is warranted due to the fact that United States-sampled Indian populations do not represent a random sample from India, these results suggest that the frequencies of many genetic variants are distinctive in India compared to other parts of the world and that the effects of population heterogeneity on the production of false positives in association studies may be smaller in Indians (and particularly in Indian-Americans) than might be expected for such a geographically and linguistically diverse subset of the human population.

Alleles↗

On statistical tests of phylogenetic tree imbalance: the Sackin and other indices revisited.

We investigate the distribution of statistical measures of tree imbalance in large phylogenies. More specifically, we study normalized versions of the Sackin's index and the number of subtrees of given sizes. Using the connection with structures from theoretical computer science, we provide precise description for the limiting distribution under the null hypothesis of Yule trees. Corrected p-values are then computed, and the statistical power of these statistics for testing the Yule model against a model of biased speciation is evaluated from simulations. As an illustration, the tests are applied to the HIV-1 reconstructed phylogeny.

Acquired Immunodeficiency Syndrome↗

Brownian models and coalescent structures.

Brownian motions on coalescent structures have a biological relevance, either as an approximation of the stepwise mutation model for microsatellites, or as a model of spatial evolution considering the locations of individuals at successive generations. We discuss estimation procedures for the dispersal parameter of a Brownian motion defined on coalescent trees. First, we consider the mean square distance unbiased estimator and compute its variance. In a second approach, we introduce a phylogenetic estimator. Given the UPGMA topology, the likelihood of the parameter is computed thanks to a new dynamical programming method. By a proper correction, an unbiased estimator is derived from the pseudomaximum of the likelihood. The last approach consists of computing the likelihood by a Markov chain Monte Carlo sampling method. In the one-dimensional Brownian motion, this method seems less reliable than pseudomaximum-likelihood.

Algorithms↗