Search PubMed⌕ Search

Biomedical subjects

Keyan Zhao

Publications and source records attributed to Keyan Zhao.

7 recordsLinked to original sources

A nonparametric test reveals selection for rapid flowering in the Arabidopsis genome.

The detection of footprints of natural selection in genetic polymorphism data is fundamental to understanding the genetic basis of adaptation, and has important implications for human health. The standard approach has been to reject neutrality in favor of selection if the pattern of variation at a candidate locus was significantly different from the predictions of the standard neutral model. The problem is that the standard neutral model assumes more than just neutrality, and it is almost always possible to explain the data using an alternative neutral model with more complex demography. Today's wealth of genomic polymorphism data, however, makes it possible to dispense with models altogether by simply comparing the pattern observed at a candidate locus to the genomic pattern, and rejecting neutrality if the pattern is extreme. Here, we utilize this approach on a truly genomic scale, comparing a candidate locus to thousands of alleles throughout the Arabidopsis thaliana genome. We demonstrate that selection has acted to increase the frequency of early-flowering alleles at the vernalization requirement locus FRIGIDA. Selection seems to have occurred during the last several thousand years, possibly in response to the spread of agriculture. We introduce a novel test statistic based on haplotype sharing that embraces the problem of population structure, and so should be widely applicable.

Alleles↗

Association mapping with single-feature polymorphisms.

We develop methods for exploiting "single-feature polymorphism" data, generated by hybridizing genomic DNA to oligonucleotide expression arrays. Our methods enable the use of such data, which can be regarded as very high density, but imperfect, polymorphism data, for genomewide association or linkage disequilibrium mapping. We use a simulation-based power study to conclude that our methods should have good power for organisms like Arabidopsis thaliana, in which linkage disequilibrium is extensive, the reason being that the noisiness of single-feature polymorphism data is more than compensated for by their great number. Finally, we show how power depends on the accuracy with which single-feature polymorphisms are called.

Algorithms↗

Fine mapping--19th century style.

BACKGROUND: There is great interest in the use of computationally intensive methods for fine mapping of marker data. In this paper we develop methods based upon ideas originally proposed 100 years ago in the context of spatial clustering. METHODS: We use spatial clustering of haplotypes as a low-dimensional surrogate for the unobserved genealogy underlying a set of genotype data. In doing so we hope to avoid the computational complexity inherent in explicitly modelling details of the ancestry of the sample, while at the same time capturing the key correlations induced by that ancestry at a much lower computational cost. RESULTS: We benchmark our methods using the simulated Genetic Analysis Workshop 14 data, using 100 replicates of 4 phenotypes to indicate the power of our method. When a functional mutation relating to a trait is actually present, we find evidence for that mutation in 97 out of 100 replicates, on average. CONCLUSION: Our results show that our method has the ability to accurately infer the location of functional mutations from unphased genotype data.

Congresses as Topic↗

Genome-wide association mapping in Arabidopsis identifies previously known flowering time and pathogen resistance genes.

There is currently tremendous interest in the possibility of using genome-wide association mapping to identify genes responsible for natural variation, particularly for human disease susceptibility. The model plant Arabidopsis thaliana is in many ways an ideal candidate for such studies, because it is a highly selfing hermaphrodite. As a result, the species largely exists as a collection of naturally occurring inbred lines, or accessions, which can be genotyped once and phenotyped repeatedly. Furthermore, linkage disequilibrium in such a species will be much more extensive than in a comparable outcrossing species. We tested the feasibility of genome-wide association mapping in A. thaliana by searching for associations with flowering time and pathogen resistance in a sample of 95 accessions for which genome-wide polymorphism data were available. In spite of an extremely high rate of false positives due to population structure, we were able to identify known major genes for all phenotypes tested, thus demonstrating the potential of genome-wide association mapping in A. thaliana and other species with similar patterns of variation. The rate of false positives differed strongly between traits, with more clinal traits showing the highest rate. However, the false positive rates were always substantial regardless of the trait, highlighting the necessity of an appropriate genomic control in association studies.

Arabidopsis↗

The pattern of polymorphism in Arabidopsis thaliana.

We resequenced 876 short fragments in a sample of 96 individuals of Arabidopsis thaliana that included stock center accessions as well as a hierarchical sample from natural populations. Although A. thaliana is a selfing weed, the pattern of polymorphism in general agrees with what is expected for a widely distributed, sexually reproducing species. Linkage disequilibrium decays rapidly, within 50 kb. Variation is shared worldwide, although population structure and isolation by distance are evident. The data fail to fit standard neutral models in several ways. There is a genome-wide excess of rare alleles, at least partially due to selection. There is too much variation between genomic regions in the level of polymorphism. The local level of polymorphism is negatively correlated with gene density and positively correlated with segmental duplications. Because the data do not fit theoretical null distributions, attempts to infer natural selection from polymorphism data will require genome-wide surveys of polymorphism in order to identify anomalous regions. Despite this, our data support the utility of A. thaliana as a model for evolutionary functional genomics.

Arabidopsis↗

The probability and chromosomal extent of trans-specific polymorphism.

Balancing selection may result in trans-specific polymorphism: the maintenance of allelic classes that transcend species boundaries by virtue of being more ancient than the species themselves. At the selected site, gene genealogies are expected not to reflect the species tree. Because of linkage, the same will be true for part of the surrounding chromosomal region. Here we obtain various approximations for the distribution of the length of this region and discuss the practical implications of our results. Our main finding is that the trans-specific region surrounding a single-locus balanced polymorphism is expected to be quite short, probably too short to be readily detectable. Thus lack of obvious trans-specific polymorphism should not be taken as evidence against balancing selection. When trans-specific polymorphism is obvious, on the other hand, it may be reasonable to argue that selection must be acting on multiple sites or that recombination is suppressed in the surrounding region.

Chromosomes↗

Haplotype structure and phenotypic associations in the chromosomal regions surrounding two Arabidopsis thaliana flowering time loci.

The feasibility of using linkage disequilbrium (LD) to fine-map loci underlying natural variation in Arabidopsis thaliana was investigated by looking for associations between flowering time and marker polymorphism in the genomic regions containing two candidate genes, FRI and FLC, both of which are known to contribute to natural variation in flowering. A sample of 196 accessions was used, and polymorphism was assessed by sequencing a total of 17 roughly 500-bp fragments. Using a novel Bayesian algorithm based on haplotype similarity, we demonstrate that LD could have been used to fine-map the FRI gene to a roughly 30-kb region and to identify two common loss-of-function alleles. Interestingly, because of genetic heterogeneity, simple single-marker associations would not have been able to map FRI with nearly the same precision. No clear evidence for previously unknown alleles at either locus was found, but the effect of population structure in causing false positives was evident.

Arabidopsis↗