Search PubMed⌕ Search

Biomedical subjects

Matthew Stephens

Publications and source records attributed to Matthew Stephens.

At least 19 recordsLinked to original sources

Conservation of hotspots for recombination in low-copy repeats associated with the NF1 microdeletion.

Several large-scale studies of human genetic variation have provided insights into processes such as recombination that have shaped human diversity. However, regions such as low-copy repeats (LCRs) have proven difficult to characterize, hindering efforts to understand the processes operating in these regions. We present a detailed study of genetic variation and underlying recombination processes in two copies of an LCR (NF1REPa and NF1REPc) on chromosome 17 involved in the generation of NF1 microdeletions and in a third copy (REP19) on chromosome 19 from which the others originated over 6.7 million years ago. We find evidence for shared hotspots of recombination among the LCRs. REP19 seems to contain hotspots in the same place as the nonallelic recombination hotspots in NF1REPa and NF1REPc. This apparent conservation of patterns of recombination hotspots in moderately diverged paralogous regions contrasts with recent evidence that these patterns are not conserved in less-diverged orthologous regions of chimpanzees.

Animals↗

Automating resequencing-based detection of insertion-deletion polymorphisms.

Structural and insertion-deletion (indel) variants have received considerable recent attention, partly because of their phenotypic consequences. Among these variants, the most common are small indels ( approximately 1-30 bp). Identifying and genotyping indels using sequence traces obtained from diploid samples requires extensive manual review, which makes large-scale studies inconvenient. We report a new algorithm, implemented in available software (PolyPhred version 6.0), to help automate detection and genotyping of indels from sequence traces. The algorithm identifies heterozygous individuals, which permits the discovery of low-frequency indels. It finds 80% of all indel polymorphisms with almost no false positives and finds 97% with a false discovery rate of 10%. Additionally, genotyping accuracy exceeds 99%, and it correctly infers indel length in 96% of the cases. Using this approach, we identify indels in the HapMap ENCODE regions, providing the first report of these polymorphisms in this data set.

Algorithms↗

Insights into recombination from population genetic variation.

Patterns of genetic variation in natural populations are shaped by, and hence carry valuable information about, the underlying recombination process. In the past five years, the increasing availability of large-scale population genetic data on dense sets of markers, coupled with advances in statistical methods for extracting information from these data, have led to several important advances in our understanding of the recombination process in humans. These advances include the identification of large numbers of 'hotspots', where recombination appears to take place considerably more frequently than in the surrounding sequence, and the identification of DNA sequence motifs that are associated with the locations of these hotspots.

Animals↗

Automating sequence-based detection and genotyping of SNPs from diploid samples.

The detection of sequence variation, for which DNA sequencing has emerged as the most sensitive and automated approach, forms the basis of all genetic analysis. Here we describe and illustrate an algorithm that accurately detects and genotypes SNPs from fluorescence-based sequence data. Because the algorithm focuses particularly on detecting SNPs through the identification of heterozygous individuals, it is especially well suited to the detection of SNPs in diploid samples obtained after DNA amplification. It is substantially more accurate than existing approaches and, notably, provides a useful quantitative measure of its confidence in each potential SNP detected and in each genotype called. Calls assigned the highest confidence are sufficiently reliable to remove the need for manual review in several contexts. For example, for sequence data from 47-90 individuals sequenced on both the forward and reverse strands, the highest-confidence calls from our algorithm detected 93% of all SNPs and 100% of high-frequency SNPs, with no false positive SNPs identified and 99.9% genotyping accuracy. This algorithm is implemented in a software package, PolyPhred version 5.0, which is freely available for academic use.

Algorithms↗

A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.

We present a statistical model for patterns of genetic variation in samples of unrelated individuals from natural populations. This model is based on the idea that, over short regions, haplotypes in a population tend to cluster into groups of similar haplotypes. To capture the fact that, because of recombination, this clustering tends to be local in nature, our model allows cluster memberships to change continuously along the chromosome according to a hidden Markov model. This approach is flexible, allowing for both "block-like" patterns of linkage disequilibrium (LD) and gradual decline in LD with distance. The resulting model is also fast and, as a result, is practicable for large data sets (e.g., thousands of individuals typed at hundreds of thousands of markers). We illustrate the utility of the model by applying it to dense single-nucleotide-polymorphism genotype data for the tasks of imputing missing genotypes and estimating haplotypic phase. For imputing missing genotypes, methods based on this model are as accurate or more accurate than existing methods. For haplotype estimation, the point estimates are slightly less accurate than those from the best existing methods (e.g., for unrelated Centre d'Etude du Polymorphisme Humain individuals from the HapMap project, switch error was 0.055 for our method vs. 0.051 for PHASE) but require a small fraction of the computational cost. In addition, we demonstrate that the model accurately reflects uncertainty in its estimates, in that probabilities computed using the model are approximately well calibrated. The methods described in this article are implemented in a software package, fastPHASE, which is available from the Stephens Lab Web site.

Calibration↗

A comparison of phasing algorithms for trios and unrelated individuals.

Knowledge of haplotype phase is valuable for many analysis methods in the study of disease, population, and evolutionary genetics. Considerable research effort has been devoted to the development of statistical and computational methods that infer haplotype phase from genotype data. Although a substantial number of such methods have been developed, they have focused principally on inference from unrelated individuals, and comparisons between methods have been rather limited. Here, we describe the extension of five leading algorithms for phase inference for handling father-mother-child trios. We performed a comprehensive assessment of the methods applied to both trios and to unrelated individuals, with a focus on genomic-scale problems, using both simulated data and data from the HapMap project. The most accurate algorithm was PHASE (v2.1). For this method, the percentages of genotypes whose phase was incorrectly inferred were 0.12%, 0.05%, and 0.16% for trios from simulated data, HapMap Centre d'Etude du Polymorphisme Humain (CEPH) trios, and HapMap Yoruban trios, respectively, and 5.2% and 5.9% for unrelated individuals in simulated data and the HapMap CEPH data, respectively. The other methods considered in this work had comparable but slightly worse error rates. The error rates for trios are similar to the levels of genotyping error and missing data expected. We thus conclude that all the methods considered will provide highly accurate estimates of haplotypes when applied to trio data sets. Running times differ substantially between methods. Although it is one of the slowest methods, PHASE (v2.1) was used to infer haplotypes for the 1 million-SNP HapMap data set. Finally, we evaluated methods of estimating the value of r(2) between a pair of SNPs and concluded that all methods estimated r(2) well when the estimated value was >or=0.8.

Algorithms↗

The effects of genotype-dependent recombination, and transmission asymmetry, on linkage disequilibrium.

A recent sperm-typing study by Jeffreys and Neumann suggested that recombination rates in different individuals at the DNA2 recombination hotspot appeared to be highly dependent on their genotype at a particular A/G SNP, FG11. Specifically, individuals who carried at least one copy of the A allele at this SNP exhibited rates of crossover considerably higher than those of individuals with no copies. Further, recombinant sperm from heterozygous individuals showed a preferential tendency to carry the G allele. We consider the effects of these phenomena on patterns of linkage disequilibrium and find them to be more subtle than might have been expected. In particular, our analysis suggests that, perhaps surprisingly, patterns of LD among chromosomes carrying the "hot" allele (in this case, A) will typically be similar to those among chromosomes carrying the "cold" allele (G).

Alleles↗

Probabilistic segmentation and intensity estimation for microarray images.

We describe a probabilistic approach to simultaneous image segmentation and intensity estimation for complementary DNA microarray experiments. The approach overcomes several limitations of existing methods. In particular, it (a) uses a flexible Markov random field approach to segmentation that allows for a wider range of spot shapes than existing methods, including relatively common 'doughnut-shaped' spots; (b) models the image directly as background plus hybridization intensity, and estimates the two quantities simultaneously, avoiding the common logical error that estimates of foreground may be less than those of the corresponding background if the two are estimated separately; and (c) uses a probabilistic modeling approach to simultaneously perform segmentation and intensity estimation, and to compute spot quality measures. We describe two approaches to parameter estimation: a fast algorithm, based on the expectation-maximization and the iterated conditional modes algorithms, and a fully Bayesian framework. These approaches produce comparable results, and both appear to offer some advantages over other methods. We use an HIV experiment to compare our approach to two commercial software products: Spot and Arrayvision.

Algorithms↗

Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.

Although many algorithms exist for estimating haplotypes from genotype data, none of them take full account of both the decay of linkage disequilibrium (LD) with distance and the order and spacing of genotyped markers. Here, we describe an algorithm that does take these factors into account, using a flexible model for the decay of LD with distance that can handle both "blocklike" and "nonblocklike" patterns of LD. We compare the accuracy of this approach with a range of other available algorithms in three ways: for reconstruction of randomly paired, molecularly determined male X chromosome haplotypes; for reconstruction of haplotypes obtained from trios in an autosomal region; and for estimation of missing genotypes in 50 autosomal genes that have been completely resequenced in 24 African Americans and 23 individuals of European descent. For the autosomal data sets, our new approach clearly outperforms the best available methods, whereas its accuracy in inferring the X chromosome haplotypes is only slightly superior. For estimation of missing genotypes, our method performed slightly better when the two subsamples were combined than when they were analyzed separately, which illustrates its robustness to population stratification. Our method is implemented in the software package PHASE (v2.1.1), available from the Stephens Lab Web site.

Algorithms↗

Assigning African elephant DNA to geographic region of origin: applications to the ivory trade.

Resurgence of illicit trade in African elephant ivory is placing the elephant at renewed risk. Regulation of this trade could be vastly improved by the ability to verify the geographic origin of tusks. We address this need by developing a combined genetic and statistical method to determine the origin of poached ivory. Our statistical approach exploits a smoothing method to estimate geographic-specific allele frequencies over the entire African elephants' range for 16 microsatellite loci, using 315 tissue and 84 scat samples from forest (Loxodonta africana cyclotis) and savannah (Loxodonta africana africana) elephants at 28 locations. These geographic-specific allele frequency estimates are used to infer the geographic origin of DNA samples, such as could be obtained from tusks of unknown origin. We demonstrate that our method alleviates several problems associated with standard assignment methods in this context, and the absolute accuracy of our method is high. Continent-wide, 50% of samples were located within 500 km, and 80% within 932 km of their actual place of origin. Accuracy varied by region (median accuracies: West Africa, 135 km; Central Savannah, 286 km; Central Forest, 411 km; South, 535 km; and East, 697 km). In some cases, allele frequencies vary considerably over small geographic regions, making much finer discriminations possible and suggesting that resolution could be further improved by collection of samples from locations not represented in our study.

Africa↗

Absence of the TAP2 human recombination hotspot in chimpanzees.

Recent experiments using sperm typing have demonstrated that, in several regions of the human genome, recombination does not occur uniformly but instead is concentrated in "hotspots" of 1-2 kb. Moreover, the crossover asymmetry observed in a subset of these has led to the suggestion that hotspots may be short-lived on an evolutionary time scale. To test this possibility, we focused on a region known to contain a recombination hotspot in humans, TAP2, and asked whether chimpanzees, the closest living evolutionary relatives of humans, harbor a hotspot in a similar location. Specifically, we used a new statistical approach to estimate recombination rate variation from patterns of linkage disequilibrium in a sample of 24 western chimpanzees (Pan troglodytes verus). This method has been shown to produce reliable results on simulated data and on human data from the TAP2 region. Strikingly, however, it finds very little support for recombination rate variation at TAP2 in the western chimpanzee data. Moreover, simulations suggest that there should be stronger support if there were a hotspot similar to the one characterized in humans. Thus, it appears that the human TAP2 recombination hotspot is not shared by western chimpanzees. These findings demonstrate that fine-scale recombination rates can change between very closely related species and raise the possibility that rates differ among human populations, with important implications for linkage-disequilibrium based association studies.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Evidence for substantial fine-scale variation in recombination rates across the human genome.

Characterizing fine-scale variation in human recombination rates is important, both to deepen understanding of the recombination process and to aid the design of disease association studies. Current genetic maps show that rates vary on a megabase scale, but studying finer-scale variation using pedigrees is difficult. Sperm-typing experiments have characterized regions where crossovers cluster into 1-2-kb hot spots, but technical difficulties limit the number of studies. An alternative is to use population variation to infer fine-scale characteristics of the recombination process. Several surveys reported 'block-like' patterns of diversity, which may reflect fine-scale recombination rate variation, but limitations of available methods made this impossible to assess. Here, we applied a new statistical method, which overcomes these limitations, to infer patterns of fine-scale recombination rate variation in 74 genes. We found extensive rate variation both within and among genes. In particular, recombination hot spots are a common feature of the human genome: 47% (35 of 74) of genes showed substantive evidence for a hot spot, and many more showed evidence for some rate variation. No primary sequence characteristics are consistently associated with precise hot-spot location, although G+C content and nucleotide diversity are correlated with local recombination rate.

Genome, Human↗

Global effect of PEG-IFN-alpha and ribavirin on gene expression in PBMC in vitro.

Using oligonucleotide microarrays, we have examined the expression of 22,000 genes in peripheral blood cells treated with pegylated interferon-alpha2b (PEG-IFN-alpha) and ribavirin. Treatment with ribavirin had very little effect on gene expression, whereas treatment with PEG-IFN-alpha had a dramatic effect, modulating the expression of approximately 1000 genes (at p < 0.001). In addition to genes previously reported to be induced by type I or type II IFNs, many novel genes were found to be upregulated, including transcription factors, such as ATF3, ATF4, properdin, a key regulator of the complement pathway, a homeobox gene (HESX1), and an RNA editing enzyme (apobec3). Chemokines CXCL10 and CXCL11 were upregulated, whereas CXCL5 was downregulated. Cytokines interleukin-15 (IL-15) and IL-18 were also significantly induced, whereas IL-1alpha and IL-1beta were downregulated. Most other interleukins were not affected. The results of the microarrays were confirmed by kinetic real-time PCR. These data indicate that IFN treatment causes upregulation of genes associated with the stress response, apoptosis, and signaling, and an equal number of genes are downregulated, including those associated with protein synthesis, specific cytokines and chemokines and other biosynthetic functions.

Cells, Cultured↗

A comparison of bayesian methods for haplotype reconstruction from population genotype data.

In this report, we compare and contrast three previously published Bayesian methods for inferring haplotypes from genotype data in a population sample. We review the methods, emphasizing the differences between them in terms of both the models ("priors") they use and the computational strategies they employ. We introduce a new algorithm that combines the modeling strategy of one method with the computational strategies of another. In comparisons using real and simulated data, this new algorithm outperforms all three existing methods. The new algorithm is included in the software package PHASE, version 2.0, available online (http://www.stat.washington.edu/stephens/software.html).

Algorithms↗

Traces of human migrations in Helicobacter pylori populations.

Helicobacter pylori, a chronic gastric pathogen of human beings, can be divided into seven populations and subpopulations with distinct geographical distributions. These modern populations derive their gene pools from ancestral populations that arose in Africa, Central Asia, and East Asia. Subsequent spread can be attributed to human migratory fluxes such as the prehistoric colonization of Polynesia and the Americas, the neolithic introduction of farming to Europe, the Bantu expansion within Africa, and the slave trade.

Africa↗

Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies.

We describe extensions to the method of Pritchard et al. for inferring population structure from multilocus genotype data. Most importantly, we develop methods that allow for linkage between loci. The new model accounts for the correlations between linked loci that arise in admixed populations ("admixture linkage disequilibium"). This modification has several advantages, allowing (1) detection of admixture events farther back into the past, (2) inference of the population of origin of chromosomal regions, and (3) more accurate estimates of statistical uncertainty when linked loci are used. It is also of potential use for admixture mapping. In addition, we describe a new prior model for the allele frequencies within each population, which allows identification of subtle population subdivisions that were not detectable using the existing method. We present results applying the new methods to study admixture in African-Americans, recombination in Helicobacter pylori, and drift in populations of Drosophila melanogaster. The methods are implemented in a program, structure, version 2.0, which is available at http://pritch.bsd.uchicago.edu.

Algorithms↗

Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data.

We introduce a new statistical model for patterns of linkage disequilibrium (LD) among multiple SNPs in a population sample. The model overcomes limitations of existing approaches to understanding, summarizing, and interpreting LD by (i) relating patterns of LD directly to the underlying recombination process; (ii) considering all loci simultaneously, rather than pairwise; (iii) avoiding the assumption that LD necessarily has a "block-like" structure; and (iv) being computationally tractable for huge genomic regions (up to complete chromosomes). We examine in detail one natural application of the model: estimation of underlying recombination rates from population data. Using simulation, we show that in the case where recombination is assumed constant across the region of interest, recombination rate estimates based on our model are competitive with the very best of current available methods. More importantly, we demonstrate, on real and simulated data, the potential of the model to help identify and quantify fine-scale variation in recombination rate from population data. We also outline how the model could be useful in other contexts, such as in the development of more efficient haplotype-based methods for LD mapping.

ATP Binding Cassette Transporter, Subfamily B, Mem↗