Search PubMed⌕ Search

Biomedical subjects

Nick Patterson

Publications and source records attributed to Nick Patterson.

At least 19 recordsLinked to original sources

Admixture mapping identifies 8q24 as a prostate cancer risk locus in African-American men.

A whole-genome admixture scan in 1,597 African Americans identified a 3.8 Mb interval on chromosome 8q24 as significantly associated with susceptibility to prostate cancer [logarithm of odds (LOD) = 7.1]. The increased risk because of inheriting African ancestry is greater in men diagnosed before 72 years of age (P < 0.00032) and may contribute to the epidemiological observation that the higher risk for prostate cancer in African Americans is greatest in younger men (and attenuates with older age). The same region was recently identified through linkage analysis of prostate cancer, followed by fine-mapping. We strongly replicated this association (P < 4.2 x 10(-9)) but find that the previously described alleles do not explain more than a fraction of the admixture signal. Thus, admixture mapping indicates a major, still-unidentified risk gene for prostate cancer at 8q24, motivating intense work to find it.

Black or African American↗

Genetic evidence for complex speciation of humans and chimpanzees.

The genetic divergence time between two species varies substantially across the genome, conveying important information about the timing and process of speciation. Here we develop a framework for studying this variation and apply it to about 20 million base pairs of aligned sequence from humans, chimpanzees, gorillas and more distantly related primates. Human-chimpanzee genetic divergence varies from less than 84% to more than 147% of the average, a range of more than 4 million years. Our analysis also shows that human-chimpanzee speciation occurred less than 6.3 million years ago and probably more recently, conflicting with some interpretations of ancient fossils. Most strikingly, chromosome X shows an extremely young genetic divergence time, close to the genome minimum along nearly its entire length. These unexpected features would be explained if the human and chimpanzee lineages initially diverged, then later exchanged genes before separating permanently.

Animals↗

A comparison of phasing algorithms for trios and unrelated individuals.

Knowledge of haplotype phase is valuable for many analysis methods in the study of disease, population, and evolutionary genetics. Considerable research effort has been devoted to the development of statistical and computational methods that infer haplotype phase from genotype data. Although a substantial number of such methods have been developed, they have focused principally on inference from unrelated individuals, and comparisons between methods have been rather limited. Here, we describe the extension of five leading algorithms for phase inference for handling father-mother-child trios. We performed a comprehensive assessment of the methods applied to both trios and to unrelated individuals, with a focus on genomic-scale problems, using both simulated data and data from the HapMap project. The most accurate algorithm was PHASE (v2.1). For this method, the percentages of genotypes whose phase was incorrectly inferred were 0.12%, 0.05%, and 0.16% for trios from simulated data, HapMap Centre d'Etude du Polymorphisme Humain (CEPH) trios, and HapMap Yoruban trios, respectively, and 5.2% and 5.9% for unrelated individuals in simulated data and the HapMap CEPH data, respectively. The other methods considered in this work had comparable but slightly worse error rates. The error rates for trios are similar to the levels of genotyping error and missing data expected. We thus conclude that all the methods considered will provide highly accurate estimates of haplotypes when applied to trio data sets. Running times differ substantially between methods. Although it is one of the slowest methods, PHASE (v2.1) was used to infer haplotypes for the 1 million-SNP HapMap data set. Finally, we evaluated methods of estimating the value of r(2) between a pair of SNPs and concluded that all methods estimated r(2) well when the estimated value was >or=0.8.

Algorithms↗

Population structure and eigenanalysis.

Current methods for inferring population structure from genetic data do not provide formal significance tests for population differentiation. We discuss an approach to studying population structure (principal components analysis) that was first applied to genetic data by Cavalli-Sforza and colleagues. We place the method on a solid statistical footing, using results from modern statistics to develop formal significance tests. We also uncover a general "phase change" phenomenon about the ability to detect structure in genetic data, which emerges from the statistical theory we use, and has an important implication for the ability to discover structure in genetic data: for a fixed but large dataset size, divergence between two populations (as measured, for example, by a statistic like FST) below a threshold is essentially undetectable, but a little above threshold, detection will be easy. This means that we can predict the dataset size needed to detect structure.

Computer Simulation↗

The case for selection at CCR5-Delta32.

The C-C chemokine receptor 5, 32 base-pair deletion (CCR5-Delta32) allele confers strong resistance to infection by the AIDS virus HIV. Previous studies have suggested that CCR5-Delta32 arose within the past 1,000 y and rose to its present high frequency (5%-14%) in Europe as a result of strong positive selection, perhaps by such selective agents as the bubonic plague or smallpox during the Middle Ages. This hypothesis was based on several lines of evidence, including the absence of the allele outside of Europe and long-range linkage disequilibrium at the locus. We reevaluated this evidence with the benefit of much denser genetic maps and extensive control data. We find that the pattern of genetic variation at CCR5-Delta32 does not stand out as exceptional relative to other loci across the genome. Moreover using newer genetic maps, we estimated that the CCR5-Delta32 allele is likely to have arisen more than 5,000 y ago. While such results can not rule out the possibility that some selection may have occurred at C-C chemokine receptor 5 (CCR5), they imply that the pattern of genetic variation seen at CCR5-Delta32 is consistent with neutral evolution. More broadly, the results have general implications for the design of future studies to detect the signs of positive selection in the human genome.

Alleles↗

A whole-genome admixture scan finds a candidate locus for multiple sclerosis susceptibility.

Multiple sclerosis is a common disease with proven heritability, but, despite large-scale attempts, no underlying risk genes have been identified. Traditional linkage scans have so far identified only one risk haplotype for multiple sclerosis (at HLA on chromosome 6), which explains only a fraction of the increased risk to siblings. Association scans such as admixture mapping have much more power, in principle, to find the weak factors that must explain most of the disease risk. We describe here the first high-powered admixture scan, focusing on 605 African American cases and 1,043 African American controls, and report a locus on chromosome 1 that is significantly associated with multiple sclerosis.

Black or African American↗

Will admixture mapping work to find disease genes?

Admixture mapping is the first experimentally practical method for carrying out a whole-genome association scan, and is thus a promising method for detecting risk factors for common disease. The goal of the community should now be to aggressively test whether the method is useful in practice for localizing disease genes, by carrying out at least three high-powered studies. We also propose a stringent criteria we believe the community should adopt before declaring a statistically significant admixture association to disease.

Chromosome Mapping↗

Relationship of stereochemical and skeletal diversity of small molecules to cellular measurement space.

Systematic and quantitative measurements of the roles of stereochemistry and skeleton-dependent conformational restriction were made using multidimensional screening. We first used diversity-oriented synthesis to synthesize the same number (122) of [10.4.0] bicyclic products (B) and their corresponding monocyclic precursors (M). We measured the ability of these compounds to modulate a broad swath of biology using 40 parallel cell-based assays. We analyzed the results using statistical methods that revealed illuminating relationships between stereochemistry, ring number, and assay outcomes. Conformational restriction by ring-closing metathesis increased the specificity of responses among active compounds and was the dominant factor in global activity patterns. Hierarchical clustering also revealed that stereochemistry was a second dominant factor; whereas the stereochemistry of macrocyclic appendages was a determinant for bicyclic compounds, the stereochemistry of the carbohydrates was a determinant for the monocyclic compounds of global activity patterns. These studies illustrate a quantitative method for measuring stereochemical and skeletal diversity of small molecules and their cellular consequences.

Animals↗

Large-scale copy number polymorphism in the human genome.

The extent to which large duplications and deletions contribute to human genetic variation and diversity is unknown. Here, we show that large-scale copy number polymorphisms (CNPs) (about 100 kilobases and greater) contribute substantially to genomic variation between normal humans. Representational oligonucleotide microarray analysis of 20 individuals revealed a total of 221 copy number differences representing 76 unique CNPs. On average, individuals differed by 11 CNPs, and the average length of a CNP interval was 465 kilobases. We observed copy number variation of 70 different genes within CNP intervals, including genes involved in neurological function, regulation of cell growth, regulation of metabolism, and several genes known to be associated with disease.

Alleles↗

Genetic signatures of strong recent positive selection at the lactase gene.

In most human populations, the ability to digest lactose contained in milk usually disappears in childhood, but in European-derived populations, lactase activity frequently persists into adulthood (Scrimshaw and Murray 1988). It has been suggested (Cavalli-Sforza 1973; Hollox et al. 2001; Enattah et al. 2002; Poulter et al. 2003) that a selective advantage based on additional nutrition from dairy explains these genetically determined population differences (Simoons 1970; Kretchmer 1971; Scrimshaw and Murray 1988; Enattah et al. 2002), but formal population-genetics-based evidence of selection has not yet been provided. To assess the population-genetics evidence for selection, we typed 101 single-nucleotide polymorphisms covering 3.2 Mb around the lactase gene. In northern European-derived populations, two alleles that are tightly associated with lactase persistence (Enattah et al. 2002) uniquely mark a common (~77%) haplotype that extends largely undisrupted for >1 Mb. We provide two new lines of genetic evidence that this long, common haplotype arose rapidly due to recent selection: (1) by use of the traditional F(ST) measure and a novel test based on p(excess), we demonstrate large frequency differences among populations for the persistence-associated markers and for flanking markers throughout the haplotype, and (2) we show that the haplotype is unusually long, given its high frequency--a hallmark of recent selection. We estimate that strong selection occurred within the past 5,000-10,000 years, consistent with an advantage to lactase persistence in the setting of dairy farming; the signals of selection we observe are among the strongest yet seen for any gene in the genome.

Gene Frequency↗

Erralpha and Gabpa/b specify PGC-1alpha-dependent oxidative phosphorylation gene expression that is altered in diabetic muscle.

Recent studies have shown that genes involved in oxidative phosphorylation (OXPHOS) exhibit reduced expression in skeletal muscle of diabetic and prediabetic humans. Moreover, these changes may be mediated by the transcriptional coactivator peroxisome proliferator-activated receptor gamma coactivator-1alpha (PGC-1alpha). By combining PGC-1alpha-induced genome-wide transcriptional profiles with a computational strategy to detect cis-regulatory motifs, we identified estrogen-related receptor alpha (Erralpha) and GA repeat-binding protein alpha as key transcription factors regulating the OXPHOS pathway. Interestingly, the genes encoding these two transcription factors are themselves PGC-1alpha-inducible and contain variants of both motifs near their promoters. Cellular assays confirmed that Erralpha and GA-binding protein a partner with PGC-1alpha in muscle to form a double-positive-feedback loop that drives the expression of many OXPHOS genes. By using a synthetic inhibitor of Erralpha, we demonstrated its key role in PGC-1alpha-mediated effects on gene regulation and cellular respiration. These results illustrate the dissection of gene regulatory networks in a complex mammalian system, elucidate the mechanism of PGC-1alpha action in the OXPHOS pathway, and suggest that Erralpha agonists may ameliorate insulin-resistance in individuals with type 2 diabetes mellitus.

Animals↗

A high-density admixture map for disease gene discovery in african americans.

Admixture mapping (also known as "mapping by admixture linkage disequilibrium," or MALD) provides a way of localizing genes that cause disease, in admixed ethnic groups such as African Americans, with approximately 100 times fewer markers than are required for whole-genome haplotype scans. However, it has not been possible to perform powerful scans with admixture mapping because the method requires a dense map of validated markers known to have large frequency differences between Europeans and Africans. To create such a map, we screened through databases containing approximately 450000 single-nucleotide polymorphisms (SNPs) for which frequencies had been estimated in African and European population samples. We experimentally confirmed the frequencies of the most promising SNPs in a multiethnic panel of unrelated samples and identified 3011 as a MALD map (1.2 cM average spacing). We estimate that this map is approximately 70% informative in differentiating African versus European origins of chromosomal segments. This map provides a practical and powerful tool, which is freely available without restriction, for screening for disease genes in African American patient cohorts. The map is especially appropriate for those diseases that differ in incidence between the parental African and European populations.

Black or African American↗

Methods for high-density admixture mapping of disease genes.

Admixture mapping (also known as "mapping by admixture linkage disequilibrium," or MALD) has been proposed as an efficient approach to localizing disease-causing variants that differ in frequency (because of either drift or selection) between two historically separated populations. Near a disease gene, patient populations descended from the recent mixing of two or more ethnic groups should have an increased probability of inheriting the alleles derived from the ethnic group that carries more disease-susceptibility alleles. The central attraction of admixture mapping is that, since gene flow has occurred recently in modern populations (e.g., in African and Hispanic Americans in the past 20 generations), it is expected that admixture-generated linkage disequilibrium should extend for many centimorgans. High-resolution marker sets are now becoming available to test this approach, but progress will require (a). computational methods to infer ancestral origin at each point in the genome and (b). empirical characterization of the general properties of linkage disequilibrium due to admixture. Here we describe statistical methods to estimate the ancestral origin of a locus on the basis of the composite genotypes of linked markers, and we show that this approach accurately estimates states of ancestral origin along the genome. We apply this approach to show that strong admixture linkage disequilibrium extends, on average, for 17 cM in African Americans. Finally, we present power calculations under varying models of disease risk, sample size, and proportions of ancestry. Studying approximately 2500 markers in approximately 2500 patients should provide power to detect many regions contributing to common disease. A particularly important result is that the power of an admixture mapping study to detect a locus will be nearly the same for a wide range of mixture scenarios: the mixture proportion should be 10%-90% from both ancestral populations.

Alleles↗

Assessing the impact of population stratification on genetic association studies.

Population stratification refers to differences in allele frequencies between cases and controls due to systematic differences in ancestry rather than association of genes with disease. It has been proposed that false positive associations due to stratification can be controlled by genotyping a few dozen unlinked genetic markers. To assess stratification empirically, we analyzed data from 11 case-control and case-cohort association studies. We did not detect statistically significant evidence for stratification but did observe that assessments based on a few dozen markers lack power to rule out moderate levels of stratification that could cause false positive associations in studies designed to detect modest genetic risk factors. After increasing the number of markers and samples in a case-cohort study (the design most immune to stratification), we found that stratification was in fact present. Our results suggest that modest amounts of stratification can exist even in well designed studies.

Case-Control Studies↗

Methods in comparative genomics: genome correspondence, gene identification and regulatory motif discovery.

In Kellis et al. (2003), we reported the genome sequences of S. paradoxus, S. mikatae, and S. bayanus and compared these three yeast species to their close relative, S. cerevisiae. Genomewide comparative analysis allowed the identification of functionally important sequences, both coding and noncoding. In this companion paper we describe the mathematical and algorithmic results underpinning the analysis of these genomes. (1) We present methods for the automatic determination of genome correspondence. The algorithms enabled the automatic identification of orthologs for more than 90% of genes and intergenic regions across the four species despite the large number of duplicated genes in the yeast genome. The remaining ambiguities in the gene correspondence revealed recent gene family expansions in regions of rapid genomic change. (2) We present methods for the identification of protein-coding genes based on their patterns of nucleotide conservation across related species. We observed the pressure to conserve the reading frame of functional proteins and developed a test for gene identification with high sensitivity and specificity. We used this test to revisit the genome of S. cerevisiae, reducing the overall gene count by 500 genes (10% of previously annotated genes) and refining the gene structure of hundreds of genes. (3) We present novel methods for the systematic de novo identification of regulatory motifs. The methods do not rely on previous knowledge of gene function and in that way differ from the current literature on computational motif discovery. Based on genomewide conservation patterns of known motifs, we developed three conservation criteria that we used to discover novel motifs. We used an enumeration approach to select strongly conserved motif cores, which we extended and collapsed into a small number of candidate regulatory motifs. These include most previously known regulatory motifs as well as several noteworthy novel motifs. The majority of discovered motifs are enriched in functionally related genes, allowing us to infer a candidate function for novel motifs. Our results demonstrate the power of comparative genomics to further our understanding of any species. Our methods are validated by the extensive experimental knowledge in yeast and will be invaluable in the study of complex genomes like that of the human.

Algorithms↗

Integrated analysis of protein composition, tissue diversity, and gene regulation in mouse mitochondria.

Mitochondria are tailored to meet the metabolic and signaling needs of each cell. To explore its molecular composition, we performed a proteomic survey of mitochondria from mouse brain, heart, kidney, and liver and combined the results with existing gene annotations to produce a list of 591 mitochondrial proteins, including 163 proteins not previously associated with this organelle. The protein expression data were largely concordant with large-scale surveys of RNA abundance and both measures indicate tissue-specific differences in organelle composition. RNA expression profiles across tissues revealed networks of mitochondrial genes that share functional and regulatory mechanisms. We also determined a larger "neighborhood" of genes whose expression is closely correlated to the mitochondrial genes. The combined analysis identifies specific genes of biological interest, such as candidates for mtDNA repair enzymes, offers new insights into the biogenesis and ancestry of mammalian mitochondria, and provides a framework for understanding the organelle's contribution to human disease.

Animals↗

An integrated haplotype map of the human major histocompatibility complex.

Numerous studies have clearly indicated a role for the major histocompatibility complex (MHC) in susceptibility to autoimmune diseases. Such studies have focused on the genetic variation of a small number of classical human-leukocyte-antigen (HLA) genes in the region. Although these genes represent good candidates, given their immunological roles, linkage disequilibrium (LD) surrounding these genes has made it difficult to rule out neighboring genes, many with immune function, as influencing disease susceptibility. It is likely that a comprehensive analysis of the patterns of LD and variation, by using a high-density map of single-nucleotide polymorphisms (SNPs), would enable a greater understanding of the nature of the observed associations, as well as lead to the identification of causal variation. We present herein an initial analysis of this region, using 201 SNPs, nine classical HLA loci, two TAP genes, and 18 microsatellites. This analysis suggests that LD and variation in the MHC, aside from the classical HLA loci, are essentially no different from those in the rest of the genome. Furthermore, these data show that multi-SNP haplotypes will likely be a valuable means for refining association signals in this region.

Chromosome Mapping↗