Search PubMed⌕ Search

Biomedical subjects

Goncalo R Abecasis

Publications and source records attributed to Goncalo R Abecasis.

7 recordsLinked to original sources

The Genetic Determinants and Genomic Consequences of Non-Leukemogenic Somatic Point Mutations.

Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as point mutations or mosaic chromosomal alterations (mCAs) in genes associated with hematologic malignancy, are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we used 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantified CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we developed a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We performed a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. After fine-mapping and variant-to-gene analyses, we identified seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1), and one locus associated with a sex-associated mutation pathway (SRGAP2C). We performed a secondary analysis excluding individuals with mCAs, finding that the genetic architecture was largely unaffected by their inclusion. Functional analyses of SMC4 and NRIP1 implicated altered HSC self-renewal and proliferation as the primary mediator of mutation burden in blood. We then performed comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we performed phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count and increased risk for incident peripheral artery disease, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS without a paired-tissue sample and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.

Journal Article↗

A comparison of phasing algorithms for trios and unrelated individuals.

Knowledge of haplotype phase is valuable for many analysis methods in the study of disease, population, and evolutionary genetics. Considerable research effort has been devoted to the development of statistical and computational methods that infer haplotype phase from genotype data. Although a substantial number of such methods have been developed, they have focused principally on inference from unrelated individuals, and comparisons between methods have been rather limited. Here, we describe the extension of five leading algorithms for phase inference for handling father-mother-child trios. We performed a comprehensive assessment of the methods applied to both trios and to unrelated individuals, with a focus on genomic-scale problems, using both simulated data and data from the HapMap project. The most accurate algorithm was PHASE (v2.1). For this method, the percentages of genotypes whose phase was incorrectly inferred were 0.12%, 0.05%, and 0.16% for trios from simulated data, HapMap Centre d'Etude du Polymorphisme Humain (CEPH) trios, and HapMap Yoruban trios, respectively, and 5.2% and 5.9% for unrelated individuals in simulated data and the HapMap CEPH data, respectively. The other methods considered in this work had comparable but slightly worse error rates. The error rates for trios are similar to the levels of genotyping error and missing data expected. We thus conclude that all the methods considered will provide highly accurate estimates of haplotypes when applied to trio data sets. Running times differ substantially between methods. Although it is one of the slowest methods, PHASE (v2.1) was used to infer haplotypes for the 1 million-SNP HapMap data set. Finally, we evaluated methods of estimating the value of r(2) between a pair of SNPs and concluded that all methods estimated r(2) well when the estimated value was >or=0.8.

Algorithms↗

Meta-analysis of genome scans of age-related macular degeneration.

A genetic contribution to the development of age-related macular degeneration (AMD) is well established. Several genome-wide linkage studies have identified a number of putative susceptibility loci for AMD but only a few of these regions have been replicated in independent studies. Here, we perform a meta-analysis of six AMD genome screens using the genome-scan meta-analysis method, which allows linkage results from several studies to be combined, providing greater power to identify regions that show only weak evidence for linkage in individual studies. Results from non-parametric analysis for a broad AMD clinical phenotype (including two studies with quantitative traits) were extracted. For each study, 120 genomic bins of approximately 30 cM were defined and ranked according to maximum evidence for linkage within each bin. Bin ranks were weighted according to study size and summed across all studies; the summed rank (SR) for each bin was assessed empirically for significance using permutation methods. A high SR indicates a region with consistent evidence for linkage across studies. The strongest evidence for an AMD susceptibility locus was found on chromosome 10q26 where genome-wide significant linkage was observed (P=0.00025). Several other regions met the empirical significance criteria for bins likely to contain linked loci including adjacent pairs of bins on chromosomes 1q, 2p, 3p and 16. Several of the regions identified here showed only weak evidence for linkage in the individual studies. These results will help prioritize regions for future positional and functional candidate gene studies in AMD.

Aging↗

Joint modeling of linkage and association: identifying SNPs responsible for a linkage signal.

Once genetic linkage has been identified for a complex disease, the next step is often association analysis, in which single-nucleotide polymorphisms (SNPs) within the linkage region are genotyped and tested for association with the disease. If a SNP shows evidence of association, it is useful to know whether the linkage result can be explained, in part or in full, by the candidate SNP. We propose a novel approach that quantifies the degree of linkage disequilibrium (LD) between the candidate SNP and the putative disease locus through joint modeling of linkage and association. We describe a simple likelihood of the marker data conditional on the trait data for a sample of affected sib pairs, with disease penetrances and disease-SNP haplotype frequencies as parameters. We estimate model parameters by maximum likelihood and propose two likelihood-ratio tests to characterize the relationship of the candidate SNP and the disease locus. The first test assesses whether the candidate SNP and the disease locus are in linkage equilibrium so that the SNP plays no causal role in the linkage signal. The second test assesses whether the candidate SNP and the disease locus are in complete LD so that the SNP or a marker in complete LD with it may account fully for the linkage signal. Our method also yields a genetic model that includes parameter estimates for disease-SNP haplotype frequencies and the degree of disease-SNP LD. Our method provides a new tool for detecting linkage and association and can be extended to study designs that include unaffected family members.

Alleles↗

A note on exact tests of Hardy-Weinberg equilibrium.

Deviations from Hardy-Weinberg equilibrium (HWE) can indicate inbreeding, population stratification, and even problems in genotyping. In samples of affected individuals, these deviations can also provide evidence for association. Tests of HWE are commonly performed using a simple chi2 goodness-of-fit test. We show that this chi2 test can have inflated type I error rates, even in relatively large samples (e.g., samples of 1,000 individuals that include approximately 100 copies of the minor allele). On the basis of previous work, we describe exact tests of HWE together with efficient computational methods for their implementation. Our methods adequately control type I error in large and small samples and are computationally efficient. They have been implemented in freely available code that will be useful for quality assessment of genotype data and for the detection of genetic association or population stratification in very large data sets.

Genetics, Population↗

Regression-based sib pair linkage analysis for binary traits.

The Haseman-Elston (HE) regression method offers a mathematically and computationally simpler alternative to variance-components (VC) models for the linkage analysis of quantitative traits. However, current versions of HE regression and VC models are not optimised for binary traits. Here, we present a modified HE regression and a liability-threshold VC model for binary-traits. The new HE method is based on the regression of a linear combination of the trait squares and the trait cross-product on the proportion of alleles identical by descent (IBD) at the putative locus, for sibling pairs. We have implemented both the new HE regression-based method and have performed analytic and simulation studies to assess its type 1 error rate and power under a range of conditions. These studies showed that the new HE method is well-behaved under the null hypothesis in large samples, is more powerful than both the original and the revisited HE methods, and is approximately equivalent in power to the liability-threshold VC model.

Data Interpretation, Statistical↗

Genetic variation in the 22q11 locus and susceptibility to schizophrenia.

An increased prevalence of microdeletions at the 22q11 locus has been reported in samples of patients with schizophrenia. 22q11 microdeletions represent the highest known genetic risk factor for the development of schizophrenia, second only to that of the monozygotic cotwin of an affected individual or the offspring of two schizophrenic parents. It is therefore clear that a schizophrenia susceptibility locus maps to chromosome 22q11. In light of evidence for suggestive linkage for schizophrenia in this region, we hypothesized that, whereas deletions of chromosome 22q11 may account for only a small proportion of schizophrenia cases in the general population (up to approximately 2%), nondeletion variants of individual genes within the 22q11 region may make a larger contribution to susceptibility to schizophrenia in the wider population. By studying a dense collection of markers (average one single nucleotide polymorphism20 kb over 1.5 Mb) in the vicinity of the 22q11 locus, in both family- and population-based samples, we present here results consistent with this assumption. Moreover, our results are consistent with contribution from more than one gene to the strikingly increased disease risk associated with this locus. Finer-scale haplotype mapping has identified two subregions within the 1.5-Mb locus that are likely to harbor candidate schizophrenia susceptibility genes.

Amino Acid Sequence↗