Search PubMed⌕ Search

Biomedical subjects

Peter M Visscher

Publications and source records attributed to Peter M Visscher.

At least 19 recordsLinked to original sources

The importance of family-based sampling for biobanks.

Biobanks aim to improve our understanding of health and disease by collecting and analysing diverse biological and phenotypic information in large samples. So far, biobanks have largely pursued a population-based sampling strategy, where the individual is the unit of sampling, and familial relatedness occurs sporadically and by chance. This strategy has been remarkably efficient and successful, leading to thousands of scientific discoveries across multiple research domains, and plans for the next wave of biobanks are underway. In this Perspective, we discuss the strengths and limitations of a complementary sampling strategy for future biobanks based on oversampling of close genetic relatives. Such family-based samples facilitate research that clarifies causal relationships between putative risk factors and outcomes, particularly in estimates of genetic effects, because they enable analyses that reduce or eliminate confounding due to familial and demographic factors. Family-based biobank samples would also shed new light on fundamental questions across multiple fields that are often difficult to explore in population-based samples. Despite the potential for higher costs and greater analytical complexity, the many advantages of family-based samples should often outweigh their potential challenges.

Humans↗

Replicated effects of sex and genotype on gene expression in human lymphoblastoid cell lines.

The expression level for 15,887 transcripts in lymphoblastoid cell lines from 19 monozygotic twin pairs (10 male, 9 female) were analysed for the effects of genotype and sex. On an average, the effect of twin pairs explained 31% of the variance in normalized gene expression levels, consistent with previous broad sense heritability estimates. The effect of sex on gene expression levels was most noticeable on the X chromosome, which contained 15 of the 20 significantly differentially expressed genes. A high concordance was observed between the sex difference test statistics and surveys of genes escaping X chromosome inactivation. Notably, several autosomal genes showed significant differences in gene expression between the sexes despite much of the cellular environment differences being effectively removed in the cell lines. A publicly available gene expression data set from the CEPH families was used to validate the results. The heritability of gene expression levels as estimated from the two data sets showed a highly significant positive correlation, particularly when both estimates were close to one and thus had the smallest standard error. There was a large concordance between the genes significantly differentially expressed between the sexes in the two data sets. Analysis of the variability of probe binding intensities within a probe set indicated that results are robust to the possible presence of polymorphisms in the target sequences.

Adult↗

Residual linkage: why do linkage peaks not disappear after an association study?

Family-based candidate gene and genome-wide association studies are a logical progression from linkage studies for the identification of gene and polymorphisms underlying complex traits. An efficient way to analyse phenotypic and genotypic data is to model linkage and association simultaneously. An important result from such an analysis is whether any evidence for linkage remains after fitting polymorphisms at candidate genes (residual linkage), because this may indicate locus and allelic heterogeneity in the population and will influence subsequent molecular strategies. Here we report that substantial residual linkage is to be expected, even under genetic homogeneity and when the underlying causal polymorphisms are genotyped and fitted in the model. We simulated a powerful design to detect linkage to quantitative trait loci, with 5, 10 or 20 causal SNPs spread throughout the genome. These SNPs were responsible for all genetic variation, and hence for both linkage and association. Residual linkage at the largest linkage peak from a genome-wide scan was substantial, with mean LOD scores of 0.4, 0.7, and 1.4 for the case of 5, 10 and 20 underlying causal SNPs, respectively. For less powerful designs, the proportion of the original LOD scores that remains after association will be even larger. All cases of 'significant' residual linkage are false positives. The reason for the apparent paradox of detecting residual linkage after fitting causal polymorphisms is that the linkage signals at the largest peaks in a genome-scan are severely inflated, even if all peaks correspond to true linkage. Our findings are general and apply to linkage mapping of any phenotype and to any pedigree structure.

Computer Simulation↗

HLA and genomewide allele sharing in dizygotic twins.

Gametic selection during fertilization or the effects of specific genotypes on the viability of embryos may cause a skewed transmission of chromosomes to surviving offspring. A recent analysis of transmission distortion in humans reported significant excess sharing among full siblings. Dizygotic (DZ) twin pairs are a special case of the simultaneous survival of two genotypes, and there have been reports of DZ pairs with excess allele sharing around the HLA locus, a candidate locus for embryo survival. We performed an allele-sharing study of 1,592 DZ twin pairs from two independent Australian cohorts, of which 1,561 pairs were informative for linkage on chromosome 6. We also analyzed allele sharing in 336 DZ twin pairs from The Netherlands. We found no evidence of excess allele sharing, either at the HLA locus or in the rest of the genome. In contrast, we found evidence of a small but significant (P=.003 for the Australian sample) genomewide deficit in the proportion of two alleles shared identical by descent among DZ twin pairs. We reconciled conflicting evidence in the literature for excess genomewide allele sharing by performing a simulation study that shows how undetected genotyping errors can lead to an apparent deficit or excess of allele sharing among sibling pairs, dependent on whether parental genotypes are known. Our results imply that gene-mapping studies based on affected sibling pairs that include DZ pairs will not suffer from false-positive results due to loci involved in embryo survival.

Adolescent↗

Quantitative trait loci (QTL) mapping of resistance to strongyles and coccidia in the free-living Soay sheep (Ovis aries).

A genome-wide scan was performed to detect quantitative trait loci (QTL) for resistance to gastrointestinal parasites and ectoparasitic keds segregating in the free-living Soay sheep population on St. Kilda (UK). The mapping panel consisted of a single pedigree of 882 individuals of which 588 were genotyped. The Soay linkage map used for the scans comprised 251 markers covering the whole genome at average spacing of 15cM. The traits here investigated were the strongyle faecal egg count (FEC), the coccidia faecal oocyst count (FOC) and a count of keds (Melophagus ovinus). QTL mapping was performed by means of variance component analysis so that the genetic parameters of the study traits were also estimated and compared with previous studies in Soay and domestic sheep. Strongyle FEC and coccidia FOC showed moderate heritability (h(2)=0.26 and 0.22, respectively) in lambs but low heritability in adults (h(2)<0.10). Ked count appeared to have very low h(2) in both lambs and adults. Genome scans were performed for the traits with moderate heritability and two genomic regions reached the level of suggestive linkage for coccidia FOC in lambs (logarithm of the odds=2.68 and 2.21 on chromosomes 3 and X, respectively). We believe this is the first study to report a QTL search for parasite resistance in a free-living animal population and therefore may represent a useful reference for similar studies aimed at understanding the genetics of host-parasite co-evolution in the wild.

Animals↗

Bias, precision and heritability of self-reported and clinically measured height in Australian twins.

Many studies of quantitative and disease traits in human genetics rely upon self-reported measures. Such measures are based on questionnaires or interviews and are often cheaper and more readily available than alternatives. However, the precision and potential bias cannot usually be assessed. Here we report a detailed quantitative genetic analysis of stature. We characterise the degree of measurement error by utilising a large sample of Australian twin pairs (857 MZ, 815 DZ) with both clinical and self-reported measures of height. Self-report height measurements are shown to be more variable than clinical measures. This has led to lowered estimates of heritability in many previous studies of stature. In our twin sample the heritability estimate for clinical height exceeded 90%. Repeated measures analysis shows that 2-3 times as many self-report measures are required to recover heritability estimates similar to those obtained from clinical measures. Bivariate genetic repeated measures analysis of self-report and clinical height measures showed an additive genetic correlation >0.98. We show that the accuracy of self-report height is upwardly biased in older individuals and in individuals of short stature. By comparing clinical and self-report measures we also showed that there was a genetic component to females systematically reporting their height incorrectly; this phenomenon appeared to not be present in males. The results from the measurement error analysis were subsequently used to assess the effects of error on the power to detect linkage in a genome scan. Moderate reduction in error (through the use of accurate clinical or multiple self-report measures) increased the effective sample size by 22%; elimination of measurement error led to increases in effective sample size of 41%.

Adult↗

Precision and bias of a normal finite mixture distribution model to analyze twin data when zygosity is unknown: simulations and application to IQ phenotypes on a large sample of twin pairs.

The classification of twin pairs based on zygosity into monozygotic (MZ) or dizygotic (DZ) twins is the basis of most twin analyses. When zygosity information is unavailable, a normal finite mixture distribution (mixture distribution) model can be used to estimate components of variation for continuous traits. The main assumption of this model is that the observed phenotypes on a twin pair are bivariately normally distributed. Any deviation from normality, in particular kurtosis, could produce biased estimates. Using computer simulations and analyses of a wide range of phenotypes from the U.K. Twins' Early Developments Study (TEDS), where zygosity is known, properties of the mixture distribution model were assessed. Simulation results showed that, if normality assumptions were satisfied and the sample size was large (e.g., 2,000 pairs), then the variance component estimates from the mixture distribution model were unbiased and the standard deviation of the difference between heritability estimates from known and unknown zygosity in the range of 0.02-0.20. Unexpectedly, the estimates of heritability of 10 variables from TEDS using the mixture distribution model were consistently larger than those from the conventional (known zygosity) model. This discrepancy was due to violation of the bivariate normality assumption. A leptokurtic distribution of pair difference was observed for all traits (except non-verbal ability scores of MZ twins), even when the univariate distribution of the trait was close to normality. From an independent sample of Australian twins, the heritability estimates for IQ variables were also larger for the mixture distribution model in six out of eight traits, consistent with the observed kurtosis of pair difference. While the known zygosity model is quite robust to the violation of the bivariate normality assumption, this novel finding of widespread kurtosis of the pair difference may suggest that this assumption for analysis of quantitative trait in twin studies may be incorrect and needs revisiting. A possible explanation of widespread kurtosis within zygosity groups is heterogeneity of variance, which could be caused by genetic or environmental factors. For the mixture distribution model, violation of the bivariate normality assumption will produce biased estimates.

Analysis of Variance↗

A simple method to localise pleiotropic susceptibility loci using univariate linkage analyses of correlated traits.

Univariate linkage analysis is used routinely to localise genes for human complex traits. Often, many traits are analysed but the significance of linkage for each trait is not corrected for multiple trait testing, which increases the experiment-wise type-I error rate. In addition, univariate analyses do not realise the full power provided by multivariate data sets. Multivariate linkage is the ideal solution but it is computationally intensive, so genome-wide analysis and evaluation of empirical significance are often prohibitive. We describe two simple methods that efficiently alleviate these caveats by combining P-values from multiple univariate linkage analyses. The first method estimates empirical pointwise and genome-wide significance between one trait and one marker when multiple traits have been tested. It is as robust as an appropriate Bonferroni adjustment, with the advantage that no assumptions are required about the number of independent tests performed. The second method estimates the significance of linkage between multiple traits and one marker and, therefore, it can be used to localise regions that harbour pleiotropic quantitative trait loci (QTL). We show that this method has greater power than individual univariate analyses to detect a pleiotropic QTL across different situations. In addition, when traits are moderately correlated and the QTL influences all traits, it can outperform formal multivariate VC analysis. This approach is computationally feasible for any number of traits and was not affected by the residual correlation between traits. We illustrate the utility of our approach with a genome scan of three asthma traits measured in families with a twin proband.

Asthma↗

Analysis of pooled DNA samples on high density arrays without prior knowledge of differential hybridization rates.

Array based DNA pooling techniques facilitate genome-wide scale genotyping of large samples. We describe a structured analysis method for pooled data using internal replication information in large scale genotyping sets. The method takes advantage of information from single nucleotide polymorphisms (SNPs) typed in parallel on a high density array to construct a test statistic with desirable statistical properties. We utilize a general linear model to appropriately account for the structured multiple measurements available with array data. The method does not require the use of additional arrays for the estimation of unequal hybridization rates and hence scales readily to accommodate arrays with several hundred thousand SNPs. Tests for differences between cases and controls can be conducted with very few arrays. We demonstrate the method on 384 endometriosis cases and controls, typed using Affymetrix Genechip(c) HindIII 50 K arrays. For a subset of this data there were accurate measures of hybridization rates available. Assuming equal hybridization rates is shown to have a negligible effect upon the results. With a total of only six arrays, the method extracted one-third of the information (in terms of equivalent sample size) available with individual genotyping (requiring 768 arrays). With 20 arrays (10 for cases, 10 for controls), over half of the information could be extracted from this sample.

Case-Control Studies↗

A simple linear regression method for quantitative trait loci linkage analysis with censored observations.

Standard quantitative trait loci (QTL) mapping techniques commonly assume that the trait is both fully observed and normally distributed. When considering survival or age-at-onset traits these assumptions are often incorrect. Methods have been developed to map QTL for survival traits; however, they are both computationally intensive and not available in standard genome analysis software packages. We propose a grouped linear regression method for the analysis of continuous survival data. Using simulation we compare this method to both the Cox and Weibull proportional hazards models and a standard linear regression method that ignores censoring. The grouped linear regression method is of equivalent power to both the Cox and Weibull proportional hazards methods and is significantly better than the standard linear regression method when censored observations are present. The method is also robust to the proportion of censored individuals and the underlying distribution of the trait. On the basis of linear regression methodology, the grouped linear regression model is computationally simple and fast and can be implemented readily in freely available statistical software.

Chromosome Mapping↗

Assumption-free estimation of heritability from genome-wide identity-by-descent sharing between full siblings.

The study of continuously varying, quantitative traits is important in evolutionary biology, agriculture, and medicine. Variation in such traits is attributable to many, possibly interacting, genes whose expression may be sensitive to the environment, which makes their dissection into underlying causative factors difficult. An important population parameter for quantitative traits is heritability, the proportion of total variance that is due to genetic factors. Response to artificial and natural selection and the degree of resemblance between relatives are all a function of this parameter. Following the classic paper by R. A. Fisher in 1918, the estimation of additive and dominance genetic variance and heritability in populations is based upon the expected proportion of genes shared between different types of relatives, and explicit, often controversial and untestable models of genetic and non-genetic causes of family resemblance. With genome-wide coverage of genetic markers it is now possible to estimate such parameters solely within families using the actual degree of identity-by-descent sharing between relatives. Using genome scans on 4,401 quasi-independent sib pairs of which 3,375 pairs had phenotypes, we estimated the heritability of height from empirical genome-wide identity-by-descent sharing, which varied from 0.374 to 0.617 (mean 0.498, standard deviation 0.036). The variance in identity-by-descent sharing per chromosome and per genome was consistent with theory. The maximum likelihood estimate of the heritability for height was 0.80 with no evidence for non-genetic causes of sib resemblance, consistent with results from independent twin and family studies but using an entirely separate source of information. Our application shows that it is feasible to estimate genetic variance solely from within-family segregation and provides an independent validation of previously untestable assumptions. Given sufficient data, our new paradigm will allow the estimation of genetic variation for disease susceptibility and quantitative traits that is free from confounding with non-genetic factors and will allow partitioning of genetic variation into additive and non-additive components.

Body Height↗

The value of relatives with phenotypes but missing genotypes in association studies for quantitative traits.

The additional statistical power of association studies for quantitative traits was derived when ungenotyped relatives with phenotypes are included in the analysis. It was shown that the extra power is a simple function of the coefficient of additive genetic relationship and the phenotypic correlation coefficient between the genotyped and ungenotyped relatives. For close relatives, such as pairs of fullsibs and identical twin pairs, gains in power in the range of 10 to 30% are achieved if only one of the pair is genotyped. The theoretical results were verified by simulations. It was shown that ignoring the error in estimating the genotype of the ungenotyped relative has little impact on the estimates and on statistical power, consistent with results from quantitative trait loci (QTL) linkage studies. For genome-wide association studies in which not all relatives with phenotypes can be genotyped, our study provides a prediction of the additional power of an analysis that includes phenotypes on ungenotyped individuals, and can be used in experimental design. We show that a two-step procedure, in which missing genotypes are imputed and subsequently an association analysis is performed, is efficient and powerful.

Genetic Predisposition to Disease↗

False disease region identification from identity-by-descent haplotype sharing in the presence of phenocopies.

Linkage analysis (either parametric or nonparametric) is commonly applied to identify chromosomal regions using related individuals affected by disease. In complex disease the incomplete relationship between phenotype and genotype can be modeled using a phenocopy parameter, the probability that an individual is affected given they do not carry the disease mutation of interest, and a nonpenetrance parameter, the probability that an individual is not affected given they do carry the disease mutation of interest. If the linkage phase between multiple markers and a putative disease locus is known, then haplotypes carrying the mutation can, in principle, be identified by comparing the chromosome segments that are shared identical-by-descent (IBD) across affected individuals. We consider here the effect of a nonzero phenocopy rate on the linkage peak and hence upon the identification of disease haplotypes that are shared IBD between affected individuals. We show, by theory and computer simulation, that in diseases for which there is a nonzero phenocopy rate, the chromosomal regions identified may not include the true disease locus. We utilize a LOD-1 confidence interval for a widely used nonparametric linkage statistic. We find that in small/moderate samples this confidence interval may be inappropriate. We give specific examples where the phenocopy rates are nonnegligible in some complex diseases. The success of further work to identify the causal mutations underlying the linkage peaks in these diseases will depend on researchers allowing for the presence of phenocopies by examining appropriately wide regions around the initial positive linkage finding.

Alleles↗

Identification of twin pairs from large population-based samples.

The basis of most twin studies is the ascertainment of twins, often through twin registries, and determination of zygosity. The current rate of twin births in many industrialized countries implies that in the near future around 3% or more of individuals will be a twin. Hence, there are and will be a lot of twins around and many of those will not participate in twin studies. However, if large population-based samples are available that include appropriate identifiers, then twins can be detected and twin studies performed, even in the absence of zygosity information. We quantified the number of twin pairs that could be detected from a longitudinal survey in the Netherlands, which aims to answer questions about educational strategies and performance in primary education in the Netherlands. We detected 2865 twin pairs if we used a coded name identifier, date of birth, school, grade and year of survey, which is 2.01% of 284,945 pupils in five cohorts. Relaxing our selection criteria increased the number of apparent twin pairs identified, most of which are false positives due to chance matching of identification criteria. We show that the intraclass correlation on measured phenotypes can be used as a quality control measure for twin identification, and quantify the proportion of false negatives (true twin pairs not identified) due to missing data and data coding errors. We compared our estimated rate of twins in the sample to census data and estimate that with our most stringent selection criteria we detect more than 80% of all twin pairs in the sample. We conclude that the identification of twin pairs from large population-based samples is feasible, rapid and accurate if the appropriate identifiers are available, and that twin pairs from such sources are a valuable resource for studies to answer scientific question about twins versus nontwins and about genetic and environmental factors of twin resemblance.

Female↗

A note on the asymptotic distribution of likelihood ratio tests to test variance components.

When using maximum likelihood methods to estimate genetic and environmental components of (co)variance, it is common to test hypotheses using likelihood ratio tests, since such tests have desirable asymptotic properties. In particular, the standard likelihood ratio test statistic is assumed asymptotically to follow a chi2 distribution with degrees of freedom equal to the number of parameters tested. Using the relationship between least squares and maximum likelihood estimators for balanced designs, it is shown why the asymptotic distribution of the likelihood ratio test for variance components does not follow a chi2 distribution with degrees of freedom equal to the number of parameters tested when the null hypothesis is true. Instead, the distribution of the likelihood ratio test is a mixture of chi2 distributions with different degrees of freedom. Implications for testing variance components in twin designs and for quantitative trait loci mapping are discussed. The appropriate distribution of the likelihood ratio test statistic should be used in hypothesis testing and model selection.

Chromosome Mapping↗

Development of a linkage map and mapping of phenotypic polymorphisms in a free-living population of Soay sheep (Ovis aries).

An understanding of the determinants of trait variation and the selective forces acting on it in natural populations would give insights into the process of evolution. The combination of long-term studies of individuals living in the wild and better genomic resources for nonmodel organisms makes achieving this goal feasible. This article reports the development of a complete linkage map in a pedigree of free-living Soay sheep on St. Kilda and its application to mapping the loci responsible for three morphological polymorphisms for which the maintenance of variation demands explanation. The map was derived from 251 microsatellite and four allozyme markers and covers 3350 cM (approximately 90% of the sheep genome) at approximately 15-cM intervals. Marker order was consistent with the published sheep map with the exception of one region on chromosome 1 and one on chromosome 12. Coat color maps to chromosome 2 where a strong candidate gene, tyrosinase-related protein 1 (TYRP1), has also been mapped. Coat pattern maps to chromosome 13, close to the candidate locus Agouti. Horn type maps to chromosome 10, a location similar to that previously identified in domestic sheep. These findings represent an advance in the dissection of the genetic diversity in the wild and provide the foundation for QTL analyses in the study population.

Animals↗

Replicated linkage for eye color on 15q using comparative ratings of sibling pairs.

: The aim of the study was to perform a genetic linkage analysis for eye color, for comparative data. Similarity in eye color of mono- and dizygotic twins was rated by the twins' mother, their father and/or the twins themselves. For 4,748 twin pairs the similarity in eye color was available on a three point scale ("not at all alike"-"somewhat alike"-"completely alike"), absolute eye color on individuals was not assessed. The probability that twins were alike for eye color was calculated as a weighted average of the different responses of all respondents on several different time points. The mean probability of being alike for eye color was 0.98 for MZ twins (2,167 pairs), whereas the mean probability for DZ twins was 0.46 (2,537 pairs), suggesting very high heritability for eye color. For 294 DZ twin pairs genome-wide marker data were available. The probability of being alike for eye color was regressed on the average amount of IBD sharing. We found a peak LOD-score of 2.9 at chromosome 15q, overlapping with the region recently implicated for absolute ratings of eye color in Australian twins [Zhu, G., Evans, D. M., Duffy, D. L., Montgomery, G. W., Medland, S. E., Gillespie, N. A., Ewen, K. R., Jewell, M., Liew, Y. W., Hayward, N. K., Sturm, R. A., Trent, J. M., and Martin, N. G. (2004). Twin Res. 7:197-210] and containing the OCA2 gene, which is the major candidate gene for eye color [Sturm, R. A. Teasdale, R. D, and Box, N. F. (2001). Gene 277:49-62]. Our results demonstrate that comparative measures on relatives can be used in genetic linkage analysis.

Adolescent↗