Search PubMed⌕ Search

Biomedical subjects

Ola Hössjer

Publications and source records attributed to Ola Hössjer.

10 recordsLinked to original sources

Estimating the parameters of the operational model of pharmacological agonism.

The aim of this work is practical. We show that the parameters of the widely used operational model of pharmacological agonism are difficult to estimate from single dose-response curves. The parameters can be estimated using pairs of dose-response curves (usually treatment and control) sharing some parameters. Confidence bands for the estimators are developed. In the case of multiple dose-response curve pairs one can employ a non-linear mixed effects model to allow for inter-individual variation. The point estimates and the confidence intervals thus obtained are similar to the more naive construction based on mean and standard errors of parameter estimates. To test for difference of certain parameters between treatment and control we employ a permutation test and Wald's test.

Animals↗

Modeling the effect of inbreeding among founders in linkage analysis.

In this paper, we present a unified mathematical model for linkage analysis that allows for inbreeding among founders in all families. The identical by descent (IBD) configuration of each pedigree is modeled as a Markov process containing two parameters; the inverse inbreeding and kinship coefficient and a rate parameter proportional to the inverse expected length of chromosome segments shared IBD by two different founder haplotypes. We use hidden Markov models and define a forward-backward algorithm for computing the conditional IBD-distribution given marker data, thereby extending the multipoint method of Lander and Green [1987. Construction of multilocus genetic maps in humans, Proc. Natl. Acad. Sci. USA 84, 2363-2367] to situations where founders are inbred. Our methodology is valid for arbitrary pedigree structures. Simulation and theoretical approximations for nonparametric linkage (NPL) analysis based on affected sib pairs reveal that NPL scores are inflated and type 1 errors increased when the inbreeding coefficient or rate parameter is underestimated. When the parents are genotyped, we present a general way of modifying the score function to drastically reduce this effect.

Algorithms↗

Methodological study of affine transformations of gene expression data with proposed robust non-parametric multi-dimensional normalization method.

BACKGROUND: Low-level processing and normalization of microarray data are most important steps in microarray analysis, which have profound impact on downstream analysis. Multiple methods have been suggested to date, but it is not clear which is the best. It is therefore important to further study the different normalization methods in detail and the nature of microarray data in general. RESULTS: A methodological study of affine models for gene expression data is carried out. Focus is on two-channel comparative studies, but the findings generalize also to single- and multi-channel data. The discussion applies to spotted as well as in-situ synthesized microarray data. Existing normalization methods such as curve-fit ("lowess") normalization, parallel and perpendicular translation normalization, and quantile normalization, but also dye-swap normalization are revisited in the light of the affine model and their strengths and weaknesses are investigated in this context. As a direct result from this study, we propose a robust non-parametric multi-dimensional affine normalization method, which can be applied to any number of microarrays with any number of channels either individually or all at once. A high-quality cDNA microarray data set with spike-in controls is used to demonstrate the power of the affine model and the proposed normalization method. CONCLUSION: We find that an affine model can explain non-linear intensity-dependent systematic effects in observed log-ratios. Affine normalization removes such artifacts for non-differentially expressed genes and assures that symmetry between negative and positive log-ratios is obtained, which is fundamental when identifying differentially expressed genes. In addition, affine normalization makes the empirical distributions in different channels more equal, which is the purpose of quantile normalization, and may also explain why dye-swap normalization works or fails. All methods are made available in the aroma package, which is a platform-independent package for R.

Algorithms↗

Combined association and linkage analysis for general pedigrees and genetic models.

A combined score test for association and linkage analysis is introduced, based on a biologically plausible model with association between markers and causal genes and penetrance between phenotypes and the causal gene. The test is based on a retrospective likelihood of marker data given phenotypes, treating the alleles of the causal gene as hidden data. It is defined for arbitrary outbred pedigrees, a wide class of genetic models including polygenic and shared environmental effects and allows for missing marker data. It is multipoint, taking marker genotypes from several loci into account simultaneously. The score vector has one association and one linkage component, which can be used to define separate tests for association and linkage. For complete marker data, we give closed form expressions for the efficiency of the linkage, association and combined tests. These are examplified for binary and quantitative phenotypes with or without polygenic effects. The conclusion is that association tests are comparatively more efficient than linkage tests for strong association, weak penetrance models, small families and non-extreme phenotypes, whereas the linkage test is more efficient for weak association, strong penetrance models, large families and extreme phenotypes. The combined test is a robust alternative, which never performs much worse than the best of the linkage and association tests, and sometimes significantly better than both of them. It should be particularly useful when little is known about the genetic model.

Journal Article↗

Improving the calculation of statistical significance in genome-wide scans.

Calculations of the significance of results from linkage analysis can be performed by simulation or by theoretical approximation, with or without the assumption of perfect marker information. Here we concentrate on theoretical approximation. Our starting point is the asymptotic approximation formula presented by Lander and Kruglyak (1995, Nature Genetics, 11, 241--247), incorporating the effect of finite marker spacing as suggested by Feingold et al. (1993, American Journal of Human Genetics, 53, 234--251). We consider two distinct ways in which this formula can be improved. Firstly, we present a formula for calculating the crossover rate rho for a pedigree of general structure. For a pedigree set, these values may then be weighted into an overall crossover rate which can be used as input to the original approximation formula. Secondly, the unadjusted p -value formula is based on the assumption of a Normally distributed nonparametric linkage (NPL) score. This leads to conservative or anti-conservative p -values of varying magnitude depending on the pedigree set structure. We adjust for non-Normality by calculating the marginal distribution of the NPL score under the null hypothesis of no linkage with an arbitrarily small error. The NPL score is then transformed to have a marginal standard Normal distribution and the transformed maximal NPL score, together with a slightly corrected value of the overall crossover rate, is inserted into the original formula in order to calculate the p -value. We use pedigrees of seven different structures to compare the performance of our suggested approximation formula to the original approximation formula, with and without skewness correction, and to results found by simulation. We also apply the suggested formula to two real pedigree set structure examples. Our method generally seems to provide improved behavior, especially for pedigree sets which show clear departure from Normality, in relation to the competing approximations.

Computer Simulation↗

Conditional likelihood score functions for mixed models in linkage analysis.

In this paper, we develop a general strategy for linkage analysis, applicable for arbitrary pedigree structures and genetic models with one major gene, polygenes and shared environmental effects. Extending work of Whittemore (1996), McPeek (1999) and Hossjer (2003d), the efficient score statistic is computed from a conditional likelihood of marker data given phenotypes. The resulting semiparametric linkage analysis is very similar to nonparametric linkage based on affected individuals. The efficient score S depends not only on identical-by-descent sharing and phenotypes, but also on a few parameters chosen by the user. We focus on (1) weak penetrance models, where the major gene has a small effect and (2) rare disease models, where the major gene has a possibly strong effect but the disease causing allele is rare. We illustrate our results for a large class of genetic models with a multivariate Gaussian liability. This class incorporates one major gene, polygenes and shared environmental effects in the liability, and allows e.g. binary, Gaussian, Poisson distributed and life-length phenotypes. A detailed simulation study is conducted for Gaussian phenotypes. The performance of the two optimal score functions S(wpairs) and S(normdom) are investigated. The conclusion is that (i) inclusion of polygenic effects into the score function increases overall performance for a wide range of genetic models and (ii) score functions based on the rare disease assumption are slightly more powerful.

Alleles↗

Information and effective number of meioses in linkage analysis.

In this paper we introduce two information criteria in linkage analysis. The setup is a sample of families with unusually high occurrence of a certain inheritable disease. Given phenotypes from all families, the two criteria measure the amount of information inherent in the sample for 1) testing existence of a disease locus harbouring a disease gene somewhere along a chromosome or 2) estimating the position of the disease locus. Both criteria have natural interpretations in terms of effective number of meioses present in the sample. Thereby they generalize classical performance measures directly counting number of informative meioses. Our approach is conditional on observed phenotypes and we assume perfect marker data. We analyze two extreme cases of complete and weak penetrance models in particular detail. Some consequences of our work for sampling of pedigrees are discussed. For instance, a large sibship family with extreme phenotypes is very informative for linkage for weak penetrance models, more informative than a number of small families of the same total size.

Chromosome Mapping↗

Using importance sampling to improve simulation in linkage analysis.

In this article we describe and discuss implementation of a weighted simulation procedure, importance sampling, in the context of nonparametric linkage analysis. The objective is to estimate genome-wide p-values, i.e. the probability that the maximal linkage score exceeds given thresholds under the null hypothesis of no linkage. In order to reduce variance of the estimate for large thresholds, we simulate linkage scores under a distribution different from the null with an artificial disease locus positioned somewhere along the genome. To compensate for the fact that we simulate under the wrong distribution, the simulated scores are reweighted using a certain likelihood ratio. If the sampling distribution are properly chosen the variance of the corresponding estimate is reduced. This results in accurate genome-wide p-value estimates for a wide range of large thresholds with a substantially smaller cost adjusted relative efficiency with respect to standard unweighted simulation. We illustrate the performance of the method for several pedigree examples, discuss implementation including the amount of variance reduction and describe some possible generalizations.

Journal Article↗

On computation of p-values in parametric linkage analysis.

Parametric linkage analysis is usually used to find chromosomal regions linked to a disease (phenotype) that is described with a specific genetic model. This is done by investigating the relations between the disease and genetic markers, that is, well-characterized loci of known position with a clear Mendelian mode of inheritance. Assume we have found an interesting region on a chromosome that we suspect is linked to the disease. Then we want to test the hypothesis of no linkage versus the alternative one of linkage. As a measure we use the maximal lod score Z(max). It is well known that the maximal lod score has asymptotically a (2 ln 10)(-1) x (1/2 chi2(0) + 1/2 chi2(1)) distribution under the null hypothesis of no linkage when only one point (one marker) on the chromosome is studied. In this paper, we show, both by simulations and theoretical arguments, that the null hypothesis distribution of Zmax has no simple form when more than one marker is used (multipoint analysis). In fact, the distribution of Zmax depends on the number of families, their structure, the assumed genetic model, marker denseness, and marker informativity. This means that a constant critical limit of Zmax leads to tests associated with different significance levels. Because of the above-mentioned problems, from the statistical point of view the maximal lod score should be supplemented by a p-value when results are reported.

Chromosome Mapping↗

Assessing accuracy in linkage analysis by means of confidence regions.

When statistical linkage to a certain chromosomal region has been found, it is of interest to develop methods quantifying the accuracy with which the disease locus can be mapped. In this paper, we investigate the performance of three different types of confidence regions, with asymptotically correct coverage probability as the number of pedigrees grows. Our setup is that of a saturated map of marker data. We allow for arbitrary combinations of pedigree structures, and treat various kinds of genetic models (e.g. binary and quantitative phenotypes) in a unified way. The linkage scores are weighted sums of the individual family scores, with NPL and lod scores as special cases. We show that the expected length of the confidence region is inversely proportional to the slope-to-noise ratio, or equivalently, inversely proportional to the product of the square of the noncentrality parameter and a certain normalized slope-to-noise ratio. Our investigations reveal that maximal expected linkage scores can be quite different from estimation-based performance criteria based on expected length of confidence regions. The main reason is that there is no simple relationship between peak height and peak slope of the mean linkage score. One application of our results is planning of linkage studies: given a certain genetic model, we can approximate the number of pedigrees needed to obtain a confidence region with given coverage probability and expected length.

Chromosome Mapping↗