Search PubMed⌕ Search

Biomedical subjects

Mikkel H Schierup

Publications and source records attributed to Mikkel H Schierup.

14 recordsLinked to original sources

Whole genome association mapping by incompatibilities and local perfect phylogenies.

BACKGROUND: With current technology, vast amounts of data can be cheaply and efficiently produced in association studies, and to prevent data analysis to become the bottleneck of studies, fast and efficient analysis methods that scale to such data set sizes must be developed. RESULTS: We present a fast method for accurate localisation of disease causing variants in high density case-control association mapping experiments with large numbers of cases and controls. The method searches for significant clustering of case chromosomes in the "perfect" phylogenetic tree defined by the largest region around each marker that is compatible with a single phylogenetic tree. This perfect phylogenetic tree is treated as a decision tree for determining disease status, and scored by its accuracy as a decision tree. The rationale for this is that the perfect phylogeny near a disease affecting mutation should provide more information about the affected/unaffected classification than random trees. If regions of compatibility contain few markers, due to e.g. large marker spacing, the algorithm can allow the inclusion of incompatibility markers in order to enlarge the regions prior to estimating their phylogeny. Haplotype data and phased genotype data can be analysed. The power and efficiency of the method is investigated on 1) simulated genotype data under different models of disease determination 2) artificial data sets created from the HapMap ressource, and 3) data sets used for testing of other methods in order to compare with these. Our method has the same accuracy as single marker association (SMA) in the simplest case of a single disease causing mutation and a constant recombination rate. However, when it comes to more complex scenarios of mutation heterogeneity and more complex haplotype structure such as found in the HapMap data our method outperforms SMA as well as other fast, data mining approaches such as HapMiner and Haplotype Pattern Mining (HPM) despite being significantly faster. For unphased genotype data, an initial step of estimating the phase only slightly decreases the power of the method. The method was also found to accurately localise the known susceptibility variants in an empirical data set--the DeltaF508 mutation for cystic fibrosis--where the susceptibility variant is already known--and to find significant signals for association between the CYP2D6 gene and poor drug metabolism, although for this dataset the highest association score is about 60 kb from the CYP2D6 gene. CONCLUSION: Our method has been implemented in the Blossoc (BLOck aSSOCiation) software. Using Blossoc, genome wide chip-based surveys of 3 million SNPs in 1000 cases and 1000 controls can be analysed in less than two CPU hours.

Chromosome Mapping↗

The transition to self-compatibility in Arabidopsis thaliana and evolution within S-haplotypes over 10 Myr.

A recent investigation found evidence that the transition of Arabidopsis thaliana from ancestral self-incompatibility (SI) to full self-compatibility occurred very recently and suggested that this occurred through a selective fixation of a nonfunctional allele (PsiSCR1) at the SCR gene, which determines pollen specificity in the incompatibility response. The main evidence is the lack of polymorphism at the SCR locus in A. thaliana. However, the nearby SRK gene, which determines stigma specificity in self-incompatible Brassicaceae species, has extremely high sequence diversity, with 3 very divergent SRK haplotypes, 2 of them present in multiple strains. Such high diversity is extremely unusual in this species, and it suggests the possibility that multiple, different SRK haplotypes may have been preserved from A. thaliana's self-incompatible ancestor. To study the evolution of S-haplotypes in the A. thaliana lineage, we searched the 2 most closely related Arabidopsis species Arabidopsis lyrata and Arabidopsis halleri, in which most populations have retained SI, and found SRK sequences corresponding to all 3 A. thaliana haplogroup sequences. Our molecular evolutionary analyses of these 3 S-haplotypes provide an independent estimate of the timing of the breakdown of SI and again exclude an ancient transition to selfing in A. thaliana. Comparing sequences of each of the 3 haplogroups between species, we find that 2 of the 3 SRK sequences (haplogroups A and B) are similar throughout their length, suggesting that little or no recombination with other SRK alleles has occurred since these species diverged. The diversity difference between the SCR and SRK loci in A. thaliana, however, suggests crossing-over, either within SRK or between the SCR and SRK loci. If the loss of SI involved fixation of the PsiSCR1 sequence, the exchange must have occurred during its fixation. Divergence between the species is much lower at the S-locus, compared with reference loci, and we discuss two contributory possibilities. Introgression may have occurred between A. lyrata and A. halleri and between their ancestral lineage and A. thaliana, at least for some period after their split. In addition, the coalescence times of sequences of individual S-haplogroups are expected to be less than those of alleles at non-S-loci.

Arabidopsis↗

The effective size of the Icelandic population and the prospects for LD mapping: inference from unphased microsatellite markers.

Characterizing the extent of linkage disequilibrium (LD) in the genome is a pre-requisite for association mapping studies. Patterns of LD also contain information about the past demography of populations. In this study, we focus on the Icelandic population where LD was investigated in 12 regions of approximately 15 cM using regularly spaced microsatellite loci displaying high heterozygosity. A total of 1753 individuals were genotyped for 179 markers. LD was estimated using a composite disequilibrium measure based on unphased data. LD decreases with distance in all 12 regions and more LD than expected by chance can be detected over approximately 4 cM in our sample. Differences in the patterns of decrease of LD with distance among genomic regions were mostly due to two regions exhibiting, respectively, higher and lower proportions of pairs in LD than average within the first 4 cM. We pooled data from all regions, except these two and summarized patterns of LD by computing the proportion of pairs of loci exhibiting significant LD (at the 5% level) as a function of distance. We compared observed patterns of LD with simulated data sets obtained under scenarios with varying demography and intensity of recombination. We show that unphased data allow to make inferences on scaled recombination rates from patterns of LD. Patterns of LD in Iceland suggest a genome-wide scaled recombination rate of rho* = 200 (130-330) per cM (or an effective size of roughly 5000), in the low range of estimates recently reported in three populations from the HapMap project.

Biological Evolution↗

GeneRecon--a coalescent based tool for fine-scale association mapping.

UNLABELLED: GeneRecon is a tool for fine-scale association mapping using a coalescence model. GeneRecon takes as input case-control data from phased or unphased SNP and microsatellite genotypes. The posterior distribution of disease locus position is obtained by Metropolis-Hastings sampling in the state space of genealogies. Input format, search strategy and the sampled statistics can be configured through the Guile Scheme programming language embedded in GeneRecon, making GeneRecon highly configurable. AVAILABILITY: The source code for GeneRecon, written in C++ and Scheme, is available under the GNU General Public License (GPL) at http://www.birc.au.dk/~mailund/GeneRecon CONTACT: mailund@birc.au.dk.

Algorithms↗

Long-term stability and effective population size in North Sea and Baltic Sea cod (Gadus morhua).

DNA from archived otoliths was used to explore the temporal stability of the genetic composition of two cod populations, the Moray Firth (North Sea) sampled in 1965 and 2002, and the Bornholm Basin (Baltic Sea) sampled in 1928 and 1997. We found no significant changes in the allele frequencies for the Moray Firth population, while subtle but significant genetic changes over time were detected for the Bornholm Basin population. Estimates of the effective population size (Ne) generally exceeded 500 for both populations when employing a number of varieties of the temporal genetic method. However, confidence intervals were very wide and Ne's most likely range in the thousands. There was no apparent loss of genetic variability and no evidence of a genetic bottleneck for either of the populations. Calculations of the expected levels of genetic variability under different scenarios of Ne showed that the number of alleles commonly reported at microsatellite loci in Atlantic cod is best explained by Ne's exceeding thousand. Recent fishery-induced bottlenecks can, however, not be ruled out as an explanation for the apparent discrepancy between high levels of variability and recently reported estimates of Ne << 1000. From life history traits and estimates of survival rates in the wild, we evaluate the compatibility of the species' biology and extremely low Ne/N ratios. Our data suggest that very small Ne's are not likely to be of general concern for cod populations and, accordingly, most populations do not face any severe threat of losing evolutionary potential due to genetic drift.

Animals↗

CoaSim: a flexible environment for simulating genetic data under coalescent models.

BACKGROUND: Coalescent simulations are playing a large role in interpreting large scale intra-specific sequence or polymorphism surveys and for planning and evaluating association studies. Coalescent simulations of data sets under different models can be compared to the actual data to test the importance of different evolutionary factors and thus get insight into these. RESULTS: We have created the CoaSim application as a flexible environment for Monte Carlo simulation of various types of genetic data under equilibrium and non-equilibrium coalescent processes for a variety of applications. Interaction with the tool is through the Guile version of the Scheme scripting language. Scheme scripts for many standard and advanced applications are provided and these can easily be modified by the user for a much wider range of applications. A graphical user interface with less functionality and flexibility is also included. It is primarily intended as an exploratory and educational tool CONCLUSION: CoaSim is a powerful tool because of its flexibility and ease of use. This is illustrated through very varied uses of the application, e.g. evaluation of association mapping methods, parametric bootstrapping, and design and choice of markers for specific questions.

Case-Control Studies↗

Evidence of recombination among early-vaccination era measles virus strains.

BACKGROUND: The advent of live-attenuated vaccines against measles virus during the 1960'ies changed the circulation dynamics of the virus. Earlier the virus was indigenous to countries worldwide, but now it is mediated by a limited number of evolutionary lineages causing sporadic outbreaks/epidemics of measles or circulating in geographically restricted endemic areas of Africa, Asia and Europe. We expect that the evolutionary dynamics of measles virus has changed from a situation where a variety of genomic variants co-circulates in an epidemic with relatively high probabilities of co-infection of the individual to a situation where a co-infection with strains from evolutionary different lineages is unlikely. RESULTS: We performed an analysis of the partial sequences of the hemagglutinin gene of 18 measles virus strains collected in Denmark between 1965 and 1983 where vaccination was first initiated in 1987. The results were compared with those obtained with strains collected from other parts of the world after the initiation of vaccination in the given place. Intergenomic recombination among pre-/early-vaccination strains is suggested by 1) estimations of linkage disequilibrium between informative sites, 2) the decay of linkage disequilibrium with distance between informative sites and 3) a comparison of the expected number of homoplasies to the number of apparent homoplasies in the most parsimonious tree. No significant evidence of recombination could be demonstrated among strains circulating at present. CONCLUSION: We provide evidence that recombination can occur in measles virus and that it has had a detectable impact on sequence evolution of pre-vaccination samples. We were not able to detect recombination from present-day sequence surveys. We believe that the decreased rate of visible recombination may be explained by changed dynamics, since divergent strains do not meet very often in current epidemics that are often spawned by a single sequence type. Signs of pre-vaccination recombination events in the present-day sequences are not strong enough to be detectable.

Base Sequence↗

Selection at work in self-incompatible Arabidopsis lyrata: mating patterns in a natural population.

Identification and characterization of the self-incompatibility genes in Brassicaceae species now allow typing of self-incompatibility haplotypes in natural populations. In this study we sampled and mapped all 88 individuals in a small population of Arabidopsis lyrata from Iceland. The self-incompatibility haplotypes at the SRK gene were typed for all the plants and some of their progeny and used to investigate the realized mating patterns in the population. The observed frequencies of haplotypes were found to change considerably from the parent generation to the offspring generation around their deterministic equilibria as determined from the known dominance relations among haplotypes. We provide direct evidence that the incompatibility system discriminates against matings among adjacent individuals. Multiple paternity is very common, causing mate availability among progeny of a single mother to be much larger than expected for single paternity.

Arabidopsis↗

Pigs in sequence space: a 0.66X coverage pig genome survey based on shotgun sequencing.

BACKGROUND: Comparative whole genome analysis of Mammalia can benefit from the addition of more species. The pig is an obvious choice due to its economic and medical importance as well as its evolutionary position in the artiodactyls. RESULTS: We have generated approximately 3.84 million shotgun sequences (0.66X coverage) from the pig genome. The data are hereby released (NCBI Trace repository with center name "SDJVP", and project name "Sino-Danish Pig Genome Project") together with an initial evolutionary analysis. The non-repetitive fraction of the sequences was aligned to the UCSC human-mouse alignment and the resulting three-species alignments were annotated using the human genome annotation. Ultra-conserved elements and miRNAs were identified. The results show that for each of these types of orthologous data, pig is much closer to human than mouse is. Purifying selection has been more efficient in pig compared to human, but not as efficient as in mouse, and pig seems to have an isochore structure most similar to the structure in human. CONCLUSION: The addition of the pig to the set of species sequenced at low coverage adds to the understanding of selective pressures that have acted on the human genome by bisecting the evolutionary branch between human and mouse with the mouse branch being approximately 3 times as long as the human branch. Additionally, the joint alignment of the shot-gun sequences to the human-mouse alignment offers the investigator a rapid way to defining specific regions for analysis and resequencing.

Animals↗

Haplotype structure of the stigmatic self-incompatibility gene in natural populations of Arabidopsis lyrata.

We describe analyses of almost full-length sequences (including both the kinase domain and the S-domain) of the putative SRK incompatibility gene of the self-incompatible plant Arabidopsis lyrata. In A. lyrata, the SRK S-domain controls the pistil recognition specificity, as in self-incompatible Brassica species. In alleles from plants derived from natural A. lyrata populations, nonsynonymous and synonymous site diversity values are very high in both domains; even in exons 3 to 7 of the kinase domain, which probably have no recognition functions, 39% of the amino acids are polymorphic. Within populations, diversity between alleles is high, as expected for an incompatibility locus, which should be under frequency-dependent selection within populations, whereas within the different putative allelic classes polymorphism is very low, as predicted from theoretical models when recombination is rare. Nonsynonymous site variability declines in the kinase domain with increasing distance from the S-domain border, although synonymous diversity remains high, and the introns are unalignable. A decline in nonsynonymous diversity is expected due to selective constraints in the kinase domain, in combination with recombination (allowing diversity to decrease at sites distant from those under balancing selection). However, it is unclear whether recombination occurs in the SRK locus, and interpretation of the observed diversity pattern is complicated by apparent gene conversion with a paralogous gene (or genes). Patterns of linkage disequilibrium in our SRK sequences do not support the conclusion that recombination occurs, which was suggested from previous analyses based on Brassica SLG sequences.

Alleles↗

Relative roles of mutation and recombination in generating allelic polymorphism at an MHC class II locus in Peromyscus maniculatus.

The MHC class II loci encoding cell surface antigens exhibit extremely high allelic polymorphism. There is considerable uncertainty in the literature over the relative roles of recombination and de novo mutation in generating this diversity. We studied class II sequence diversity and allelic polymorphism in two populations of Peromyscus maniculatus, which are among the most widespread and abundant mammals of North America. We find that intragenic recombination (or gene conversion) has been the predominant mode for the generation of allelic polymorphism in this species, with the amount of population recombination per base pair exceeding mutation by at least an order of magnitude during the history of the sample. Despite this, patchwork motifs of sites with high linkage disequilibrium are observed. This does not appear to be consistent with the much larger amount of recombination versus mutation in the history of the sample, unless the recombination rate is highly non-uniform over the sequence or selection maintains certain sites in linkage disequilibrium. We conclude that selection is most likely to be responsible for preserving sequence motifs in the presence of abundant recombination.

Alleles↗

Diversity and linkage of genes in the self-incompatibility gene family in Arabidopsis lyrata.

We report studies of seven members of the S-domain gene family in Arabidopsis lyrata, a member of the Brassicaceae that has a sporophytic self-incompatibility (SI) system. Orthologs for five loci are identifiable in the self-compatible relative A. thaliana. Like the Brassica stigmatic incompatibility protein locus (SRK), some of these genes have kinase domains. We show that several of these genes are unlinked to the putative A. lyrata SRK, Aly13. These genes have much lower nonsynonymous and synonymous polymorphism than Aly13 in the S-domains within natural populations, and differentiation between populations is higher, consistent with balancing selection at the Aly13 locus. One gene (Aly8) is linked to Aly13 and has high diversity. No departures from neutrality were detected for any of the loci. Comparing different loci within A. lyrata, sites corresponding to hypervariable regions in the Brassica S-loci (SLG and SRK) and in comparable regions of Aly13 have greater replacement site divergence than the rest of the S-domain. This suggests that the high polymorphism in these regions of incompatibility loci is due to balancing selection acting on sites within or near these regions, combined with low selective constraints.

Alleles↗

Sequence analysis of measles virus strains collected during the pre- and early-vaccination era in Denmark reveals a considerable diversity of ancient strains.

A total of 199 serum samples from patients with measles collected in Denmark, Greenland and the Faroe Islands from 1964 to 1983 were analysed by PCR. Measles virus (MV) RNA could be detected in 38 (19%) of the samples and a total of 18 strains were subjected to partial sequence analysis of the hemagglutinin gene. The strains exhibited a considerable genomic diversity, which is at odds with the assumption that one genome type prevailed among globally circulating MV strains prior to the advent of live-attenuated vaccines. Our data indicate that the similarity of the various vaccine strains is attributed to their having originated from the same primary isolate. Consequently, it is implied that a small number of clinical manifestations of MV worldwide from which strains similar to the vaccine strain were identified were vaccine related rather than being caused by members of a persistently circulating ancient genome type. The Danish pre- and early-vaccination era MV strains seem to change the evolutionary spectrum of genome types A, C2 and E into one coherent group, suggesting that the genome types of MV strains circulating in the world at present do not represent far ranging evolutionary lineages but merely members of an evolutionary continuum of pre-vaccination era MV strains which by chance or due to an improved capability survived the worldwide partial herd immunity accomplished through vaccination.

Child↗