Search PubMed⌕ Search

Biomedical subjects

Kenneth Lange

Publications and source records attributed to Kenneth Lange.

15 recordsLinked to original sources

Accommodating chromosome inversions in linkage analysis.

This work develops a population-genetics model for polymorphic chromosome inversions. The model precisely describes how an inversion changes the nature of and approach to linkage equilibrium. The work also describes algorithms and software for allele-frequency estimation and linkage analysis in the presence of an inversion. The linkage algorithms implemented in the software package Mendel estimate recombination parameters and calculate the posterior probability that each pedigree member carries the inversion. Application of Mendel to eight Centre d'Etude du Polymorphisme Humain pedigrees in a region containing a common inversion on 8p23 illustrates its potential for providing more-precise estimates of the location of an unmapped marker or trait gene. Our expanded cytogenetic analysis of these families further identifies inversion carriers and increases the evidence of linkage.

Algorithms↗

Variance component models for X-linked QTLs.

This paper discusses the theory and implementation of a model for mapping X-linked quantitative trait loci (QTL). As a result of X inactivation, a female's body is subdivided into a number of patches. In each patch one of her two X chromosomes is randomly switched off. This smooths the allelic contributions in a heterozygote and implies that females should show less trait variation than males for an X-linked trait. The latest version of the genetic analysis program Mendel incorporates a simple variance component version of this model. An application to head circumference in autistic children illustrates Mendel in action.

Analysis of Variance↗

Reconstructing ancestral haplotypes with a dictionary model.

We propose a dictionary model for haplotypes. According to the model, a haplotype is constructed by randomly concatenating haplotype segments from a given dictionary of segments. A haplotype block is defined as a set of haplotype segments that begin and end with the same pair of markers. In this framework, haplotype blocks can overlap, and the model provides a setting for testing the accuracy of simpler models invoking only nonoverlapping blocks. Each haplotype segment in a dictionary has an assigned probability and alternate spellings that account for genotyping errors and mutation. The model also allows for missing data, unphased genotypes, and prior distribution of parameters. Likelihood evaluations rely on forward and backward recurrences similar to the ones encountered in hidden Markov models. Parameter estimation is carried out with an EM algorithm. The search for the optimal dictionary is particularly difficult because of the variable dimension of the model space. We define a minimum description length criteria to evaluate each dictionary and use a combination of greedy search and careful initialization to select a best dictionary for a given dataset. Application of the model to simulated data gives encouraging results. In a real dataset, we are able to reconstruct a parsimonious dictionary that captures patterns of linkage disequilibrium well.

Algorithms↗

Strong correlation between meiotic crossovers and haplotype structure in a 2.5-Mb region on the long arm of chromosome 21.

Although the haplotype structure of the human genome has been studied in great detail, very little is known about the mechanisms underlying its formation. To investigate the role of meiotic recombination on haplotype block formation, single nucleotide polymorphisms were selected at a high density from a 2.5-Mb region of human chromosome 21. Direct analysis of meiotic recombination by high-throughput multiplex genotyping of 662 single sperm identifies 41 recombinants. The crossovers were nonrandomly distributed within 16 small areas. All, except one, of these crossovers fall in areas where the haplotype structure exhibits breakdown, displaying a strong statistically positive association between crossovers and haplotype block breaks. The data also indicate a particular clustered distribution of recombination hotspots within the region. This finding supports the hypothesis that meiotic recombination makes a primary contribution to haplotype block formation in the human genome.

Chromosome Mapping↗

Association testing in a linked region using large pedigrees.

This report describes computer implementation of a scheme for joint linkage and association analysis. The model implemented in the computer package Mendel estimates both recombination and linkage-disequilibrium parameters and conducts likelihood-ratio tests for (1) linkage alone, (2) linkage and association simultaneously, and (3) association in the presence of linkage. Application of the method to data from Finnish pedigrees with familial combined hyperlipidemia illustrates its potential for identification of associated SNP haplotypes in the presence of linkage. For the test results to be valid, good estimates of haplotype frequencies must be used in the analysis.

Female↗

Association testing with Mendel.

This report presents an overview of association testing strategies from a user's perspective, with particular attention to the capabilities of the computer program Mendel. Association testing is driven by the nature of the study sample, the nature of the disease trait, and the kind of markers employed. The practicing statistician must also choose whether to conduct parametric or nonparametric tests. Because of the complexities involved, Mendel offers users several analysis options. The different options are tied together by shared input and output conventions and a shared language for defining models. Mendel also features new statistics and theory found in no other genetics software. The most important innovations include: association testing by penetrance estimation, expansion of matched-pair designs to permutation unit designs, and a rigorous implementation of the measured genotype approach for quantitative trait loci. This report explains how Mendel imputes allele counts and conducts both asymptotic and permutation tests in the measured genotype framework.

Analysis of Variance↗

Vocabulon: a dictionary model approach for reconstruction and localization of transcription factor binding sites.

MOTIVATION: Gene expression arrays enable measurements of transcription values for a large number or all genes in the genome. In order to better interpret these results and to use them to reconstruct transcription networks, information on location of binding sites for regulatory proteins in the entire genome is needed. In particular, this represents an open problem in Escherichia coli. RESULTS: We describe the first implementation of dictionary-style models to the study of transcription factors binding sites in an entire genome. Vocabulon's unique feature is that it can both reconstruct binding sites characterized by unknown motifs and impute locations of known binding sites in long sequences by simultaneous search. On one hand, the dictionary model specifies a probability for the entire sequence taking simultaneously into account all the possible binding sites. This greatly reduces the number of false positives. On the other hand, the possibility of refining motif description, as an increasing number of binding sites are identified, augments the sensitivity of the method. We illustrate these properties with examples in E.coli. The results of gene expression arrays are used both to guide the search and corroborate it.

Algorithms↗

Locus for quantitative HDL-cholesterol on chromosome 10q in Finnish families with dyslipidemia.

Decreased HDL-cholesterol (HDL-C) and familial combined hyperlipidemia (FCHL) are the two most common familial dyslipidemias predisposing to premature coronary heart disease (CHD). These dyslipidemias share many phenotypic features, suggesting a partially overlapping molecular pathogenesis. This was supported by our previous pooled data analysis of the genome scans for low HDL-C and FCHL, which identified three shared chromosomal regions for a qualitative HDL-C trait on 8q23.1, 16q23.3, and 20q13.32. This study further investigates these regions as well as two other loci we identified earlier for premature CHD on 2q31 and Xq24 and a locus for high serum triglycerides (TGs) on 10q11. We analyzed 67 microsatellite markers in an extended study sample of 1,109 individuals from 92 low HDL-C or FCHL families using both qualitative and quantitative lipid phenotypes. These analyses provided evidence for linkage (a logarithm of odds score of 3.2) on 10q11 using a quantitative HDL-C trait. Importantly, this region, previously linked to TGs, body mass index, and obesity, provided evidence for association for quantitative TGs (P = 0.0006) and for a combined trait of HDL-C and TGs (P = 0.008) with marker D10S546. Suggestive evidence for linkage also emerged for HDL-C on 2q31 and for TGs on 20q13.32. Finnish families ascertained for dyslipidemias thus suggest that 10q11, 2q31, and 20q13.32 harbor loci for HDL-C and TGs.

Adult↗

Powerful allele sharing statistics for nonparametric linkage analysis.

Nonparametric linkage analysis is widely used to map susceptibility genes for complex diseases. This paper introduces six nonparametric statistics for measuring marker allele sharing among the affected members of a pedigree. We compare the power of these new statistics and three previous statistics to detect linkage with Mendelian diseases having recessive, additive, and dominant modes of inheritance. The nine statistics represent all possible combinations of three different IBD scoring functions and three different schemes for sampling genes among affecteds. Our results strongly suggest that the statistic T(rec)(blocks) is best for recessive traits, while the two statistics T(kin)(pairs) and T(all)(kin) vie for best for an additive trait. The best statistic for a dominant trait is less clear. The statistics T(kin)(pairs) and T(all)(kin) are equally promising for small sibships, but in extended pedigrees the statistics T(dom)(blocks) and T(dom)(pairs) appear best. For a complex trait, we advocate computing several of these statistics.

Alleles↗

The pedigree trimming problem.

This report mathematically validates a fast algorithm for trimming irrelevant members from a pedigree. These individuals are typically dead or otherwise unavailable for study. Left in the pedigree, they slow likelihood evaluation. Of course, each pedigree of interest will have some core people who must be retained. The described algorithm retains just enough of the pedigree to maintain the proper relationships among the core people.

Algorithms↗

Detection and integration of genotyping errors in statistical genetics.

Detection of genotyping errors and integration of such errors in statistical analysis are relatively neglected topics, given their importance in gene mapping. A few inopportunely placed errors, if ignored, can tremendously affect evidence for linkage. The present study takes a fresh look at the calculation of pedigree likelihoods in the presence of genotyping error. To accommodate genotyping error, we present extensions to the Lander-Green-Kruglyak deterministic algorithm for small pedigrees and to the Markov-chain Monte Carlo stochastic algorithm for large pedigrees. These extensions can accommodate a variety of error models and refrain from simplifying assumptions, such as allowing, at most, one error per pedigree. In principle, almost any statistical genetic analysis can be performed taking errors into account, without actually correcting or deleting suspect genotypes. Three examples illustrate the possibilities. These examples make use of the full pedigree data, multiple linked markers, and a prior error model. The first example is the estimation of genotyping error rates from pedigree data. The second-and currently most useful-example is the computation of posterior mistyping probabilities. These probabilities cover both Mendelian-consistent and Mendelian-inconsistent errors. The third example is the selection of the true pedigree structure connecting a group of people from among several competing pedigree structures. Paternity testing and twin zygosity testing are typical applications.

Algorithms↗

Spline methods for the comparison of physical and genetic maps.

The first genetic maps were constructed by linkage analysis. Physical mapping techniques, such as radiation hybrids and complete sequencing, produce a different picture. For the purposes of population genetics, clinical genetics, and genetic epidemiology, it is important to harmonize and amalgamate existing genetic and physical maps. Among other things, comparisons of the two kinds of maps promotes better understanding of the wide variation in local recombination rates per unit physical length of DNA. The current paper presents methods for estimating recombination intensity as a function of physical distance along a chromosome. Genetic map distance is the integral of intensity. We derive fast reliable estimation algorithms based on a Poisson process model, penalized likelihoods, and cubic spline interpolation. Our methods provide a rigorous and statistically sound foundation for comparing physical and genetic maps. To illustrate the possibilities, we apply the methods to published recombination data on CEPH families and the complete sequences of chromosomes 21 and 22. Our results are in good agreement with previous studies and the biological data.

Algorithms↗

Codon and rate variation models in molecular phylogeny.

This article generalizes previous models for codon substitution and rate variation in molecular phylogeny. Particular attention is paid to (1) reversibility, (2) acceptance and rejection of proposed codon changes, (3) varying rates of evolution among codon sites, and (4) the interaction of these sites in determining evolutionary rates. To accommodate spatial variation in rates, Markov random fields rather than Markov chains are introduced. Because these innovations complicate maximum likelihood estimation in phylogeny reconstruction, it is necessary to formulate new algorithms for the evaluation of the likelihood and its derivatives with respect to the underlying kinetic, acceptance, and spatial parameters. To derive the most from maximum likelihood analysis of sequence data, it is useful to compute posterior probabilities assigning residues to internal nodes and evolutionary rate classes to codon sites. It is also helpful to search through tree space in a way that respects accepted phylogenetic relationships. Our phylogeny program LINNAEUS implements algorithms realizing these goals. Readers may consult our companion article in this issue for several examples.

Algorithms↗

Applications of codon and rate variation models in molecular phylogeny.

The current article illustrates the practical advantages of some new models and statistical algorithms for codon substitution and spatial rate variation in molecular phylogeny. Our companion paper in this issue discusses at length the mathematical properties of these models for nucleotide and codon substitution, for site-to-site and branch-to-branch heterogeneity in rates of evolution, and for spatial correlation in the assignment of rates. In this study we summarize the theoretical background and apply the models and algorithms to data on beta-globin, the complete HIV genome, and the mitochondrial genome. Our complex but realistic models enhance biological interpretation of sequence data and show substantial improvements in model fit over existing models. All the new statistical algorithms applied are incorporated in our phylogeny software LINNAEUS, which is tuned for performance and modeling flexibility.

Animals↗

Merging microsatellite data.

Genotype calling procedures vary from laboratory to laboratory for many microsatellite markers. Even within the same laboratory, application of different experimental protocols often leads to ambiguities. The impact of these ambiguities ranges from irksome to devastating. Resolving the ambiguities can increase effective sample size and preserve evidence in favor of disease-marker associations. Because different data sets may contain different numbers of alleles, merging is unfortunately not a simple process of matching alleles one to one. Merging data sets manually is difficult, time-consuming, and error-prone due to differences in genotyping hardware, binning methods, molecular weight standards, and curve fitting algorithms. Merging is particularly difficult if few or no samples occur in common, or if samples are drawn from ethnic groups with widely varying allele frequencies. It is dangerous to align alleles simply by adding a constant number of base pairs to the alleles of one of the data sets. To address these issues, we have developed a Bayesian model and a Markov chain Monte Carlo (MCMC) algorithm for sampling the posterior distribution under the model. Our computer program, MicroMerge, implements the algorithm and almost always accurately and efficiently finds the most likely correct alignment. Common allele frequencies across laboratories in the same ethnic group are the single most important cue in the model. MicroMerge computes the allelic alignments with the greatest posterior probabilities under several merging options. It also reports when data sets cannot be confidently merged. These features are emphasized in our analysis of simulated and real data.

Algorithms↗