Search PubMed⌕ Search

Biomedical subjects

K Roeder

Publications and source records attributed to K Roeder.

26 records · Page 2Linked to original sources

Binning clones by hybridization with complex probes: statistical refinement of an inner product mapping method.

Molecular methods that use long-range information to solve genomics problems (i.e., top-down strategies) efficiently have become increasingly prominent in the genomics literature. One such method, an implementation of inner product mapping (IPM), uses noisy, long-range radiation hybrid (RH)/YAC overlap data and relatively noise-free RH/STS overlap data to localize clones to specific chromosomal regions. Because the molecular data are rarely noise-free, statistical models tailored to the top-down molecular methods make the methods far more effective. We develop two statistical models for IPM (or any other top-down strategy of similar form), a parametric logit model and a nonparametric order-restricted model, and show how these models can be implemented within a hierarchical Bayes framework. Using these models, we refine the chromosome 11 map reported in M. Perlin et al. (1995, Genomics 28: 315-327). Our analyses improve the IPM map, both in terms of successful localization of clones and in terms of the confidence with which they are localized.

Chromosome Mapping↗

Disequilibrium mapping: composite likelihood for pairwise disequilibrium.

The pattern of linkage disequilibrium between a disease locus and a set of marker loci has been shown to be a useful tool for geneticists searching for disease genes. Several methods have been advanced to utilize the pairwise disequilibrium between the disease locus and each of a set of marker loci. However, none of the methods take into account the information from all pairs simultaneously while also modeling the variability in the disequilibrium values due to the evolutionary dynamics of the population. We propose a Composite Likelihood (CL) model that has these features when the physical distances between the marker loci are known or can be approximated. In this instance, and assuming that there is a single disease mutation, the CL model depends on only three parameters, the recombination fraction between the disease locus and an arbitrary marker locus, theta, the age of the mutation, and a variance parameter. When the CL is maximized over a grid of theta, it provides a graph that can direct the search for the disease locus. We also show how the CL model can be generalized to account for multiple disease mutations. Evolutionary simulations demonstrate the power of the analyses, as well as their potential weaknesses. Finally, we analyze the data from two mapped diseases, cystic fibrosis and diastrophic dysplasia, finding that the CL method performs well in both cases.

Computer Simulation↗

Comments on the statistical aspects of the NRC's report on DNA typing.

The goal of the NRC report on DNA typing was to answer a "crescendo of questions concerning DNA typing," many of them in the areas of population genetics and statistics. Unfortunately, few of these questions were answered adequately. In lieu of answering these questions, the panel proposed another conservative method of forensic inference, the "ceiling principle." Aside from its extreme conservativeness, this new method is difficult to justify because it is based on inadequate population genetics and statistical theory. Moreover, in its ultimate implementation, the panel's method will depend on a population genetics study whose rationale is questionable. In this article, we elaborate some of the general comments we made about the NRC report in a recent article [1]. Specifically we cover three topics. First we question the statistical basis for the ceiling principle, showing that the empirical results that motivated the method are likely to be misinterpreted and showing, by power calculations, that the effects of population substructure cannot be substantial. Second, we show that the study design to determine "ceiling" allele frequencies has several undesirable statistical properties. Finally, we discuss the estimation of handling errors from the statistical perspective, a subject treated inadequately by the report.

DNA Fingerprinting↗

Estimation of allele frequencies for VNTR loci.

VNTR loci provide valuable information for a number of fields of study involving human genetics, ranging from forensics (DNA fingerprinting and paternity testing) to linkage analysis and population genetics. Alleles of a VNTR locus are simply fragments obtained from a particular portion of the DNA molecule and are defined in terms of their length. The essential element of a VNTR fragment is the repeat, which is a short sequence of basepairs. The core of the fragment is composed of a variable number of identical repeats that are linked in tandem. A sample of fragments from a population of individuals exhibits substantial variation in length because of variation in the number of repeats. Each distinct fragment length defines an allele, but any given fragment is measured with error. Therefore the observed distribution of fragment lengths is not discrete but is continuous, and determination of distinct allele classes is not straightforward. A mixture model is the natural statistical method for estimating the allele frequencies of VNTR loci. In this article we develop nonparametric methods for obtaining the distribution of allele sizes and estimates of their frequencies. Methods for obtaining maximum-likelihood estimates are developed. In addition, we suggest an empirical Bayes method to improve the maximum-likelihood estimates of the gene frequencies; the empirical Bayes procedure effects a local smoothing. The latter method works particularly well when measurement error is large relative to the repeat size, because the estimated distribution of allele frequencies when maximum likelihood is used is unreliable because of an alternating pattern of over- and underestimation. We define alleles and estimate the allele frequencies for two VNTR loci from the human genome (D17S79 and D2S44), from data obtained from Lifecodes, Inc.

Alleles↗

No excess of homozygosity at loci used for DNA fingerprinting.

Variable number of tandem repeat (VNTR) loci are extremely valuable for the forensic technique known as DNA fingerprinting because of their hypervariability. Nevertheless, the use of these loci in forensics has been controversial. One criticism of DNA fingerprinting is that the VNTR loci used for the "fingerprints" violate the assumption of Hardy-Weinberg equilibrium (H-W), making it difficult to calculate the probability of observing a genotype in the population. If one can assume H-W, the probability of observing the pair of alleles constituting an individual's genotype can be calculated by taking the product of the alleles' frequencies in the population and multiplying by two if the alleles are different. The evidence cited against assuming H-W is homozygote excess, which is presumed to be caused by an undetected mixture of two or more populations with limited interpopulational mating and distinct allele frequencies. For most VNTR loci, measurement error makes it impossible to test these claims by standard methods. The Lifecodes database of three VNTR loci used for forensics was used to show that the claimed excess of homozygotes is not necessarily real because many heterozygotes with similar allele sizes are misclassified as homozygotes. A simple test of H-W that takes such misclassifications into account was developed to test for an overall excess or dearth of heterozygotes in the sample (the complement of homozygote dearth or excess). The application of this test to the Lifecodes database revealed that there was no consistent evidence of violation of H-W for the Caucasian, black, or Hispanic populations.

DNA↗