A likelihood-based analysis of consistent linkage of a disease locus to two nonsyntenic marker loci: osteogenesis imperfecta versus COL1A1 and COL1A2.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to D E Weeks.
Explore the source record for details and available documents.
The EM algorithm is an iterative method for finding maximum-likelihood estimates. Its advantages often include numerical stability, simplicity of computer implementation, and natural incorporation of parameter constraints. However, the EM algorithm must be tailored to each specific problem. Smith (1957) and Ott (1977, 1979) have accomplished this for a variety of problems in human pedigree analysis. The present paper clarifies their theory by presenting it from a modern perspective. Five practical numerical examples are also given in an attempt to assess the value of the EM algorithm in realistic genetic modelling. These examples deal with racial admixture, linkage homogeneity, classical segregation analysis, a Mendelian latent trait model for schizophrenia, and a heterozygote detection assay for Ataxia-telangiectasia. Comparison with a quasi-Newton method of optimization reveals that the EM algorithm generally converges more slowly, but also more stably.
Calculation of multilocus lod scores presents challenging problems in numerical analysis, combinatories, programming, and genetics. It is possible to accelerate these computations by exploiting the simple pedigree structure of a CEPH-type pedigree consisting of a nuclear family plus all four grandparents. Lathrop et al. (1986) have done this by introducing likelihood factorization and transformation rules and Lander & Green (1987) by the method of 'hidden Markov chains'. The present paper explores an alternative approach based on genotype redefinition in the grandparents and systematic phase elimination in all pedigree members. All three approaches accelerate the computation of a single likelihood. Equally relevant to multilocus mapping are search strategies for finding the maximum likelihood estimates of recombination fractions. Hybrid algorithms that start with the EM algorithm and switch midway to quasi-Newton algorithms show promise. These issues are investigated in the context of a simulated 10 locus example. This same example allows us to illustrate a simple strategy for determining locus order.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This paper describes a generalization of the affected-sib-pair method of linkage analysis to pedigrees. By substituting identity-by-state relations for identity-by-descent relations, we develop a test statistic for detecting departures from independent segregation of disease and marker phenotypes. The statistic is based on the marker phenotypes of affected pedigree members only. Since it is more striking for distantly affected relatives to share a rare marker allele than a common marker allele, the statistic also includes a weighting factor based on allele frequency. The distributional properties of the statistic are investigated theoretically and by simulation. Part of the theoretical treatment entails generalizing Karigl's multiple-person kinship coefficients. When the test statistic is applied to pedigree data on Huntington disease, the null hypothesis of independent segregation between the marker locus and the disease locus is firmly rejected. In this case, as expected, there is a loss of power when compared with standard lod-score analysis. However, our statistic possesses the advantage of requiring no explicit assumptions about the mode of inheritance of the disease. This point is illustrated by application of the test statistic to data on rheumatoid arthritis.
N linked loci can be arranged in N!/2 possible orders. We describe two criteria for providing a preliminary ranking of the possible orders based on the N(N-1)/2 pairwise lod score curves for the loci. For a given order the first criterion is the sum of the N-1 maximal lod scores corresponding to the adjacent pairs of loci in the order. The second criterion is the minimum of a least-squares problem due to J.M. Lalouel (1977, Heredity 38(1): 61-77). This least-squares problem requires the maximum likelihood recombination fraction estimates and their standard errors. For N small it is feasible to evaluate these measures for every possible order. For N large we use a simulated annealing algorithm. This gives a fairly complete listing of the best-candidate orders without sampling every possible order. These ranking methods are applied to data from linkage groups on chromosomes 1, 6, 11, and 13.
We have tested thirty-two phenotypic blood markers on sixteen families with with ataxia-telangiectasia (AT) in an attempt to identify the chromosomal location of the AT gene(s). Although at least five complementation groups have been defined, it is not known whether the corresponding AT genes are clustered or dispersed in the genome. Both clustered and dispersed genetic models were considered in linkage analyses. No significant linkages were found. The data exclude approximately 7 per cent of the autosomal genome for a 'clustered' model and 2 per cent of the autosomal genome for a 'dispersed' model. Several genomic areas were identified which warrant further study.
Ataxia-telangiectasia (AT) is a multifaceted autosomal recessive disorder, inherited as a single gene in each family, presumably due to a defective DNA processing protein such as a recombinase, endonuclease or even a regulatory DNA-binding protein. We are attempting to identify the chromosomal location of the AT gene(s) by performing linkage analyses on a variety of genetic models. At least five AT complementation groups have been defined. This genetic heterogeneity complicates linkage analysis. Model I assumes that the complementation genes are clustered into a single genomic region and, therefore, lod scores of linkage data from all families can be added. Model II assumes that the AT complementation genes are dispersed throughout the genome and the lod scores cannot be added. This model necessitates assigning the complementation group of every family that is included in the linkage analyses and reduces the number of families in each data base. Model III utilizes heterozygote identification to follow the AT gene (in a Group A pedigree of 61 members) as a dominant trait, thereby increasing the amount of linkage information that can be derived from that family. Model IV will focus only on consanguineous offspring of first-cousin marriages, seeking to identify the location of the AT gene(s) by the increased degree of homozygosity of genetic markers in close proximity. This model has several advantages, including that much smaller numbers of patients are required. Model V assumes that a subset of our patients will carry deletions and can be used to confirm the relationship of a candidate gene to the AT phenotype. Progress: Models I and II have been used to survey 7% and 2% of the genome, respectively. (An additional 5% of the genome can be added for exclusion of the X chromosome on clinical grounds). Model III is intended to survey the entire genome. Our initial studies have surveyed approximately 30% of the genome. Several areas of increased lod scores have been identified and are under further investigation.
Between 3 June and 15 July 1967 four explosive outbreaks of acute poisoning with the insecticide endrin occurred in Doha in Qatar and Hofuf in Saudi Arabia. Altogether 874 persons were hospitalized and 26 died. It is estimated that many others were poisoned whose symptoms were not so severe as to cause them to seek medical care or to enter hospital.The author describes the course of the outbreaks and the measures taken to ascertain their cause and prevent their extension and recurrence. It was found that the victims had eaten bread made from flour contaminated with endrin. In two different ships, both of them loaded and off-loaded at different ports, flour and endrin had been stowed in the same hold, with the endrin above the flour. In both ships the endrin containers had leaked and penetrated the sacks of flour which was later used to make bread.These two unconnected but nearly simultaneous mass poisonings emphasize the importance of regulating the carriage of insecticides and other toxic chemicals in such a way as to prevent the contamination of foodstuffs and similar substances during transport; both the World Health Organization and the Inter-Governmental Maritime Consultative Organization are working towards the establishment of regulations and practices to that end.
This paper (i) reviews the current clinical and molecular genetic data which strongly suggest that endometriosis has a genetic basis; (ii) outlines the general principles of affected-sib pair analysis; and (iii) describes the Oxford Endometriosis Gene (OXEGENE) Study which aims, using a positional cloning approach, to identify susceptibility genes involved in the development of the disease.
Given the DNA fingerprints of two individuals with some bands being shared by both individuals, we define a new measure of the degree of similarity between the DNA profiles of two individuals. We use this measure to calculate the expected DNA similarity of two unrelated individuals of a randomly mating population; this similarity is due to chance only. Then, the expected similarity between two related individuals is obtained; this similarity is due to chance and relatedness. From these results, the degree of similarity due to relatedness alone may be calculated.
Genetic chiasma interference occurs when one crossover influences the probability of another crossover occurring nearby. While interference is known to occur in humans, it is typically ignored when computing multipoint likelihoods for genetic mapping. This biologically unsound assumption of no interference facilitates the calculation of the likelihoods at the expense of reduced power to accurately construct a genetic map. We have developed a computer program that calculates multipoint likelihoods of three-generation nuclear families while taking interference into account. In our program, interference is modelled by using a map function to convert genetic distances into recombination fractions. We can determine which of several map functions best fits the data by comparing the multipoint likelihoods of the data under each map function. Since the distribution of the difference between likelihoods is unknown, we use a simulation approach to determine the statistical significance of our results. When our program is applied to six loci, D10S34, D10S19, D10S16, D10S14, D10S4, and D10S20, from the CEPH consortium map of chromosome 10, we find significant evidence in favor of positive interference as modelled by the Sturt map function.
Explore the source record for details and available documents.
The affected-pedigree-member (APM) method of linkage analysis is a nonparametric statistic for testing for nonindependent segregation of a marker to affected members of a pedigree. We present here results of a simulation study evaluating the power of the APM method to detect linkage. We have systematically explored, by computer simulation, the effect of a variety of factors on the power to detect linkage using the single-locus APM statistic. These factors include mode of inheritance, marker polymorphism, the distance between marker and disease, phenocopy rate, heterogeneity, and misspecified marker allele frequencies. We also evaluated the relative power obtained under fixed-structure sampling and sequential sampling. For a dominant disease, sequential sampling led to increased power as compared to fixed-structure sampling, while for a recessive disease, there was no clear advantage in sampling beyond nuclear families.
The affected-pedigree-member (APM) method of linkage analysis detects linkage by testing for increased marker similarity between affected individuals in a pedigree. Thus, the APM method poses a 'one-locus' question about the inheritance of the marker; it requires and makes no assumptions about the inheritance of the disease. It only requires that one be able to accurately determine who is affected. Since the APM method is model free, it is an important tool in testing for linkage of genetically complex diseases for which it is difficult to determine or define an accurate genetic model. In addition, the APM method may be used as a model-free confirmatory test to complement traditional lod score analyses. The X-linked version of the APM method developed here now permits researchers to apply this method to test for linkage of complex diseases to X-linked markers.
We have developed a version of the CRI-MAP computer program for genetic likelihood computations that runs the FLIPS and ALL functions of CRI-MAP in parallel on a distributed network of workstations. The performance of CRI-MAP-PVM was assessed in several linkage analyses using the FLIPS option of CRI-MAP on a map of 85 microsatellite markers for human chromosome 1. These analyses showed excellent speedup and efficiency and low distribution overhead. In addition, we have adapted the MultiMap program for automated construction of linkage maps to use CRI-MAP-PVM. These improvements significantly reduce the time required to compare likelihoods of different marker orders. Thus, the construction of linkage maps can proceed in a more timely fashion, in keeping with recent advances in genotyping technology.
While much effort has gone into developing efficient algorithms for calculating multipoint likelihoods, these calculations still form a significant bottleneck in the construction of genetic linkage maps. Our approach to this problem is based on incremental processing techniques, which attempt to reduce the time required to perform iterative computations by storing intermediate results during the initial iteration, so that they may be reused with little extra computation in subsequent iterations. We have developed an incremental program which provides a more efficient substitute for the CMAP program of the LINKAGE package. Our incremental approach stores intermediate results of the computations in the form of a rational function. Thus, computing the likelihood for one position of an unmapped marker locus requires only the reevaluation of the rational function. Timing data suggest that when pedigrees are fully or nearly fully typed, our program runs about 3-fold faster than CMAP to compute the likelihood for one position of a marker locus. Additional positions do not add any appreciable time to our program; thus, speedups become more pronounced as more marker locus positions are considered.