Search PubMedSearch

Biomedical subjects

T P Speed

Publications and source records attributed to T P Speed.

At least 19 recordsLinked to original sources

The limits of random fingerprinting.

Various random fingerprinting methods are sometimes used to detect overlap between pairs of clones as a first step toward producing a minimal tiling path of clones for subsequent mapping and sequencing efforts. This paper evaluates and compares various statistical procedures for detecting pairwise overlap between clones when the fingerprints arise from any random process meeting simple, plausible assumptions about the relationship between overlap and the resulting fingerprint. Examples of such random processes include, but are not limited to, large-scale hybridization procedures designed to prepare tiling paths of clones for subsequent large-scale genomic sequencing. Our goals are to assess how well random fingerprinting can possibly detect overlap, to assess the effects of inevitable fingerprinting errors on statistical detection, to determine how one can make the best use of the data random fingerprinting provides, and to evaluate how well simple, heuristic techniques for overlap detection compare to more complex, likelihood-based approaches. The paper provides a quantitative assessment of the ability of any random fingerprinting procedure to detect various proportions of clonal overlap and shows the extent to which a small amount of experimental error will vitiate the performance of such techniques. The paper outlines a simple approximation method for constructing Bayesian overlap detectors, while concluding that detectors constructed from linear combinations of fingerprint data can be designed that will perform nearly as well as more complex, likelihood-based approaches.

DNA Fingerprinting

Over- and underrepresentation of short DNA words in herpesvirus genomes.

The relative abundance and rarity of DNA words have been recognized in previous biological studies to have implications for the regulation, repair, and evolutionary mechanisms of a genome. In this paper, we review several different measures of abundance and rarity of DNA words, including z-scores, representation ratios, and cross-ratios, that have appeared in the recent literature, and examine the concordance among them using the human cytomegalovirus genome sequence. We then rank all words of length k = 2, ..., 5 of seven herpesvirus genomes according to their abundance, as measured by one of the z-scores based upon a stationary Markov model of order k-2. Using a simple metric on the ranks of 2-words of the seven herpesvirus sequences, we construct an evolutionary tree. Several 3-words are observed to be consistently over- or underrepresented in all seven herpesviruses. Furthermore, clusters of some of the most over- and underrepresented 4- and 5-words in the genomes are identified with functional sites such as the origins of replication and regulatory signals of individual viruses.

Algorithms

On genetic map functions.

Various genetic map functions have been proposed to infer the unobservable genetic distance between two loci from the observable recombination fraction between them. Some map functions were found to fit data better than others. When there are more than three markers, multilocus recombination probabilities cannot be uniquely determined by the defining property of map functions, and different methods have been proposed to permit the use of map functions to analyze multilocus data. If for a given map function, there is a probability model for recombination that can give rise to it, then joint recombination probabilities can be deduced from this model. This provides another way to use map functions in multilocus analysis. In this paper we show that stationary renewal processes give rise to most of the map functions in the literature. Furthermore, we show that the interevent distributions of these renewal processes can all be approximated quite well by gamma distributions.

Chromosome Mapping

A note on the combination of estimates of a recombination fraction.

A number of ways of combining two or more independent estimates of the same recombination fraction can be found in the literature. We revisit this topic in the context of human gene mapping, and explore the value of transforming the recombination fraction to a new parameter whose log-likelihood function is more nearly quadratic. It is shown that the arcsine of the cube-root is one such function. These observations lead naturally to a way of summarizing and combining the summarized set of log-likelihood functions of a common recombination fraction. This idea is illustrated using pedigree data concerning six loci on chromosome 10 from the CEPH consortium. A comparison is also made with the method of summarizing and combining using 'equivalent numbers' of recombinants and informative meioses.

Chromosome Mapping

Relative efficiencies of chi 2 models of recombination for exclusion mapping and gene ordering.

In multilocus linkage analysis, it is common to assume chiasma interference is absent. While this assumption provides mathematical tractability, there is substantial biological evidence contradicting it, particularly when the loci are closely spaced. The chi 2 class of recombination models, recently described by Foss et al. (Genetics 133: 681-691, 1993), has a plausible biological basis and provides a dramatically improved fit over virtually all other models currently in use. Here, a simulation study is performed to assess the relative efficiency of a no interference model analysis to an analysis with a chi 2 model which allows for interference. The results presented show that analysis with the no interference model is inefficient in the presence of interference.

Chi-Square Distribution

Modeling interference in genetic recombination.

In analyzing genetic linkage data it is common to assume that the locations of crossovers along a chromosome follow a Poisson process, whereas it has long been known that this assumption does not fit the data. In many organisms it appears that the presence of a crossover inhibits the formation of another nearby, a phenomenon known as "interference." We discuss several point process models for recombination that incorporate position interference but assume no chromatid interference. Using stochastic simulation, we are able to fit the models to a multilocus Drosophila dataset by the method of maximum likelihood. We find that some biologically inspired point process models incorporating one or two additional parameters provide a dramatically better fit to the data than the usual "no-interference" Poisson model.

Animals

Statistical analysis of crossover interference using the chi-square model.

The chi-square model (also known as the gamma model with integer shape parameter) for the occurrence of crossovers along a chromosome was first proposed in the 1940's as a description of interference that was mathematically tractable but without biological basis. Recently, the chi-square model has been reintroduced into the literature from a biological perspective. It arises as a result of certain hypothesized constraints on the resolution of randomly distributed crossover intermediates. In this paper under the assumption of no chromatid interference, the probability for any single spore or tetrad joint recombination pattern is derived under the chi-square model. The method of maximum likelihood is then used to estimate the chi-square parameter m and genetic distances among marker loci. We discuss how to interpret the goodness-of-fit statistics appropriately when there are some recombination classes that have only a small number of observations. Finally, comparisons are made between the chi-square model and some other tractable models in the literature.

Animals

Statistical analysis of chromatid interference.

The nonrandom occurrence of crossovers along a single strand during meiosis can be caused by either chromatid interference, crossover interference or both. Although crossover interference has been consistently observed in almost all organisms since the time of the first linkage studies, chromatid interference has not been as thoroughly discussed in the literature, and the evidence provided for it is inconsistent. In this paper with virtually no restrictions on the nature of crossover interference, we describe the constraints that follow from the assumption of no chromatid interference for single spore data. These constraints are necessary consequences of the assumption of no chromatid interference, but their satisfaction is not sufficient to guarantee no chromatid interference. Models can be constructed in which chromatid interference clearly exists but is not detectable with single spore data. We then extend our analysis to cover tetrad data, which permits more powerful tests of no chromatid interference. We note that the traditional test of no chromatid interference based on tetrad data does not make full use of the information provided by the data, and we offer a statistical procedure for testing the no chromatid interference constraints that does make full use of the data. The procedure is then applied to data from several organisms. Although no strong evidence of chromatid interference is found, we do observe an excess of two-strand double recombinations, i.e., negative chromatid interference.

Animals

Alveolar lining layer is thin and continuous: low-temperature scanning electron microscopy of rat lung.

The low-temperature electron microscope, which preserves aqueous structures as solid water at liquid nitrogen temperature, was used to image the alveolar lining layer, including surfactant and its aqueous subphase, of air-filled lungs frozen in anesthetized rats at 15-cmH2O transpulmonary pressure. Lining layer thickness was measured on cross fractures of walls of the outermost subpleural alveoli that could be solidified with metal mirror cryofixation at rates sufficient to limit ice crystal growth to 10 nm and prevent appreciable water movement. The thickness of the liquid layer averaged 0.14 micron over relatively flat portions of the alveolar walls, 0.89 micron at the alveolar wall junctions, and 0.09 micron over the protruding features (9 rats, 20 walls, 16 junctions, and 146 areas), for an area-weighted average thickness of 0.2 micron. The alveolar lining layer appears continuous, submerging epithelial cell microvilli and intercellular junctional ridges; varies from a few nanometers to several micrometers in thickness, and serves to smooth the alveolar air-liquid interface in lungs inflated to zone 1 or 2 conditions.

Animals

Tests of random mating for a highly polymorphic locus: application to HLA data.

Testing for random mating at an HLA locus is a difficult problem because of the highly polymorphic nature of the HLA loci. We discuss some methodological issues and propose several tests. A simulation study is conducted to evaluate these tests. The single allele test and the shared allele test deal with small sample sizes by aggregating the data in different ways. The shared allele test is found to be a more powerful method of detecting non-random mating patterns involving a deficiency or an excess of similar genotypes than the single allele test. We show that random mating of couple at the genotype level implies the random mating of couple at the allele level. Several multi-allele approaches are proposed for large population-based data sets. Among them, the corrected allele-table test performs better than the generalized Wald test in terms of power and size. These methods are then applied to an HLA data set of Caucasian couples, and no solid evidence for non-random mating at the HLA A, B, and DR loci is found.

Alleles

Reproductive failure and the major histocompatibility complex.

The association between HLA sharing and recurrent spontaneous abortion (RSA) was tested in 123 couples and the association between HLA sharing, and the outcome of treatment for unexplained infertility by in vitro fertilization (IVF) was tested in 76 couples, by using a new shared-allele test in order to identify more precisely the region of the major histocompatibility complex (MHC) influencing these reproductive defects. The shared-allele test circumvents the problem of rare alleles at HLA loci and at the same time provides a substantial gain in power over the simple chi 2 test. Two statistical methods, a corrected homogeneity test and a bootstrap approach, were developed to compare the allele frequencies at each of the HLA-A, HLA-B, HLA-DR, and HLA-DQ loci; they were not statistically different among the three patient groups and the control group. There was a significant excess of HLA-DR sharing in couples with RSA and a significant excess of HLA-DQ sharing in couples with unexplained infertility who failed treatment by IVF. These findings indicate that genes located in different parts of the class II region of the MHC affect different aspects of reproduction and strongly suggest that the sharing of HLA antigens per se is not the mechanism involved in the reproductive defects. The segment of the MHC that has genes affecting reproduction also has genes associated with different autoimmune diseases, and this juxtaposition may explain the association between reproductive defects and autoimmune diseases.

Abortion, Habitual

Predicting progress in directed mapping projects.

Several recent mapping efforts have used so-called "directed" approaches to construct their maps. However, most, but not all, published methods for modeling the progress in physical mapping projects have been focused on random approaches, such as bottom-up fingerprinting and STS-content mapping. In addition, those few efforts that did model directed approaches used methods that required assuming that all insert lengths were the same. This assumption is unnecessary. Using properties of stationary processes, one can derive simple asymptotic formulas that apply equally to constant and variable clone lengths. Also, in the case of constant clone lengths, these results are equivalent to, and extend, those published results for directed mapping derived by other methods. Simulations show that these methods provide estimates well within the limits of uncertainty inherent in any mapping project.

Chromosome Mapping

Atypical regions in large genomic DNA sequences.

Large genomic DNA sequences contain regions with distinctive patterns of sequence organization. We describe a method using logarithms of probabilities based on seventh-order Markov chains to rapidly identify genomic sequences that do not resemble models of genome organization built from compilations of octanucleotide usage. Data bases have been constructed from Escherichia coli and Saccharomyces cerevisiae DNA sequences of > 1000 nt and human sequences of > 10,000 nt. Atypical genes and clusters of genes have been located in bacteriophage, yeast, and primate DNA sequences. We consider criteria for statistical significance of the results, offer possible explanations for the observed variation in genome organization, and give additional applications of these methods in DNA sequence analysis.

Bacteriophage lambda

Testing for segregation distortion in the HLA complex.

One of the long-standing issues in HLA research is whether there is segregation distortion in the HLA complex in human populations. In this paper we study some simple statistical models aimed at detecting segregation distortion. We present a statistic to test the Mendelian null hypothesis of equal transmission probabilities. To assess the possible contribution of multiple alleles to segregation distortion, we employ a specific log-linear model for transmission probabilities equivalent to the Bradley-Terry model in the literature of paired comparisons. We also provide a simple method for detecting a single allele effect, if present.

Alleles

Factors associated with human immunodeficiency virus seroconversion in homosexual men in three San Francisco cohort studies, 1984-1989.

A total of 83 HIV seroconversions occurred between 1984 and 1989 in three San Francisco cohorts of homosexual and bisexual men. A nested case-control analysis was performed to assess the risk of seroconversion associated with sexual practices. Strong associations were found with total number of intercourse partners and receptive anal intercourse. Weaker, but significant associations were found with receptive oral intercourse. Individuals reporting condom use with some partners were more likely to become infected than those reporting no condom use. Some of this difference was due to increased numbers of sexual partners among those reporting some condom use, but the association remained significant in multivariate analyses. This observation indicates there is some unmeasured risk associated with those using condoms, such as more HIV seropositive partners.

Cohort Studies

Robustness of the no-interference model for ordering genetic markers.

Under the assumption of no chromatid interference, we derive constraints on the probabilities of the different recombination patterns among m + 1 genetic loci. An application of these constraints is a proof that the ordering of the loci that maximizes the likelihood under the assumption of no interference is, in fact, a consistent estimator of the true order even when there is interference.

Chromatids

Estimating the fraction of invariable codons with a capture-recapture method.

A codon-based approach to estimating the number of variable sites in a protein is presented. When first and second positions of codons are assumed to be replacement positions, a capture-recapture model can be used to estimate the number of variable codons from every pair of homologous and aligned sequences. The capture-recapture estimate is compared to a maximum likelihood estimate of the number of variable codons and to previous approaches that estimate the number of variable sites (not codons) in a sequence. Computer simulations are presented that show under which circumstances the capture-recapture estimate can be used to correct biases in distance matrices. Analysis of published sequences of two genes, calmodulin and serum albumin, shows that distance corrections that employ a capture-recapture estimate of the number of variable sites may be considerably different from corrections that assume that the number of variable sites is equal to the total number of positions in the sequence.

Biological Evolution