Search PubMedSearch

Biomedical subjects

G A Churchill

Publications and source records attributed to G A Churchill.

17 recordsLinked to original sources

Genetic analysis of susceptibility to dextran sulfate sodium-induced colitis in mice.

The genetic basis for differential sensitivity of inbred mice to inflammatory bowel disease induced by dextran sulfate sodium (DSS) is unknown. Susceptible C3H/HeJ were outcrossed to partially resistant C57BL/6J mice. F2 and N2 progeny were phenotyped by evaluating histopathologic lesions in large intestine detected 16 days after a 5-day period of feeding 3.5% DSS. Screening for DSS colitis (Dssc) loci revealed quantitative trait loci (QTL) on Chr 5 (Dssc1) and Chr 2 (Dssc2). These traits contributed additively, explaining 17.5% of the variation in total colonic lesions. Additional QTL on Chr 18 and 1 that collectively explained 11% of the variation in total colon lesions were indicated. In the cecum, only a putative QTL on Chr 11 was associated with pathology (lesion severity) in the cecum. Reduced DSS susceptibility was observed in congenic stocks in which the highly susceptible NOD/Lt strain carried putative resistance alleles from either B6 on Chr 2 or from the highly resistant NON/Lt strain on Chr 9. We conclude that multiple genes control susceptibility to DSS colitis in mice. Possible Dssc candidate genes are discussed in terms of current knowledge of inflammatory bowel disease susceptibility loci in humans.

Animals

Multigenic and imprinting control of ovarian granulosa cell tumorigenesis in mice.

Spontaneous juvenile ovarian granulosa cell (GC) tumors that occur in young girls are similar to GC carcinomas that develop in SWR-derived inbred mice. We analyzed female offspring from a series of matings among SWR and SJL inbred mice for chromosomal loci underlying tumor susceptibility. Intercross F2 female mice were produced by reciprocal matings of (SWR x SJL)F1 and (SJL x SWR)F1 parents. Tumorigenesis in these F2 mice as well as in SWXJ recombinant inbred and congenic strains of mice derived from SWR and SJL showed significant (P < 0.001) association with Gct1, a dominant susceptibility locus on chromosome (CHR) 4 and with Gct2 on CHR 12. Suggestive (P < 0.01) association was found with Gct3 on CHR 15. A fourth susceptibility locus, Gct4 on CHR X, was demonstrated with a strong parent-of-origin effect associated with the paternal genotype. Imprinting and complex interactions among these four loci combine to establish the probability for GC tumorigenesis in this mouse model.

Alleles

Biases in amino acid replacement matrices and alignment scores due to rate heterogeneity.

Empirically derived amino acid replacement matrices are widely used in sequence comparison and database searches. We consider an extension of the usual Markov process model of protein evolution that admits site to site rate heterogeneity and demonstrates that rate heterogeneity can introduce a bias in estimated replacement probabilities and the corresponding alignment scores derived from these matrices. We suggest an approach to obtain unbiased estimates of replacement probabilities and alignment scores and derive the details for the case where rates are assumed to vary according to a gamma distribution.

Amino Acid Sequence

Permutation tests for multiple loci affecting a quantitative character.

The problem of detecting minor quantitative trait loci (QTL) responsible for genetic variation not explained by major QTL is of importance in the complete dissection of quantitative characters. Two extensions of the permutation-based method for estimating empirical threshold values are presented. These methods, the conditional empirical threshold (CET) and the residual empirical threshold (RET), yield critical values that can be used to construct tests for the presence of minor QTL effects while accounting for effects of known major QTL. The CET provides a completely nonparametric test through conditioning on markers linked to major QTL. It allows for general nonadditive interactions among QTL, but its practical application is restricted to regions of the genome that are unlinked to the major QTL. The RET assumes a structural model for the effect of major QTL, and a threshold is constructed using residuals from this structural model. The search space for minor QTL is unrestricted, and RET-based tests may be more powerful than the CET-based test when the structural model is approximately true.

Alleles

Heterogeneity in rates of recombination across the mouse genome.

If loci are randomly distributed on a physical map, the density of markers on a genetic map will be inversely proportional to recombination rate. First, proposed by Mary Lyon, we have used this idea to estimate recombination rates from the Drosophila melanogaster linkage map. These results were compared with results of two other studies that estimated regional recombination rates in D. melanogaster using both physical and genetic maps. The three methods were largely concordant in identifying large-scale genomic patterns of recombination. The marker density method was then applied to the Mus musculus microsatellite linkage map. The distribution of microsatellites provided evidence for heterogeneity in recombination rates. Centromeric regions for several mouse chromosomes had significantly greater numbers of markers than expected, suggesting that recombination rates were lower in these regions. In contrast, most telomeric regions contained significantly fewer markers than expected. This indicates that recombination rates are elevated at the telomeres of many mouse chromosomes and is consistent with a comparison of the genetic and cytogenetic maps in these regions. The density of markers on a genetic map may provide a generally useful way to estimate regional recombination rates in species for which genetic, but not physical, maps are available.

Animals

A Hidden Markov Model approach to variation among sites in rate of evolution.

The method of Hidden Markov Models is used to allow for unequal and unknown evolutionary rates at different sites in molecular sequences. Rates of evolution at different sites are assumed to be drawn from a set of possible rates, with a finite number of possibilities. The overall likelihood of phylogeny is calculated as a sum of terms, each term being the probability of the data given a particular assignment of rates to sites, times the prior probability of that particular combination of rates. The probabilities of different rate combinations are specified by a stationary Markov chain that assigns rate categories to sites. While there will be a very large number of possible ways of assigning rates to sites, a simple recursive algorithm allows the contributions to the likelihood from all possible combinations of rates to be summed, in a time proportional to the number of different rates at a single site. Thus with three rates, the effort involved is no greater than three times that for a single rate. This "Hidden Markov Model" method allows for rates to differ between sites and for correlations between the rates of neighboring sites. By summing over all possibilities it does not require us to know the rates at individual sites. However, it does not allow for correlation of rates at nonadjacent sites, nor does it allow for a continuous distribution of rates over sites. It is shown how to use the Newton-Raphson method to estimate branch lengths of a phylogeny and to infer from a phylogeny what assignment of rates to sites has the largest posterior probability. An example is given using beta-hemoglobin DNA sequences in eight mammal species; the regions of high and low evolutionary rates are inferred and also the average length of patches of similar rates.

Animals

Properties of statistical tests of neutrality for DNA polymorphism data.

A class of statistical tests based on molecular polymorphism data is studied to determine size and power properties. The class includes Tajima's D statistic as well as the D* and F* tests proposed by Fu and Li. A new method of constructing critical values for these tests is described. Simulations indicate that Tajima's test is generally most powerful against the alternative hypotheses of selective sweep, population bottleneck, and population subdivision, among tests within this class. However, even Tajima's test can detect a selective sweep or bottleneck only if it has occurred within a specific interval of time in the recent past or population subdivision only when it has persisted for a very long time. For greatest power against the particular alternatives studied here, it is better to sequence more alleles than more sites.

Computer Simulation

Estimation and reliability of molecular sequence alignments.

The problem of estimating the relatedness of a pair of biological sequences is addressed. A stochastic model of sequence evolution is described that allows insertion and deletion as well as replacement of amino acid residues (or substitution of nucleotides) over time. An expectation-maximization (EM) algorithm that obtains maximum likelihood estimates of the model parameters is introduced. The method assumes that the sequences are related by descent from a common ancestor but the alignment (i.e., the precise evolutionary correspondence between residues in each sequence) is unknown. Results from the E-step of the EM algorithm are used to assess the likelihood that any two residues are related by direct descent from a common ancestor.

Algorithms

Empirical threshold values for quantitative trait mapping.

The detection of genes that control quantitative characters is a problem of great interest to the genetic mapping community. Methods for locating these quantitative trait loci (QTL) relative to maps of genetic markers are now widely used. This paper addresses an issue common to all QTL mapping methods, that of determining an appropriate threshold value for declaring significant QTL effects. An empirical method is described, based on the concept of a permutation test, for estimating threshold values that are tailored to the experimental data at hand. The method is demonstrated using two real data sets derived from F(2) and recombinant inbred plant populations. An example using simulated data from a backcross design illustrates the effect of marker density on threshold values.

Chromosome Mapping

Pooled-sampling makes high-resolution mapping practical with DNA markers.

A pooled-sample approach to the construction of high-resolution genetic maps is described. The strategy depends on the existence of an easily selectable target locus and the ability to produce large segregating populations. If these requirements are met, the pooled-sample mapping approach allows tightly linked markers (e.g., restriction fragment length polymorphisms) to be mapped relative to the target with a great economy of effort. The recombination fractions among loci can be estimated by the maximum likelihood method and a simple approximate estimator is derived. The order of loci is deduced using a Bayesian statistical framework to yield posterior probabilities for all possible orderings of a marker set. Optimal pooling strategies and the effects of misclassification of selected individuals are discussed and studied by computer simulation. The feasibility of this method is demonstrated by the high-resolution mapping of a region on chromosome 5 of tomato that contains a gene regulating fruit ripening.

Animals

Network models for sequence evolution.

We introduce a general class of models for sequence evolution that includes network phylogenies. Networks, a generalization of strictly tree-like phylogenies, are proposed to model situations where multiple lineages contribute to the observed sequences. An algorithm to compute the probability distribution of binary character-state configurations is presented and statistical inference for this model is developed in a likelihood framework. A stepwise procedure based on likelihood ratios is used to explore the space of models. Starting with a star phylogeny, new splits (nontrivial bipartitions of the sequence set) are successively added to the model until no significant change in the likelihood is observed. A novel feature of our approach is that the new splits are not necessarily constrained to be consistent with a treelike mode of evolution. The fraction of invariable sites is estimated by maximum likelihood simultaneously with other model parameters and is essential to obtain a good fit to the data. The effect of finite sequence length on the inference methods is discussed. Finally, we provide an illustrative example using aligned VP1 genes from the foot and mouth disease viruses (FMDV). The different serotypes of the FMDV exhibit a range of treelike and network evolutionary relationships.

Aphthovirus

Phylogenetic inference: linear invariants and maximum likelihood.

We develop a new statistical method for inferring phylogenies, based on a likelihood ratio test. This method does not require parameter constraints but does require identical evolutionary processes in the sites considered. Another method of phylogenetic inference is the method of linear invariants, described by Cavender (1989, Molecular Biology and Evolution 6, 301-316), based on a notion of Lake (1987, Molecular Biology and Evolution 4, 167-191). We describe a sound mathematical basis for the use of linear invariants. We show that the validity of the method requires parameter constraints, but does not require that the evolutionary processes in differing sites be identical. We show that the method of linear invariants is asymptotically equivalent to a less powerful version of our likelihood ratio test, and is thus essentially a maximum likelihood technique.

Animals

The accuracy of DNA sequences: estimating sequence quality.

In this paper we describe a method for the statistical reconstruction of a large DNA sequence from a set of sequenced fragments. We assume that the fragments have been assembled and address the problem of determining the degree to which the reconstructed sequence is free from errors, i.e., its accuracy. A consensus distribution is derived from the assembled fragment configuration based upon the rates of sequencing errors in the individual fragments. The consensus distribution can be used to find a minimally redundant consensus sequence that meets a prespecified confidence level, either base by base or across any region of the sequence. A likelihood-based procedure for the estimation of the sequencing error rates, which utilizes an iterative EM algorithm, is described. Prior knowledge of the error rates is easily incorporated into the estimation procedure. The methods are applied to a set of assembled sequence fragments from the human G6PD locus. We close the paper with a brief discussion of the relevance and practical implications of this work.

Algorithms

Sample size for a phylogenetic inference.

The objective of this work is to describe sample-size calculations for the inference of a nonzero central branch length in an unrooted four-species phylogeny. Attention is restricted to independent binary characters, such as might be obtained from an alignment of the purine-pyrimidine sequences of a nucleic acid molecule. A statistical test based on a multinomial model for character-state configurations is described. The importance of including invariable sites in models for sequence change is demonstrated, and their effect on sample size is quantified. The methods are applied to a four-species alignment of small-subunit rRNA sequences derived from two archaebacteria, a eubacteria and a eukaryote. We conclude that the information in these sequences is not sufficient to resolve the branching order of this tree. Estimates of the number of aligned nucleotide positions required to provide a reasonably powerful test are given.

Archaea

Methods for inferring phylogenies from nucleic acid sequence data by using maximum likelihood and linear invariants.

Likelihood methods and methods using invariants are procedures for inferring the evolutionary relationships among species through statistical analysis of nucleic acid sequences. A likelihood-ratio test may be used to determine the feasibility of any tree for which the maximum likelihood can be computed. The method of linear invariants described by Cavender, which includes Lake's method of evolutionary parsimony as a special case, is essentially a form of the likelihood-ratio method. In the case of a small number of species (four or five), these methods may be used to find a confidence set for the correct tree. An exact version of Lake's asymptotic chi 2 test has been mentioned by Holmquist et al. Under very general assumptions, a one-sided exact test is appropriate, which greatly increases power.

Animals

The distribution of restriction enzyme sites in Escherichia coli.

A statistical analysis of physical map data for eight restriction enzymes covering nearly the entire genome of E. coli is presented. The methods of analysis are based on a top-down modeling approach which requires no knowledge of the statistical properties of the base sequence. For most enzymes, the distribution of mapped sites is found to be fairly homogeneous. Some heterogeneity in the distribution of sites is observed for the enzymes Pstl and HindIII. In addition, BamHI sites are found to be more evenly dispersed than we would expect for random placement and we speculate on a possible mechanism. A consistent departure from a uniform distribution, observed for each of the eight enzymes, is found to be due to a lack of closely spaced sites. We conclude from our analysis that this departure can be accounted for by deficiencies in the physical map data rather than non-random placement of actual restriction sites. Estimates of the numbers of sites missing from the map are given, based both on the map data itself and on the site frequencies in a sample of sequenced E. coli DNA. We conclude that 5 to 15% of the mapped sites represent multiple sites in the DNA sequence.

DNA, Bacterial

Stochastic models for heterogeneous DNA sequences.

The composition of naturally occurring DNA sequences is often strikingly heterogeneous. In this paper, the DNA sequence is viewed as a stochastic process with local compositional properties determined by the states of a hidden Markov chain. The model used is a discrete-state, discrete-outcome version of a general model for non-stationary time series proposed by Kitagawa (1987). A smoothing algorithm is described which can be used to reconstruct the hidden process and produce graphic displays of the compositional structure of a sequence. The problem of parameter estimation is approached using likelihood methods and an EM algorithm for approximating the maximum likelihood estimate is derived. The methods are applied to sequences from yeast mitochondrial DNA, human and mouse mitochondrial DNAs, a human X chromosomal fragment and the complete genome of bacteriophage lambda.

Base Sequence