Search PubMed⌕ Search

Biomedical subjects

Shizhong Xu

Publications and source records attributed to Shizhong Xu.

At least 19 recordsLinked to original sources

Bayesian shrinkage estimation of quantitative trait loci parameters.

Mapping multiple QTL is a typical problem of variable selection in an oversaturated model because the potential number of QTL can be substantially larger than the sample size. Currently, model selection is still the most effective approach to mapping multiple QTL, although further research is needed. An alternative approach to analyzing an oversaturated model is the shrinkage estimation in which all candidate variables are included in the model but their estimated effects are forced to shrink toward zero. In contrast to the usual shrinkage estimation where all model effects are shrunk by the same factor, we develop a Bayesian method that allows the shrinkage factor to vary across different effects. The new shrinkage method forces marker intervals that contain no QTL to have estimated effects close to zero whereas intervals containing notable QTL have estimated effects subject to virtually no shrinkage. We demonstrate the method using both simulated and real data for QTL mapping. A simulation experiment with 500 backcross (BC) individuals showed that the method can localize closely linked QTL and QTL with effects as small as 1% of the phenotypic variance of the trait. The method was also used to map QTL responsible for wound healing in a family of a (MRL/MPJ x SJL/J) cross with 633 F(2) mice derived from two inbred lines.

Animals↗

Mapping quantitative trait loci using naturally occurring genetic variance among commercial inbred lines of maize (Zea mays L.).

Many commercial inbred lines are available in crops. A large amount of genetic variation is preserved among these lines. The genealogical history of the inbred lines is usually well documented. However, quantitative trait loci (QTL) responsible for the genetic variances among the lines are largely unexplored due to lack of statistical methods. In this study, we show that the pedigree information of the lines along with the trait values and marker information can be used to map QTL without the need of further crossing experiments. We develop a Monte Carlo method to estimate locus-specific identity-by-descent (IBD) matrices. These IBD matrices are further incorporated into a mixed-model equation for variance component analysis. QTL variance is estimated and tested at every putative position of the genome. The actual QTL are detected by scanning the entire genome. Applying this new method to a well-documented pedigree of maize (Zea mays L.) that consists of 404 inbred lines, we mapped eight QTL for the maize male flowering trait, growing degree day heat units to pollen shedding (GDUSHD). These detected QTL contributed >80% of the variance observed among the inbred lines. The QTL were then used to evaluate all the inbred lines using the best linear unbiased prediction (BLUP) technique. Superior lines were selected according to the estimated QTL allelic values, a technique called marker-assisted selection (MAS). The MAS procedure implemented via BLUP may be routinely used by breeders to select superior lines and line combinations for development of new cultivars.

Alleles↗

Clustering expressed genes on the basis of their association with a quantitative phenotype.

Cluster analyses of gene expression data are usually conducted based on their associations with the phenotype of a particular disease. Many disease traits have a clearly defined binary phenotype (presence or absence), so that genes can be clustered based on the differences of expression levels between the two contrasting phenotypic groups. For example, cluster analysis based on binary phenotype has been successfully used in tumour research. Some complex diseases have phenotypes that vary in a continuous manner and the method developed for a binary trait is not immediately applicable to a continuous trait. However, understanding the role of gene expression in these complex traits is of fundamental importance. Therefore, it is necessary to develop a new statistical method to cluster expressed genes based on their association with a quantitative trait phenotype. We developed a model-based clustering method to classify genes based on their association with a continuous phenotype. We used a linear model to describe the relationship between gene expression and the phenotypic value. The model effects of the linear model (linear regression coefficients) represent the strength of the association. We assumed that the model effects of each gene follow a mixture of several multivariate Gaussian distributions. Parameter estimation and cluster assignment were accomplished via an Expectation-Maximization (EM) algorithm. The method was verified by analysing two simulated datasets, and further demonstrated using real data generated in a microarray experiment for the study of gene expression associated with Alzheimer's disease.

Algorithms↗

Identification of QTL for production traits in chickens.

If the poultry industry hopes to continue to flourish, the identification of potential quantitative trait loci (QTL) for production-related traits must be pursued This remains true despite the sequencing of the chicken genome. In view of this need, a scan of the chicken genome using 72 microsatellite markers was carried out on a meat-type x egg-type resource population measured for production and egg quality traits. Using a Bayesian analysis, potential QTL for a number of traits were identified on several chromosomes. Evidence of eight QTL regions associated with a total of eight traits (specific gravity, albumin height, Haugh score, shell shape, total number of eggs, final body weight, gain, and feed efficiency) was found. Two of these regions, one spanning the area of 263/287 cM on GAA01 and the other spanning the area of 23/28 cM on GAA02, were associated with multiple QTL.

Animals↗

Joint mapping of quantitative trait Loci for multiple binary characters.

Joint mapping for multiple quantitative traits has shed new light on genetic mapping by pinpointing pleiotropic effects and close linkage. Joint mapping also can improve statistical power of QTL detection. However, such a joint mapping procedure has not been available for discrete traits. Most disease resistance traits are measured as one or more discrete characters. These discrete characters are often correlated. Joint mapping for multiple binary disease traits may provide an opportunity to explore pleiotropic effects and increase the statistical power of detecting disease loci. We develop a maximum-likelihood method for mapping multiple binary traits. We postulate a set of multivariate normal disease liabilities, each contributing to the phenotypic variance of one disease trait. The underlying liabilities are linked to the binary phenotypes through some underlying thresholds. The new method actually maps loci for the variation of multivariate normal liabilities. As a result, we are able to take advantage of existing methods of joint mapping for quantitative traits. We treat the multivariate liabilities as missing values so that an expectation-maximization (EM) algorithm can be applied here. We also extend the method to joint mapping for both discrete and continuous traits. Efficiency of the method is demonstrated using simulated data. We also apply the new method to a set of real data and detect several loci responsible for blast resistance in rice.

Algorithms↗

Supervised cluster analysis for microarray data based on multivariate Gaussian mixture.

MOTIVATION: Grouping genes having similar expression patterns is called gene clustering, which has been proved to be a useful tool for extracting underlying biological information of gene expression data. Many clustering procedures have shown success in microarray gene clustering; most of them belong to the family of heuristic clustering algorithms. Model-based algorithms are alternative clustering algorithms, which are based on the assumption that the whole set of microarray data is a finite mixture of a certain type of distributions with different parameters. Application of the model-based algorithms to unsupervised clustering has been reported. Here, for the first time, we demonstrated the use of the model-based algorithm in supervised clustering of microarray data. RESULTS: We applied the proposed methods to real gene expression data and simulated data. We showed that the supervised model-based algorithm is superior over the unsupervised method and the support vector machines (SVM) method. AVAILABILITY: The program written in the SAS language implementing methods I-III in this report is available upon request. The software of SVMs is available in the website http://svm.sdsc.edu/cgi-bin/nph-SVMsubmit.cgi

Algorithms↗

Mapping QTLs for traits measured as percentages.

Many quantitative traits are measured as percentages. As a result, the assumption of a normal distribution for the residual errors of such percentage data is often violated. However, most quantitative trait locus (QTL) mapping procedures assume normality of the residuals. Therefore, proper data transformation is often recommended before statistical analysis is conducted. We propose the probit transformation to convert percentage data into variables with a normal distribution. The advantage of the probit transformation is that it can handle measurement errors with heterogeneous variance and correlation structure in a statistically sound manner. We compared the results of this data transformation with other transformations and found that this method can substantially increase the statistical power of QTL detection. We develop the QTL mapping procedure based on the maximum likelihood methodology implemented via the expectation-maximization algorithm. The efficacy of the new method is demonstrated using Monte Carlo simulation.

Chromosome Mapping↗

Mapping multiple quantitative trait Loci for ordinal traits.

Many complex traits in humans and other organisms show ordinal phenotypic variation but do not follow a simple Mendelian pattern of inheritance. These ordinal traits are presumably determined by many factors, including genetic and environmental components. Several statistical approaches to mapping quantitative trait loci (QTL) for such traits have been developed based on a single-QTL model. However, statistical methods for mapping multiple QTL are not well studied as continuous traits. In this paper, we propose a Bayesian method implemented via the Markov chain Monte Carlo (MCMC) algorithm to map multiple QTL for ordinal traits in experimental crosses. We model the ordinal traits under the multiple threshold model, which assumes a latent continuous variable underlying the ordinal phenotypes. The ordinal phenotype and the latent continuous variable are linked through some fixed but unknown thresholds. We adopt a standardized threshold model, which has several attractive features. An efficient sampling scheme is developed to jointly generate the threshold values and the values of latent variable. With the simulated latent variable, the posterior distributions of other unknowns, for example, the number, locations, genetic effects, and genotypes of QTL, can be computed using existing algorithms for normally distributed traits. To this end, we provide a unified approach to mapping multiple QTL for continuous, binary, and ordinal traits. Utility and flexibility of the method are demonstrated using simulated data.

Algorithms↗

Mapping quantitative trait loci in F2 incorporating phenotypes of F3 progeny.

In plants and laboratory animals, QTL mapping is commonly performed using F(2) or BC individuals derived from the cross of two inbred lines. Typical QTL mapping statistics assume that each F(2) individual is genotyped for the markers and phenotyped for the trait. For plant traits with low heritability, it has been suggested to use the average phenotypic values of F(3) progeny derived from selfing F(2) plants in place of the F(2) phenotype itself. All F(3) progeny derived from the same F(2) plant belong to the same F(2:3) family, denoted by F(2:3). If the size of each F(2:3) family (the number of F(3) progeny) is sufficiently large, the average value of the family will represent the genotypic value of the F(2) plant, and thus the power of QTL mapping may be significantly increased. The strategy of using F(2) marker genotypes and F(3) average phenotypes for QTL mapping in plants is quite similar to the daughter design of QTL mapping in dairy cattle. We study the fundamental principle of the plant version of the daughter design and develop a new statistical method to map QTL under this F(2:3) strategy. We also propose to combine both the F(2) phenotypes and the F(2:3) average phenotypes to further increase the power of QTL mapping. The statistical method developed in this study differs from published ones in that the new method fully takes advantage of the mixture distribution for F(2:3) families of heterozygous F(2) plants. Incorporation of this new information has significantly increased the statistical power of QTL detection relative to the classical F(2) design, even if only a single F(3) progeny is collected from each F(2:3) family. The mixture model is developed on the basis of a single-QTL model and implemented via the EM algorithm. Substantial computer simulation was conducted to demonstrate the improved efficiency of the mixture model. Extension of the mixture model to multiple QTL analysis is developed using a Bayesian approach. The computer program performing the Bayesian analysis of the simulated data is available to users for real data analysis.

Algorithms↗

An EM algorithm for mapping binary disease loci: application to fibrosarcoma in a four-way cross mouse family.

Many diseases show dichotomous phenotypic variation but do not follow a simple Mendelian pattern of inheritance. Variances of these binary diseases are presumably controlled by multiple loci and environmental variants. A least-squares method has been developed for mapping such complex disease loci by treating the binary phenotypes (0 and 1) as if they were continuous. However, the least-squares method is not recommended because of its ad hoc nature. Maximum Likelihood (ML) and Bayesian methods have also been developed for binary disease mapping by incorporating the discrete nature of the phenotypic distribution. In the ML analysis, the likelihood function is usually maximized using some complicated maximization algorithms (e.g. the Newton-Raphson or the simplex algorithm). Under the threshold model of binary disease, we develop an Expectation Maximization (EM) algorithm to solve for the maximum likelihood estimates (MLEs). The new EM algorithm is developed by treating both the unobserved genotype and the disease liability as missing values. As a result, the EM iteration equations have the same form as the normal equation system in linear regression. The EM algorithm is further modified to take into account sexual dimorphism in the linkage maps. Applying the EM-implemented ML method to a four-way-cross mouse family, we detected two regions on the fourth chromosome that have evidence of QTLs controlling the segregation of fibrosarcoma, a form of connective tissue cancer. The two QTLs explain 50-60% of the variance in the disease liability. We also applied a Bayesian method previously developed (modified to take into account sex-specific maps) to this data set and detected one additional QTL on chromosome 13 that explains another 26% of the variance of the disease liability. All the QTLs detected primarily show dominance effects.

Algorithms↗

Correcting the bias in estimation of genetic variances contributed by individual QTL.

In addition to locating chromosomal positions of quantitative trait loci (QTL), estimating the sizes of identified QTL is also an important component in QTL mapping. The size of a QTL is usually measured by the proportion of the phenotypic variance contributed by the QTL. However, the genetic variance may be overestimated in a small line crossing experiment. In this study, we investigate this bias and develop a simple method to correct the bias. The bias correction, however, requires the error of the estimated genetic effect, which is not trivial if the genetic effect is estimated using the Expectation and Maximization (EM) algorithm. Therefore, we also develop a simple method to estimate the standard error of the estimated genetic effect, which is subsequently used to correct the bias in the variance estimate.

Algorithms↗

Quantitative trait loci responsible for variation in sexually dimorphic traits in Drosophila melanogaster.

To understand the mechanisms of morphological evolution and species divergence, it is essential to elucidate the genetic basis of variation in natural populations. Sexually dimorphic characters, which evolve rapidly both within and among species, present attractive models for addressing these questions. In this report, we map quantitative trait loci (QTL) responsible for variation in sexually dimorphic traits (abdominal pigmentation and the number of ventral abdominal bristles and sex comb teeth) in a natural population of Drosophila melanogaster. To capture the pattern of genetic variation present in the wild, a panel of recombinant inbred lines was created from two heterozygous flies taken directly from nature. High-resolution mapping was made possible by cytological markers at the average density of one per 2 cM. We have used a new Bayesian algorithm that allows QTL mapping based on all markers simultaneously. With this approach, we were able to detect small-effect QTL that were not evident in single-marker analyses. Our results show that at least for some sexually dimorphic traits, a small number of QTL account for the majority of genetic variation. The three strongest QTL account for >60% of variation in the number of ventral abdominal bristles. Strikingly, a single QTL accounts for almost 60% of variation in female abdominal pigmentation. This QTL maps to the chromosomal region that Robertson et al. have found to affect female abdominal pigmentation in other populations of D. melanogaster. Using quantitative complementation tests, we demonstrate that this QTL is allelic to the bric a brac gene, whose expression has previously been shown to correlate with interspecific differences in pigmentation. Multiple bab alleles that confer distinct phenotypes appear to segregate in natural populations at appreciable frequencies, suggesting that intraspecific and interspecific variation in abdominal pigmentation may share a similar genetic basis.

Animals↗

Estimating polygenic effects using markers of the entire genome.

Molecular markers have been used to map quantitative trait loci. However, they are rarely used to evaluate effects of chromosome segments of the entire genome. The original interval-mapping approach and various modified versions of it may have limited use in evaluating the genetic effects of the entire genome because they require evaluation of multiple models and model selection. Here we present a Bayesian regression method to simultaneously estimate genetic effects associated with markers of the entire genome. With the Bayesian method, we were able to handle situations in which the number of effects is even larger than the number of observations. The key to the success is that we allow each marker effect to have its own variance parameter, which in turn has its own prior distribution so that the variance can be estimated from the data. Under this hierarchical model, we were able to handle a large number of markers and most of the markers may have negligible effects. As a result, it is possible to evaluate the distribution of the marker effects. Using data from the North American Barley Genome Mapping Project in double-haploid barley, we found that the distribution of gene effects follows closely an L-shaped Gamma distribution, which is in contrast to the bell-shaped Gamma distribution when the gene effects were estimated from interval mapping. In addition, we show that the Bayesian method serves as an alternative or even better QTL mapping method because it produces clearer signals for QTL. Similar results were found from simulated data sets of F(2) and backcross (BC) families.

Bayes Theorem↗

Bayesian model choice and search strategies for mapping interacting quantitative trait Loci.

Most complex traits of animals, plants, and humans are influenced by multiple genetic and environmental factors. Interactions among multiple genes play fundamental roles in the genetic control and evolution of complex traits. Statistical modeling of interaction effects in quantitative trait loci (QTL) analysis must accommodate a very large number of potential genetic effects, which presents a major challenge to determining the genetic model with respect to the number of QTL, their positions, and their genetic effects. In this study, we use the methodology of Bayesian model and variable selection to develop strategies for identifying multiple QTL with complex epistatic patterns in experimental designs with two segregating genotypes. Specifically, we develop a reversible jump Markov chain Monte Carlo algorithm to determine the number of QTL and to select main and epistatic effects. With the proposed method, we can jointly infer the genetic model of a complex trait and the associated genetic parameters, including the number, positions, and main and epistatic effects of the identified QTL. Our method can map a large number of QTL with any combination of main and epistatic effects. Utility and flexibility of the method are demonstrated using both simulated data and a real data set. Sensitivity of posterior inference to prior specifications of the number and genetic effects of QTL is investigated.

Bayes Theorem↗

Theoretical basis of the Beavis effect.

The core of statistical inference is based on both hypothesis testing and estimation. The use of inferential statistics for QTL identification thus includes estimation of genetic effects and statistical tests. Typically, QTL are reported only when the test statistics reach a predetermined critical value. Therefore, the estimated effects of detected QTL are actually sampled from a truncated distribution. As a result, the expectations of detected QTL effects are biased upward. In a simulation study, William D. Beavis showed that the average estimates of phenotypic variances associated with correctly identified QTL were greatly overestimated if only 100 progeny were evaluated, slightly overestimated if 500 progeny were evaluated, and fairly close to the actual magnitude when 1000 progeny were evaluated. This phenomenon has subsequently been called the Beavis effect. Understanding the theoretical basis of the Beavis effect will help interpret QTL mapping results and improve success of marker-assisted selection. This study provides a statistical explanation for the Beavis effect. The theoretical prediction agrees well with the observations reported in Beavis's original simulation study. Application of the theory to meta-analysis of QTL mapping is discussed.

Chromosome Mapping↗

Chromosomal regions harboring genes for the work to femur failure in mice.

The work to failure is defined as the maximum energy bone can absorb before breaking, and therefore is a direct test of the risk of fracture. To determine the genetic loci influencing work to failure, we have performed a high density genome-wide scan in 633 (MRL x SJL) F(2) female mice. Five loci ( P<0.005) with significant effects on work to failure were found on chromosomes 2, 7, 8, 9, and X, which collectively explained around 20% variance of work to femur failure in F(2) mice. Of those, only the QTL on chromosome 9 was concordant with bone mineral density (BMD) QTLs. Eight significant interactions ( P<0.01) between marker loci were identified, which accounted for an equivalent amount of F(2) variance (23%) to combined single QTL effects. Our results demonstrate that most of the genetic loci regulating work to failure are different from those for BMD in the 7-week-old female mice. If this is also true in humans, this finding will challenge the predictive value of BMD for the risk of fracture.

Animals↗

Mapping quantitative trait loci with epistatic effects.

Epistatic variance can be an important source of variation for complex traits. However, detecting epistatic effects is difficult primarily due to insufficient sample sizes and lack of robust statistical methods. In this paper, we develop a Bayesian method to map multiple quantitative trait loci (QTLs) with epistatic effects. The method can map QTLs in complicated mating designs derived from the cross of two inbred lines. In addition to mapping QTLs for quantitative traits, the proposed method can even map genes underlying binary traits such as disease susceptibility using the threshold model. The parameters of interest are various QTL effects, including additive, dominance and epistatic effects of QTLs, the locations of identified QTLs and even the number of QTLs. When the number of QTLs is treated as an unknown parameter, the dimension of the model becomes a variable. This requires the reversible jump Markov chain Monte Carlo algorithm. The utility of the proposed method is demonstrated through analysis of simulation data.

Algorithms↗