Search PubMed⌕ Search

Biomedical subjects

Shuanglin Zhang

Publications and source records attributed to Shuanglin Zhang.

12 recordsLinked to original sources

Analytical correction for multiple testing in admixture mapping.

Admixture mapping, using unrelated individuals from the admixture populations that result from recent mating between members of each parental population, is an efficient approach to localize disease-causing variants that differ in frequency between two or more historically separated populations. Recently, several methods have been proposed to test linkage between a susceptibility gene and a disease locus by using admixture-generated linkage disequilibrium (LD) for each of the genotyped markers. In a genome scan, admixture mapping usually tests 2,000 to 3,000 markers across the genome. Currently, either a very conservative Sidak (or Bonferroni) correction or a very time consuming simulation-based method is used to correct for the multiple tests and evaluate the overall p value. In this report, we propose a computationally efficient analytical approach for correction of the multiple tests and for calculating the overall p value for an admixture genome scan. Except for the Sidak (or Bonferroni) correction, our proposed method is the first analytical approach for correction of the multiple tests and for calculating the overall p value for a genome scan. Our simulation studies show that the proposed method gives correct overall type I error rates for genome scans in all cases, and is much more computationally efficient than simulation-based methods.

Computational Biology↗

A classical likelihood based approach for admixture mapping using EM algorithm.

Several disease-mapping methods have been proposed recently, which use the information generated by recent admixture of populations from historically distinct geographic origins. These methods include both classic likelihood and Bayesian approaches. In this study we directly maximize the likelihood function from the hidden Markov Model for admixture mapping using the EM algorithm, allowing for uncertainty in model parameters, such as the allele frequencies in the parental populations. We determined the robustness of the proposed method by examining the ancestral allele frequency estimate and individual marker-location specific ancestry when the data were generated by different population admixture models and no learning sample was used. The proposed method outperforms a widely used Bayesian MCMC strategy for data generated from various population admixture models. The multipoint information content for ancestry was derived based on the map provided by Smith et al. (2004) and the associated statistical power was calculated. We examined the distribution of admixture LD across the genome for both real and simulated data and established a threshold for genome wide significance applicable to admixture mapping studies. The software ADMIXPROGRAM for performing admixture mapping is available from authors.

Algorithms↗

A combinatorial searching method for detecting a set of interacting loci associated with complex traits.

Complex diseases are presumed to be the results of the interaction of several genes and environmental factors, with each gene only having a small effect on the disease. Mapping complex disease genes therefore becomes one of the greatest challenges facing geneticists. Most current approaches of association studies essentially evaluate one marker or one gene (haplotype approach) at a time. These approaches ignore the possibility that effects of multilocus functional genetic units may play a larger role than a single-locus effect in determining trait variability. In this article, we propose a Combinatorial Searching Method (CSM) to detect a set of interacting loci (may be unlinked) that predicts the complex trait. In the application of the CSM, a simple filter is used to filter all the possible locus-sets and retain the candidate locus-sets, then a new objective function based on the cross-validation and partitions of the multi-locus genotypes is proposed to evaluate the retained locus-sets. The locus-set with the largest value of the objective function is the final locus-set and a permutation procedure is performed to evaluate the overall p-value of the test for association between the final locus-set and the trait. The performance of the method is evaluated by simulation studies as well as by being applied to a real data set. The simulation studies show that the CSM has reasonable power to detect high-order interactions. When the CSM is applied to a real data set to detect the locus-set (among the 13 loci in the ACE gene) that predicts systolic blood pressure (SBP) or diastolic blood pressure (DBP), we found that a four-locus gene-gene interaction model best predicts SBP with an overall p-value = 0.033, and similarly a two-locus gene-gene interaction model best predicts DBP with an overall p-value = 0.045.

Blood Pressure↗

Haplotype sharing transmission/disequilibrium tests that allow for genotyping errors.

The present study introduces new Haplotype Sharing Transmission/Disequilibrium Tests (HS-TDTs) that allow for random genotyping errors. We evaluate the type I error rate and power of the new proposed tests under a variety of scenarios and perform a power comparison among the proposed tests, the HS-TDT and the single-marker TDT. The results indicate that the HS-TDT shows a significant increase in type I error when applied to data in which either Mendelian inconsistent trios are removed or Mendelian inconsistent markers are treated as missing genotypes, and the magnitude of the type I error increases both with an increase in sample size and with an increase in genotyping error rate. The results also show that a simple strategy, that is, merging each rare haplotype to a most similar common haplotype, can control the type I error inflation for a wide range of genotyping error rates, and after merging rare haplotypes, the power of the test is very similar to that without merging the rare haplotypes. Therefore, we conclude that a simple strategy may make the HS-TDT robust to genotyping errors. Our simulation results also show that this strategy may also be applicable to other haplotype-based TDTs.

Algorithms↗

A haplotype similarity based transmission/disequilibrium test under founder heterogeneity.

Taking advantage of increasingly available high-density single nucleotide polymorphisms (SNP) markers across the genome, various types of transmission/disequilibrium tests (TDT) using haplotype information have been developed. A practical challenge arising in such studies is the possibility that transmitted haplotypes have inherited disease-causing mutations from different ancestral chromosomes, or do not bear any disease-causing mutations (founder heterogeneity). To reduce the loss of signal strength due to founder heterogeneity, we propose an SP-TDT test that combines a sequential peeling procedure with the haplotype similarity based TDT method. The proposed SP-TDT method is applicable to any size of nuclear family with or without ambiguous phase information. Simulation studies suggest that the SP-TDT method has the correct type I error rate in stratified populations, and enhanced power compared with some existing haplotype similarity based TDT methods. Finally, we apply the proposed method to study the association of the leptin gene with obesity from the National Heart, Lung, and Blood Institute Family Heart Study.

Algorithms↗

Tests of association between quantitative traits and haplotypes in a reduced-dimensional space.

Candidate gene association tests are currently performed using several intragenic SNPs simultaneously, by testing SNP haplotype or genotype effects in multifactorial diseases or traits. The number of haplotypes drastically increases with an increase in the number of typed SNPs. As a result, large numbers of haplotypes will introduce large degrees of freedom in haplotype-based tests, and thus limit the power of the tests. In this study we propose using the principal component method to reduce the dimension, and then construct association tests on the lower-dimensional space to test the association between haplotypes and a quantitative trait using population-based samples. The proposed method allows ambiguous haplotypes. We use simulation studies to evaluate the type I error rate of the tests, and compare the power of the proposed tests with that of the tests without dimension reduction, and the tests with dimension reduction by merging rare haplotypes. The simulation results show that the proposed tests have correct type I error rates and are more powerful than other tests in most cases considered in our simulation studies.

Computer Simulation↗

Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes.

BACKGROUND: Due to the high cost and low reproducibility of many microarray experiments, it is not surprising to find a limited number of patient samples in each study, and very few common identified marker genes among different studies involving patients with the same disease. Therefore, it is of great interest and challenge to merge data sets from multiple studies to increase the sample size, which may in turn increase the power of statistical inferences. In this study, we combined two lung cancer studies using microarray GeneChip, employed two gene shaving methods and a two-step survival test to identify genes with expression patterns that can distinguish diseased from normal samples, and to indicate patient survival, respectively. RESULTS: In addition to common data transformation and normalization procedures, we applied a distribution transformation method to integrate the two data sets. Gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) were then applied separately to the joint data set for cancer gene selection. The two methods discovered 13 and 10 marker genes (5 in common), respectively, with expression patterns differentiating diseased from normal samples. Among these marker genes, 8 and 7 were found to be cancer-related in other published reports. Furthermore, based on these marker genes, the classifiers we built from one data set predicted the other data set with more than 98% accuracy. Using the univariate Cox proportional hazard regression model, the expression patterns of 36 genes were found to be significantly correlated with patient survival (p < 0.05). Twenty-six of these 36 genes were reported as survival-related genes from the literature, including 7 known tumor-suppressor genes and 9 oncogenes. Additional principal component regression analysis further reduced the gene list from 36 to 16. CONCLUSION: This study provided a valuable method of integrating microarray data sets with different origins, and new methods of selecting a minimum number of marker genes to aid in cancer diagnosis. After careful data integration, the classification method developed from one data set can be applied to the other with high prediction accuracy.

Adenocarcinoma↗

Transmission/disequilibrium test based on haplotype sharing for tightly linked markers.

Studies using haplotypes of multiple tightly linked markers are more informative than those using a single marker. However, studies based on multimarker haplotypes have some difficulties. First, if we consider each haplotype as an allele and use the conventional single-marker transmission/disequilibrium test (TDT), then the rapid increase in the degrees of freedom with an increasing number of markers means that the statistical power of the conventional tests will be low. Second, the parental haplotypes cannot always be unambiguously reconstructed. In the present article, we propose a haplotype-sharing TDT (HS-TDT) for linkage or association between a disease-susceptibility locus and a chromosome region in which several tightly linked markers have been typed. This method is applicable to both quantitative traits and qualitative traits. It is applicable to any size of nuclear family, with or without ambiguous phase information, and it is applicable to any number of alleles at each of the markers. The degrees of freedom (in a broad sense) of the test increase linearly as the number of markers considered increases but do not increase as the number of alleles at the markers increases. Our simulation results show that the HS-TDT has the correct type I error rate in structured populations and that, in most cases, the power of HS-TDT is higher than the power of the existing single-marker TDTs and haplotype-based TDTs.

Computer Simulation↗

On a semiparametric test to detect associations between quantitative traits and candidate genes using unrelated individuals.

Although genetic association studies using unrelated individuals may be subject to bias caused by population stratification, alternative methods that are robust to population stratification such as family-based association designs may be less powerful. Recently, various statistical methods robust to population stratification were proposed for association studies, using unrelated individuals to identify associations between candidate markers and traits of interest (both qualitative and quantitative). Here, we propose a semiparametric test for association (SPTA). SPTA controls for population stratification through a set of genomic markers by first deriving a genetic background variable for each sampled individual through his/her genotypes at a series of independent markers, and then modeling the relationship between trait values, genotypic scores at the candidate marker, and genetic background variables through a semiparametric model. We assume that the exact form of relationship between the trait value and the genetic background variable is unknown and estimated through smoothing techniques. We evaluate the performance of SPTA through simulations both with discrete subpopulation models and with continuous admixture population models. The simulation results suggest that our procedure has a correct type I error rate in the presence of population stratification and is more powerful than statistical association tests for family-based association designs in all the cases considered. Moreover, SPTA is more powerful than the Quantitative Similarity-Based Association Test (QSAT) developed by us under continuous admixture populations, and the number of independent markers needed by SPTA to control for population stratification is substantially fewer than that required by QSAT.

Analysis of Variance↗

Linkage and linkage disequilibrium mapping of genes influencing human obesity in chromosome region 7q22.1-7q35.

Linkage results suggest that the region of chromosome 7 containing the leptin gene cosegregates with extreme obesity; however, leptin coding region mutations are rare. To investigate whether the leptin flanking sequence and/or a larger 40-cM region (7q22.1-7q35) contributes to obesity, we genotyped individuals from 200 European American families segregating extreme obesity and normal weight (1,020 subjects) using 21 microsatellite markers and two single nucleotide polymorphisms (SNPs) and conducted nonparametric linkage (NPL) analyses. We also carried out transmission disequilibrium tests for 135 European American triads using 27 markers (including eight SNPs). Both quantitative (MERLIN-regress) and qualitative (GENEHUNTER and MERLIN-npl) analyses provided evidence for linkage for BMI (GENEHUNTER NPL = 2.98, 20 cM centromeric to leptin at the marker D7S692; MERLIN Z score = 3.56). Results for several other regions in 7q gave weak linkage. Transmission disequilibrium test (TDT) and quantitative TDT (and quantitative pedigree disequilibrium test) analyses suggest linkage disequilibrium near leptin and other regions of 7q. Our results suggested that there could be two or more genes in chromosome region 7q22.1-7q35 that influence obesity. A new region found by this study (D7S692-D7S523, 7q31.1) has the most consistent linkage results and could harbor obesity-related genes.

Adolescent↗

Linkage disequilibrium mapping with genotype data.

Linkage disequilibrium mapping has proven a powerful tool for locating disease genes. Although all existing linkage disequilibrium mapping methods implicitly assume that individual haplotypes can be inferred, only genotypes are directly observable in practice, and haplotypes cannot always be uniquely resolved based on genotype data. In this article, we propose a likelihood-based linkage disequilibrium mapping approach to analyzing multilocus genotype data arising from case-control studies. Results from extensive simulation studies suggest that this approach may be a useful tool to fine map disease genes using case-control data.

Chromosome Mapping↗

On a family-based haplotype pattern mining method for linkage disequilibrium mapping.

Linkage disequilibrium mapping is an important tool in disease gene mapping. Recently, Toivonen et al. [1] introduced a haplotype mining (HPM) method that is applicable to data consisting of unrelated high-risk and normal haplotypes. The HPM method orders haplotypes by their strength of association with trait values, and uses all haplotypes exceeding a given threshold of strength of association to predict the gene location. In this study, we extend the HPM method to pedigree data by measuring the strength of association between a haplotype and quantitative traits of interest using the Quantitative Pedigree Disequilibrium Test proposed by Zhang et al. [2]. This family-based HPM (F-HPM) method can incorporate haplotype information across a set of markers and allow both missing marker data and ambiguous haplotype information. We use a simulation procedure to evaluate the statistical significance of the patterns identified from the F-HPM method. When the F-HPM method is applied to analyze the sequence data from the seven candidate genes in the simulated data sets in the 12th Genetic Analysis Workshop, the association between genes and traits can be detected with high power, and the estimated locations of the trait loci are close to the true sites.

Chromosome Mapping↗