Search PubMed⌕ Search

Biomedical subjects

Ao Yuan

Publications and source records attributed to Ao Yuan.

8 recordsLinked to original sources

Detecting disease gene in DNA haplotype sequences by nonparametric dissimilarity test.

Association studies for complex diseases based on haplotype data have received increasing attention in the last few years. A commonly used nonparametric method, which takes haplotype structure into consideration, is to use the U-statistic to compare the similarities between genetic compositions in the case and control populations. Although the method and its variants are convenient to use in practice, there are some areas where the tests cannot detect even large differences between cases and controls. To overcome this problem and enhance the power, we propose a new form of the weighted U-statistic, which directly compares the dissimilarity between the haplotype structures in the case and control populations. We show that this test statistic is asymptotically a linear combination of the absolute values of normal random variables under the null hypothesis, and shifts strictly toward the right under the alternative, and therefore has no blind areas of detection. Simulation studies indicate that our test statistic overcomes the weakness of the existing ones and is robust and powerful as well.

Base Sequence↗

Variance components model with disequilibria.

The variance components (VC) model has been popular for genetic analysis. It has received wide applications in a variety of genetic practices, and been extended to various forms for different settings. However, most of the existing VC models are, explicitly or implicitly, under the assumption of the Hardy-Weinberg and/or linkage equilibria, which is impractical in some realistic settings since more or less deviations from this assumption are common. We propose a new VC model that incorporates both these disequilibria, and includes the existing models as special cases. The corresponding variance components are computed for some commonly used relative pairs conditional on the observed marker identity-by-descent data. Parameters can be estimated by the traditional methods such as the maximum likelihood estimate. Simulation studies suggest that this extended model improves inference significantly over the existing models when deviations of these disequilibria are present.

Analysis of Variance↗

A study of genetic association with electrophysiological measures related to alcoholism: GAW14 data.

Recently, alcohol-related traits have been shown to have a genetic component. Here, we study the association of specific genetic measures in one of the three sets of electrophysiological measures in families with alcoholism distributed as part of the Genetic Analysis Workshop 14 data, the NTTH (non-target case of Visual Oddball experiment for 4 electrode placements) phenotypes: ntth1, ntth2, ntth3, and ntth4. We focused on the analysis of the 786 Affymetrix markers on chromosome 4. Our desire was to find at least a partial answer to the question of whether ntth1, ntth2, ntth3, and ntth4 are separately or jointly genetically controlled, so we studied the principal components that explain most of the covariation of the four quantitative traits. The first principal component, which explains 70% of the covariation, showed association but not genetic linkage to two markers: tsc0272102 and tsc0560854. On the other hand, ntth1 appeared to be the trait driving the variation in the second principal component, which showed association and genetic linkage at markers in four regions: tsc0045058, tsc1213381, tsc0055068, and tsc0051777 at map distances 53.26, 85.42, 89.31, and 172.86, respectively. These results show that the partial answer to our starting question for this brief analysis is that the NTTH phenotypes are not jointly genetically controlled. The component ntth1 displays marked genetic linkage.

Alcoholism↗

Genome scan linkage analysis comparing microsatellites and single-nucleotide polymorphisms markers for two measures of alcoholism in chromosomes 1, 4, and 7.

BACKGROUND: We analyzed 143 pedigrees (364 nuclear families) in the Collaborative Study on the Genetics of Alcoholism (COGA) data provided to the participants in the Genetic Analysis Workshop 14 (GAW14) with the goal of comparing results obtained from genome linkage analysis using microsatellite and with results obtained using SNP markers for two measures of alcoholism (maximum number of drinks -MAXDRINK and an electrophysiological measure from EEG -TTTH1). First, we constructed haplotype blocks by using the entire set of single-nucleotide polymorphisms (SNP) in chromosomes 1, 4, and 7. These chromosomes have shown linkage signals for MAXDRINK or EEG-TTTH1 in previous reports. Second, we randomly selected one, two, three, four, and five SNPs from each block (referred to as Rep1 - Rep5, respectively) to conduct linkage analysis using variance component approach. Finally, results of all SNP analyses were compared with those obtained using microsatellite markers. RESULTS: The LOD scores obtained from SNPs were slightly higher but the curves were not radically different from those obtained from microsatellite analyses. The peaks of linkage regions from SNP sets were slightly shifted to the left when compared to those from microsatellite markers. The reduced sets of SNPs provide signals in the same linkage regions but with a smaller LOD score suggesting a significant impact of the decrease in information content on linkage results. The widths of 1 LOD support interval of linkage regions from SNP sets were smaller when compared to those of microsatellite markers. However, two linkage regions obtained from the microsatellite linkage analysis on chromosome 7 for LOG of TTTH1 were not detected in the SNP based analyses. CONCLUSION: The linkage results from SNPs showed narrower linkage regions and slightly higher LOD scores when compared to those of microsatellite markers. The different builds of the genetic maps used in microsatellite and SNPs markers or/and errors in genotyping may account for the microsatellite linkage signals on chromosome 7 that were not identified using SNPs. Also, unresolved map issues between SNPs and microsatellite markers may be partly responsible for the shifted linkage peaks when comparing the two types of markers.

Alcoholism↗

A statistical framework for haplotype block inference.

The existence of haplotype blocks transmitted from parents to offspring has been suggested recently. This has created an interest in the inference of the block structure and length. The motivation is that haplotype blocks that are characterized well will make it relatively easier to quickly map all the genes carrying human diseases. To study the inference of haplotype block systematically, we propose a statistical framework. In this framework, the optimal haplotype block partitioning is formulated as the problem of statistical model selection; missing data can be handled in a standard statistical way; population strata can be implemented; block structure inference/hypothesis testing can be performed; prior knowledge, if present, can be incorporated to perform a Bayesian inference. The algorithm is linear in the number of loci, instead of NP-hard for many such algorithms. We illustrate the applications of our method to both simulated and real data sets.

Algorithms↗

Identifying the susceptibility gene(s) in a set of trait-linked genes using genotype data.

There are generally three steps to isolate a disease linkage-susceptibility gene: genome-wide scan, fine mapping, and, last, positional cloning. The last step is time consuming and involves intensive laboratory work. In some cases, fine mapping cannot proceed further on a set of markers because they are tightly linked. For years, genetic statisticians have been trying different ways to narrow the fine-mapping results to provide some guidance for the next step of laboratory work. Although these methods are practical and efficient, most of them are based on IBD data, which usually can be inferred only from the genotype data with some uncertainty. The corresponding methods thus have no greater power than one using genotype data directly. Also, IBD-based methods apply only to relative pair data. Here, using genotype data, we have developed a statistical hypothesis-testing method to pinpoint a SNP, or SNPs, suspected of responsibility for a disease trait linkage among a set of SNPs tightly linked in a region. Our method uses genotype data of affected individuals or case-control studies, which are widely available in the laboratory. The testing statistic can be constructed using any genotype-based disease-marker disequilibrium measure and is asymptotically distributed as a chi-square mixture. This method can be used for singleton data, relative pair data, or general pedigree data. We have applied the method to simulated data as well as a real data set; it gives satisfactory results.

Chromosome Mapping↗

Exact test of Hardy-Weinberg equilibrium by Markov chain Monte Carlo.

The assumption of Hardy-Weinberg equilibrium (HWE) among alleles is of fundamental importance in genetic studies. There are numerous testing methods for it using genotype counts data. The exact test is used when the sample size is not large enough for asymptotic approximations. There are several numerical methods to carry out this test, such as complete enumeration, Monte Carlo and Markov chain Monte Carlo simulations. Complete enumeration is impractical in many applications, especially when the table counts are large. The Monte Carlo method is simple to use but still difficult when the table counts become large. The Markov chain Monte Carlo method, by sampling a sub-table each time, is suitable for this latter situation. Based on switches among a few (no more than four) cells, the existing Markov chain samplers are highly dependent and inefficient for large tables. Here we consider a new Markov chain sampling, in which a sub-table of user-specified size is updated at each iteration. The resulting chain is less dependent, and the sampling is flexible and efficient. The conventional test for HWE is based on a few test statistics, such as the likelihood and the chi-squared statistic. To expand the family of test statistics, we consider a class of divergence measures for the departure of HWE. Examples are given as illustrations.

Alleles↗

Two new recursive likelihood calculation methods for genetic analysis.

Recursive likelihood calculations for genetic analysis with ungenotyped pedigree data employ variations of the Elston-Stewart (ES) or the Lander-Green (LG) algorithms. With the ES algorithm, the number of loci may be limited but not the pedigree size. With the LG algorithm, the reverse is the case. We introduce two new algorithms for the computation of regressive likelihoods for pedigrees with multivariate traits. The first is an alternative formulation of our existing model, which leads to a simpler form in the binary trait, polygenic and mixed model cases. The second is an approximation model, which is computationally efficient. These methods apply to both continuous and binary traits, in the oligogenic and polygenic cases. Both methods coincide in the binary case. We considered these methods for cases in which all the traits are controlled by a single locus, with each trait controlled by one locus independent to the others. Simulation studies and analysis of a real data are presented for segregation analysis as illustrations. These methods can also be used in other model-based analyses. These methods are implemented in G.E.M.S., the genetic epidemiology models software.

Animals↗