Search PubMed⌕ Search

Biomedical subjects

Hsin-Chou Yang

Publications and source records attributed to Hsin-Chou Yang.

8 recordsLinked to original sources

PDA: Pooled DNA analyzer.

BACKGROUND: Association mapping using abundant single nucleotide polymorphisms is a powerful tool for identifying disease susceptibility genes for complex traits and exploring possible genetic diversity. Genotyping large numbers of SNPs individually is performed routinely but is cost prohibitive for large-scale genetic studies. DNA pooling is a reliable and cost-saving alternative genotyping method. However, no software has been developed for complete pooled-DNA analyses, including data standardization, allele frequency estimation, and single/multipoint DNA pooling association tests. This motivated the development of the software, 'PDA' (Pooled DNA Analyzer), to analyze pooled DNA data. RESULTS: We develop the software, PDA, for the analysis of pooled-DNA data. PDA is originally implemented with the MATLAB language, but it can also be executed on a Windows system without installing the MATLAB. PDA provides estimates of the coefficient of preferential amplification and allele frequency. PDA considers an extended single-point association test, which can compare allele frequencies between two DNA pools constructed under different experimental conditions. Moreover, PDA also provides novel chromosome-wide multipoint association tests based on p-value combinations and a sliding-window concept. This new multipoint testing procedure overcomes a computational bottleneck of conventional haplotype-oriented multipoint methods in DNA pooling analyses and can handle data sets having a large pool size and/or large numbers of polymorphic markers. All of the PDA functions are illustrated in the four bona fide examples. CONCLUSION: PDA is simple to operate and does not require that users have a strong statistical background. The software is available at http://www.ibms.sinica.edu.tw/%7Ecsjfann/first%20flow/pda.htm.

Algorithms↗

A comparison of major histocompatibility complex SNPs in Han Chinese residing in Taiwan and Caucasians.

Genetic dissection of complex diseases is both important and challenging. The human major histocompatibility complex is involved in many human diseases and genetic mechanisms. This highly polymorphic chromosome region has been extensively studied in Caucasians but not as well in Asians. Thus, we compared genotypic distributions, linkage disequilibria and haplotype blocks between Caucasian and Taiwan's Han Chinese populations. Moreover, we investigated the population admixture and phylogenetic system in Han Chinese residing in Taiwan. The results show that Taiwan's Han Chinese differ drastically in genotypic information compared with Caucasians but are relatively homogeneous among the three major ethnic subgroups, Minnan, Hakka and Mainlanders. Differences in allele frequency (AF) between Taiwanese and Caucasians in some disease-associated loci may reveal clues to differences in disease prevalence. The results of ethnic heterogeneity imply that public databases should be used with caution in cases where the study population(s) differs from the population characterized in the database. The high homogeneity we observed among the Taiwanese subpopulations mitigates the possibility of spurious association caused by ignoring population stratification in Taiwanese disease gene association studies. These results are useful for understanding our genetic background and designing future disease gene mapping studies.

Asian People↗

A sliding-window weighted linkage disequilibrium test.

Multilocus linkage disequilibrium (LD) tests that consider inter-marker (LD) are more powerful than single-locus tests when disease etiology is contributed simultaneously by several linked and correlated loci. However, inclusion of redundant non-informative markers may result in reduced testing power and/or inflated false-positive rate, therefore selection of proper marker sets is important in such tests. We introduce a unified LD test based on a convenient marker-selection procedure (sliding window) combined with an adjustment approach (marker weighting) to dilute the impact of nuisance markers on tests. The proposed procedure includes several conventional p-value combination methods as its special cases. Simulation studies were performed to evaluate the impact of inclusion of nuisance markers and performance of the procedure. The results showed that testing power was often inversely proportional to the quantity of nuisance markers. Among a class of p-value combination methods, the product p-value method had the highest testing power. P-value truncation somewhat reduced the testing power but controlled the false-positive rate well. Compared with conventional unweighted approaches, the weighted strategy alleviated the false-positive rate and/or increased testing power when nuisance markers were included. Analyses of two authentic data sets for psoriasis and Alzheimer's disease using our proposed method confirmed previous findings.

Alleles↗

A genome-wide scanning and fine mapping study of COGA data.

A thorough genetic mapping study was performed to identify predisposing genes for alcoholism dependence using the Collaborative Study on the Genetics of Alcoholism (COGA) data. The procedure comprised whole-genome linkage and confirmation analyses, single locus and haplotype fine mapping analyses, and gene x environment haplotype regression. Stratified analysis was considered to reduce the ethnic heterogeneity and simultaneously family-based and case-control study designs were applied to detect potential genetic signals. By using different methods and markers, we found high linkage signals at D1S225 (253.7 cM), D1S547 (279.2 cM), D2S1356 (64.6 cM), and D7S2846 (56.8 cM) with nonparametric linkage scores of 3.92, 4.10, 4.44, and 3.55, respectively. We also conducted haplotype and odds ratio analyses, where the response was the dichotomous status of alcohol dependence, explanatory variables were the inferred individual haplotypes and the three statistically significant covariates were age, gender, and max drink (the maximum number of drinks consumed in a 24-hr period). The final model identified important AD-related haplotypes within a candidate region of NRXN1 at 2p21 and a few others in the inter-gene regions. The relative magnitude of risks to the identified risky/protective haplotypes was elucidated.

Alcoholism↗

Breast cancer risk associated with genotypic polymorphism of the mitosis-regulating gene Aurora-A/STK15/BTAK.

Aneuploidy, an abnormal number of chromosomes, is relatively common and occurs early in breast cancer development. This observation supports a breast tumorigenic contribution of mechanisms responsible for maintaining chromosome number stability in which centrosomes play an essential role. We therefore speculated that the Aurora-A/STK15/BTAK gene, implicated in the regulation of centrosome duplication, may be associated with breast tumorigenesis. To test this hypothesis, we conducted a case-control study of 709 primary breast cancer patients and 1,972 healthy controls, examining single-nucleotide polymorphisms (SNPs), including a suggested functional Phe31Ile SNP, in Aurora-A. We were also interested in knowing whether any association between Aurora-A and breast cancer was modified by reproductive risk factors reflecting susceptibility to estrogen exposure. Our hypothesis is that, since estrogen is known to promote breast cancer development via its mitogenic effect leading to malignant proliferation on breast epithelium and since Aurora-A is involved in regulating mitosis, the discovery of a joint effect between the Aurora-A genotype and reproductive risk factors on cancer risk might yield valuable clues to the association of breast tumorigenesis with estrogen. Support for this hypothesis came from the following observations. (i) Two SNPs in Aurora-A were significantly associated with breast cancer risk (p < 0.05). (ii) Haplotype analyses, based on different combinations of multiple SNPs in Aurora-A, revealed a strong association with breast cancer risk; interestingly, the genotypic distribution of the suggested functional Phe31Ile SNP was not significantly different between breast cancer patients and controls, but the specific haplotype containing the putative at-risk Ile allele was more common in patients. (iii) This association between risk and putative high-risk genotypes was stronger and more significant in women thought to be more susceptible to estrogen, i.e., those with a longer interval between menarche and first full-term pregnancy. (iv) The protective effect conferred by a history of full-term pregnancy was significant only in women with a putative low-risk genotype of Aurora-A. Our study provides new findings supporting the mutator role of Aurora-A in breast cancer development, suggesting that breast cancer could be driven by genomic instability associated with variant Aurora-A, the tumorigenic contribution of which could be enhanced as a result of increased mitosis due to estrogen exposure.

Aurora Kinase A↗

Modeling animals' behavioral response by Markov chain models for capture-recapture experiments.

A bivariate Markov chain approach that includes both enduring (long-term) and ephemeral (short-term) behavioral effects in models for capture-recapture experiments is proposed. The capture history of each animal is modeled as a Markov chain with a bivariate state space with states determined by the capture status (capture/noncapture) and marking status (marked/unmarked). In this framework, a conditional-likelihood method is used to estimate the population size and the transition probabilities. The classical behavioral model that assumes only an enduring behavioral effect is included as a special case of the bivariate Markovian model. Another special case that assumes only an ephemeral behavioral effect reduces to a univariate Markov chain based on capture/noncapture status. The model with the ephemeral behavioral effect is extended to incorporate time effects; in this model, in contrast to extensions of the classical behavioral model, all parameters are identifiable. A data set is analyzed to illustrate the use of the Markovian models in interpreting animals' behavioral response. Simulation results are reported to examine the performance of the estimators.

Animals↗

New adjustment factors and sample size calculation in a DNA-pooling experiment with preferential amplification.

In the post-genome era, disease gene mapping using dense genetic markers has become an important tool for dissecting complex inheritable diseases. Locating disease susceptibility genes using DNA-pooling experiments is a potentially economical alternative to those involving individual genotyping. The foundation of a successful DNA-pooling association test is a precise and accurate estimation of allele frequency. In this article, we propose two new adjustment methods that correct for preferential amplification of nucleotides when estimating the allele frequency of single-nucleotide polymorphisms. We also discuss the effect of sample size when calibrating unequal allelic amplification. We conducted simulation studies to assess the performance of different adjustment procedures and found that our proposed adjustments are more reliable with respect to the estimation bias and root mean square error compared with the current approach. The improved performance not only improves the accuracy and precision of allele frequency estimations but also leads to more powerful disease gene mapping.

Algorithms↗

Estimation of the size of an open population using local estimating equations II: a partially parametric approach.

Kernel smoothing methods are applied to extend a modification of the closed population approach of Lloyd and Yip (1991, in Estimating Equations, 65-88) to open populations with frequent capture occasions. The method complements previous nonparametric methods and, when the parametric assumptions are met, simulations show the new method has a smaller integrated mean squared error than the previous fully nonparametric method. The method is applied to capture-recapture data on short-tailed shearwaters collected annually for 48 years.

Animal Identification Systems↗