Search PubMed⌕ Search

Biomedical subjects

Jing Hua Zhao

Publications and source records attributed to Jing Hua Zhao.

11 recordsLinked to original sources

Pedigree-drawing with R and graphviz.

UNLABELLED: Two functions for pedigree-drawing available in R (http://www.r-project.org): plot.pedigree in kinship and pedtodot in gap are described. The latter requires graphviz (http://www.graphviz.org). They can produce many pedigree diagrams quickly into a single file, serving as alternatives to programs that only offer interactive use. AVAILABILITY: Packages kinship and gap are available from http://cran.r-project.org.

Computer Graphics↗

Integrated analysis of genetic data with R.

Genetic data are now widely available. There is, however, an apparent lack of concerted effort to produce software systems for statistical analysis of genetic data compared with other fields of statistics. It is often a tremendous task for end-users to tailor them for particular data, especially when genetic data are analysed in conjunction with a large number of covariates. Here, R (http://www.r-project.org), a free, flexible and platform-independent environment for statistical modelling and graphics is explored as an integrated system for genetic data analysis. An overview of some packages currently available for analysis of genetic data is given. This is followed by examples of package development and practical applications. With clear advantages in data management, graphics, statistical analysis, programming, internet capability and use of available codes, it is a feasible, although challenging, task to develop it into an integrated platform for genetic analysis; this will require the joint efforts of many researchers.

Algorithms↗

Selecting cases from nuclear families for case-control association analysis.

We examine the efficiency of a number of schemes to select cases from nuclear families for case-control association analysis using the Genetic Analysis Workshop 14 simulated dataset. We show that with this simulated dataset comparing all affected siblings with unrelated controls is considerably more powerful than all of the other approaches considered. We find that the test statistic is increased by almost 3-fold compared to the next best sampling schemes of selecting all affected sibs only from families with affected parents (AF aff), one affected sib with most evidence of allele-sharing from each family (SF), and all affected sibs from families with evidence for linkage (AF L). We consider accounting for biological relatedness of samples in the association analysis to maintain the correct type I error. We also discuss the relative efficiencies of increasing the ratio of unrelated cases to controls, methods to confirm associations and issues to consider when applying our conclusions to other complex disease datasets.

Case-Control Studies↗

Estimating haplotype relative risks on human survival in population-based association studies.

Association-based linkage disequilibrium (LD) mapping is an increasingly important tool for localizing genes that show potential influence on human aging and longevity. As haplotypes contain more LD information than single markers, a haplotype-based LD approach can have increased power in detecting associations as well as increased robustness in statistical testing. In this paper, we develop a new statistical model to estimate haplotype relative risks (HRRs) on human survival using unphased multilocus genotype data from unrelated individuals in cross-sectional studies. Based on the proportional hazard assumption, the model can estimate haplotype risk and frequency parameters, incorporate observed covariates, assess interactions between haplotypes and the covariates, and investigate the modes of gene function. By introducing population survival information available from population statistics, we are able to develop a procedure that carries out the parameter estimation using a nonparametric baseline hazard function and estimates sex-specific HRRs to infer gene-sex interaction. We also evaluate the haplotype effects on human survival while taking into account individual heterogeneity in the unobserved genetic and nongenetic factors or frailty by introducing the gamma-distributed frailty into the survival function. After model validation by computer simulation, we apply our method to an empirical data set to measure haplotype effects on human survival and to estimate haplotype frequencies at birth and over the observed ages. Results from both simulation and model application indicate that our survival analysis model is an efficient method for inferring haplotype effects on human survival in population-based association studies.

Aged↗

Haplotype association analysis of human disease traits using genotype data of unrelated individuals.

Haplotype inference has become an important part of human genetic data analysis due to its functional and statistical advantages over the single-locus approach in linkage disequilibrium mapping. Different statistical methods have been proposed for detecting haplotype - disease associations using unphased multi-locus genotype data, ranging from the early approach by the simple gene-counting method to the recent work using the generalized linear model. However, these methods are either confined to case - control design or unable to yield unbiased point and interval estimates of haplotype effects. Based on the popular logistic regression model, we present a new approach for haplotype association analysis of human disease traits. Using haplotype-based parameterization, our model infers the effects of specific haplotypes (point estimation) and constructs confidence interval for the risks of haplotypes (interval estimation). Based on the estimated parameters, the model calculates haplotype frequency conditional on the trait value for both discrete and continuous traits. Moreover, our model provides an overall significance level for the association between the disease trait and a group or all of the haplotypes. Featured by the direct maximization in haplotype estimation, our method also facilitates a computer simulation approach for correcting the significance level of individual haplotype to adjust for multiple testing. We show, by applying the model to an empirical data set, that our method based on the well-known logistic regression model is a useful tool for haplotype association analysis of human disease traits.

Computer Simulation↗

2LD, GENECOUNTING and HAP: Computer programs for linkage disequilibrium analysis.

UNLABELLED: Computer programs are introduced which calculate pair-wise linkage disequilibrium statistics and conduct haplotype frequency estimation, including X chromosome data, and using a heuristic algorithm to handle multiple genetic markers and missing data. AVAILABILITY: Programs 2LD, GENECOUNTING and HAP are available on Internet from http://www.hgmp.mrc.ac.uk/~jzhao and http://www.iop.kcl.ac.uk/IoP/Departments/PsychMed/GEpiBSt/software.shtml

Alcoholism↗

Linkage disequilibrium analysis of polymorphisms in the gene for myelin oligodendrocyte glycoprotein in Tourette's syndrome patients from a Chinese sample.

Gilles de la Tourette syndrome (GTS) is a neuropsychiatric disorder characterised by multiple motor and phonic tics, which wax and wane. Recently, evidence has accumulated supporting the role of autoimmune mechanisms in the aetiology of GTS, suggesting that it is within the paediatric autoimmune neuropsychiatric disorders associated with streptococcal infection (PANDAS) spectrum of childhood neurobehavioural disorders. An immunopathogenic role of antibodies against myelin oligodendrocyte glycoprotein (MOG) has been suggested in this syndrome. In this study, we investigate the association of three microsatellite polymorphisms (MOGa, MOGb, MOGc) in the gene for MOG with GTS in 197 family trios collected from southwest China. Linkage disequilibrium between these three markers was observed with the strongest between MOGa and MOGc (D' = 0.541, P = 0.000). We did not find overall significant evidence for distorted transmission of any of these three markers of MOG gene in GTS, although we observed a weak preferential transmission of the 148 bp allele of MOGc (chi(2) = 4.000, P = 0.046) which did not survive correction for multiple testing. Our results suggest that there is no association between the MOG gene polymorphisms we tested and GTS.

China↗

Generic number systems and haplotype analysis.

Three simple and elegant algorithms involving binary and mixed-radix numbers are presented as C subroutines and applied to gene-counting procedure. The first, a multikey radix-sorting subroutine, is used to tally individuals with similar genetic marker information. The second, a subroutine for N-ary number addition, is used to enumerate all possible phases of a heterozygote. The third, a mixed-radix number subroutine, is used to generate all haplotypes and indexing single array of haplotype frequencies. Examples exposing these algorithms are also given. The sorting algorithm entails broad application while the N-ary and mixed-radix number algorithms are very efficient for generic looping. Implementation of gene-counting using these algorithms avoids use of multilocus genotype identifier and improves its portability to other analysis.

Algorithms↗

Social environment, ethnicity and schizophrenia. A case-control study.

BACKGROUND: There is accumulating evidence that genetic and neurodevelopmental factors cannot solely account for the pathogenesis of schizophrenia. In view of the reportedly increased incidence of schizophrenia among the African-Caribbean population in Britain, we sought to establish the socio-environmental influences which distinguished African-Caribbean patients from white British and Asian patients with schizophrenia, as well as from normal population controls of the same community. METHOD: A matched case-control study was conducted in London between 1991 and 1993. Inclusion criteria for patients was a first onset psychosis between the ages of 18 and 64. Symptoms were recorded using the Present State Examination (PSE), and a research diagnosis of schizophrenia was made using the CATEGO program. Comparisons were made on a range of demographic and socio-environmental measures between patients (n = 100: 38 African-Caribbean, 38 white and 24 Asian) and the same number of normal controls. RESULTS: Three socio-environmental variables differentiated the African-Caribbean cases from their peers and their normal controls: unemployment, living alone and a long period of separation from either or both parents as a minor. Though all patients were much more likely than controls to be unemployed at first contact with the services (odds ratio 5.5, 95 % CI 2.59, 11.68), the odds ratio was highest among African-Caribbeans, and further conditional logistic regression analysis demonstrated that unemployment was significantly associated with the high rate of caseness among African-Caribbeans. However, the direction of cause and effect cannot be determined from this type of study. Despite the fact that African-Caribbean cases were more likely than their peers and same group controls to live alone (p < 0.05), this did not achieve significance using Fisher's Exact Test. Separation from both parents in childhood distinguished African-Caribbean cases from their controls and from cases and controls of the other ethnic groups (odds ratio 5.0, 95 % CI 1.09, 22.82). This event cannot be attributed to the premorbid manifestations of schizophrenia, nor to psychoses in the parents, and hence is a possible explanatory factor for the high incidence of schizophrenia among African-Caribbeans in Britain. CONCLUSIONS: These findings indicate that unemployment and early separation from both parents distinguish African-Caribbeans diagnosed with schizophrenia from their counterparts of other ethnic groups as well as their normal peers, and imply that more attention needs to be focussed on socio-environmental variables in schizophrenia research.

Adult↗

Faster haplotype frequency estimation using unrelated subjects.

Linkage disequilibrium (LD) between tightly linked loci provides fine mapping information of disease-predisposing allelic variants. The most common method of LD analysis involves unrelated cases and controls. We have previously proposed model-free and permutation tests for diseases with unknown mode of inheritance that can be applied to several highly polymorphic loci. However, performing such analyses remained computer intensive. In this report we propose a speed-up of both the gene-counting procedure and the permutation procedure. We demonstrate the improved method with an analysis of schizophrenia and human leucocyte antigen markers, and an analysis of alcoholism and mitochondrial aldehyde dehydrogenase markers. Our implementation also allows the rapid calculation of permutation-based LD measures and related statistics.

Alcoholism↗