Search PubMedSearch

Biomedical subjects

L P Zhao

Publications and source records attributed to L P Zhao.

At least 19 recordsLinked to original sources

An efficient, robust and unified method for mapping complex traits (III): combined linkage/linkage-disequilibrium analysis.

Extending the method for linkage analysis [Zhao et al., 1998a: Am. J. Med. Genet. 77:366-383; 1998b: Am. J. Med. Genet. 79:49-61], this article describes a method for the linkage-disequilibrium analysis, and for combining linkage and linkage-disequilibrium analyses. As highly dense markers are increasingly used in genome scans, one or more markers are not only linked with the disease genes if they exist, but also likely in linkage-disequilibrium with those putative genes. Hence, linkage-disequilibrium analysis potentially offers additional information about positions of putative disease genes. Combining both linkage and linkage-disequilibrium signals, this approach is able to improve positional signals. As before, the proposed method is a model-based approach, but semiparametric via the estimating equation technique. Under the assumptions of penetrance and allele frequency, this method efficiently estimates recombination fractions for linkage analysis and odds ratios for linkage-disequilibrium analysis. As described in two previous papers, this method is relatively more robust than the lod score methods, since it requires weaker assumption than conditional independence. While the estimated recombination fractions are used for inference as part of linkage analysis, the estimated odds ratios are used for linkage-disequilibrium inference and combined linkage, and linkage-disequilibrium parameters can be used to test combined linkage/linkage-disequilibrium analysis. This approach has been implemented, named gSCAN, and its compiled version is available for trial on request via the web site (http:/lynx.fhcrc.org/qge). We applied this new approach to affected sib-pair data collected for the genome scan to localize type 1 diabetes genes. Under an assumed autosomal dominant gene model, the linkage analysis confirms an earlier suggestion of one major gene around D6S281. Interestingly, the linkage-disequilibrium analysis suggests several additional signals around D6S250, GATA30, D6S311, D6S441, D6S442, D6S415, D6S411, D6S305, and a290xh9. The linkage analysis, on the other hand, suggests a signal around D6S281, while providing supporting evidence for several other marker loci. However, the combined analysis did not provide strong support for any of the findings, implying that linkage and linkage-disequilibrium findings are not consistent.

Alleles

On the assessment of statistical significance in disease-gene discovery.

One of the major challenges facing genome-scan studies to discover disease genes is the assessment of the genomewide significance. The assessment becomes particularly challenging if the scan involves a large number of markers collected from a relatively small number of meioses. Typically, this assessment has two objectives: to assess genomewide significance under the null hypothesis of no linkage and to evaluate true-positive and false-positive prediction error rates under alternative hypotheses. The distinction between these goals allows one to formulate the problem in the well-established paradigm of statistical hypothesis testing. Within this paradigm, we evaluate the traditional criterion of LOD score 3.0 and a recent suggestion of LOD score 3.6, using the Monte Carlo simulation method. The Monte Carlo experiments show that the type I error varies with the chromosome length, with the number of markers, and also with sample sizes. For a typical setup with 50 informative meioses on 50 markers uniformly distributed on a chromosome of average length (i.e., 150 cM), the use of LOD score 3.0 entails an estimated chromosomewide type I error rate of.00574, leading to a genomewide significance level >.05. In contrast, the corresponding type I error for LOD score 3.6 is.00191, giving a genomewide significance level of slightly <.05. However, with a larger sample size and a shorter chromosome, a LOD score between 3.0 and 3.6 may be preferred, on the basis of proximity to the targeted type I error. In terms of reliability, these two LOD-score criteria appear not to have appreciable differences. These simulation experiments also identified factors that influence power and reliability, shedding light on the design of genome-scan studies.

Chromosome Mapping

An efficient, robust, and unified method for mapping complex traits (II): multipoint linkage analysis.

Extending the method for two-point linkage analysis [Zhao et al., 1998: Am J Med Genet 77:366-383], this paper introduces a semiparametric method for multipoint linkage analysis, expected to gain efficiency by using multiple markers simultaneously. Overcoming the longstanding statistical and computational challenge to the parametric approaches (or lod score methods) for multipoint linkage analysis, this semiparametric approach, based on the estimating equation technique, yields statistically efficient and yet robust estimates and enjoys the computational efficiency in processing multiple markers from large pedigrees. Its computational burden increases linearly with the sizes of pedigrees and with the number of marker loci. To illustrate this semiparametric method, we apply it to marker data gathered for the Breast Cancer Consortium. The result supports the earlier finding of the positive linkage with BRCA1 and has also shown that the multipoint linkage analysis has an improved power. In addition, we have applied this method to analyze genome scanning data that have been used to localize genes responsible for type 1 diabetes. In support of the earlier findings, the genome scanning detects the linkage signals on chromosome 6 but does not support the earlier suggestions of two major genes in that genome segment. Through sensitivity analysis, it appears that the results are robust to misspecification of penetrance and allele frequency.

Automation

Embryonic cardiomyocyte hypoplasia and craniofacial defects in G alpha q/G alpha 11-mutant mice.

Heterotrimeric G proteins of the Gq class have been implicated in signaling pathways regulating cardiac growth under physiological and pathological conditions. Knockout mice carrying inactivating mutations in both of the widely expressed G alpha q class genes, G alpha q and G alpha 11, demonstrate that at least two active alleles of these genes are required for extrauterine life. Mice carrying only one intact allele [G alpha q(-/+);G alpha 11(-/-) or G alpha q(-/-);G alpha 11(-/+)] died shortly after birth. These mutants showed a high incidence of cardiac malformation. In addition, G alpha q(-/-);G alpha 11(-/+) newborns suffered from craniofacial defects. Mice lacking both G alpha q and G alpha 11 [G alpha q(-/-);G alpha 11(-/-)] died at embryonic day 11 due to cardiomyocyte hypoplasia. These data demonstrate overlap in G alpha q and G alpha 11 gene functions and indicate that the Gq class of G proteins plays a crucial role in cardiac growth and development.

Alleles

Efficient, robust, and unified method for mapping complex traits (I): two-point linkage analysis.

The completion of a preliminary human genome map and development of molecular methods have enabled researchers to assay a large number of polymorphic markers that are evenly spaced along the entire human genome. Among many applications, marker data are valuable for mapping complex traits through linkage or linkage-disequilibrium analysis, the former of which is the focus of this paper, the first in a series on this subject. Formalizing the concept and computation for linkage analysis, Elston and Stewart [1971; Human Heredity 21:523-542] introduced a likelihood function to capture relevant genetic information and a recursive algorithm for computing the likelihood function. However, the computing burden is prohibitive in processing complex pedigrees. Since that fundamental development, improving the computational algorithm and extending the method has been a dynamic area of research. The primary objective of this communication is to introduce a semiparametric method for linkage analysis. It is a particularly suitable approach with desirable properties for mapping complex traits that may be binary, continuous, and partially observed (i.e., censored). It incorporates candidate genes, environmental factors, and their interactions with the putative gene and is expected to be robust and efficient in comparison with likelihood-based methods. The properties of the estimates have been studied in finite samples with a limited simulation study. This method is illustrated with an application to family data contributed to the Breast Cancer Consortium.

Algorithms

Mapping of complex traits by single-nucleotide polymorphisms.

Molecular geneticists are developing the third-generation human genome map with single-nucleotide polymorphisms (SNPs), which can be assayed via chip-based microarrays. One use of these SNP markers is the ability to locate loci that may be responsible for complex traits, via linkage/linkage-disequilibrium analysis. In this communication, we describe a semiparametric method for combined linkage/linkage-disequilibrium analysis using SNP markers. Asymptotic results are obtained for the estimated parameters, and the finite-sample properties are evaluated via a simulation study. We also applied this technique to a simulated genome-scan experiment for mapping a complex trait with two major genes. This experiment shows that separate linkage and linkage-disequilibrium analyses correctly detected the signals of both major genes; but the rates of false-positive signals seem high. When linkage and linkage-disequilibrium signals were combined, the analysis yielded much stronger and clearer signals for the presence of two major genes than did two separate analyses.

Biosensing Techniques

An efficient protocol for rare mutation genotyping in a large population.

We introduce a method to efficiently detect rare mutations for individual subjects in a large population by pooling samples and retesting subgroups of positive pooled samples. We conducted computer simulations of this method and discovered that it seems efficient for mutation prevalences less than 0.1, regardless of the number of samples. The simulations also indicate that splitting the pooled samples into three to five subgroups at each level is optimal. The expected number of necessary tests and relative efficiency of this method are given, by mutation prevalence and sample size.

Computer Simulation

Population-based family study designs: an interdisciplinary research framework for genetic epidemiology.

Most complex traits such as cancer and coronary heart diseases are attributed either to heritable factors or to environmental factors or to both. Dissecting the genetic and environmental etiology of complex traits thus requires an interdisciplinary research strategy. Genetic studies generally involve families and investigate familial aggregations of traits, segregation of major disease genes, and locations of disease genes on the human genome, the latter of which can be identified via linkage analysis. Epidemiologic studies often use population-based case-control studies to establish the role of specific environmental factors. Integrating both objectives, genetic epidemiology is to assess the associations of environmental factors with disease status, to quantify the aggregation of cases within families, to characterize putative disease genes via segregation analysis, and to localize disease genes via linkage analysis with genetic markers. To accomplish these objectives through designed studies, we propose a class of population-based family study designs, which are formed by choosing among sampling designs at three stages. The objectives of sampling at these three stages are 1) combined aggregation and association analysis, 2) combined segregation, aggregation, and association analysis, and 3) combined linkage, segregation, aggregation, and association analysis. These designs form an interdisciplinary research framework for genetic epidemiology. Our preliminary exploration of this framework and related analytic methods indicates that population-based family study designs retain the efficiency of linkage analysis for localizing disease genes without losing the property of being population-based, and they will therefore allow an assessment of a joint contribution of genetic and environmental factors to complex traits.

Case-Control Studies

A population-based family study (II): Segregation analysis.

We used segregation analysis to investigate the genetic etiology of the disease in Problem 2A. Under the assumption of a dominant major gene, our analysis suggests a major gene with relative risk of 58 and an allele frequency of 0.013. Under an additive gene assumption, it appears that there may be two genes with relative risks of 39 and 17 and allele frequencies of 0.015 and 0.075, respectively.

Case-Control Studies

Family history and risk of colorectal cancer in the multiethnic population of Hawaii.

Increased risk of colorectal cancer in individuals with family history of the disease has been observed consistently in past studies. However, limited attention has been given to the influence of ethnicity, the characteristics of the proband's tumor, and kinship. A population-based case-control study was conducted between 1987 and 1991 in Hawaii among 1,192 incident colorectal cancer cases and 1,192 sex-, age-, and ethnicity-matched population controls. The study identified 7,673 relatives for the cases and 7,823 relatives for the controls. With an estimating equation-based regression method, relatives of cases were found to have a 2.5-fold increased risk of colorectal cancer compared with relatives of controls (95% confidence interval (CI) 1.8-3.4) after adjustment for covariates. This increase in risk was greater for Japanese (odds ratio (OR) = 3.0, 95% CI 1.7-5.4) than Caucasians (OR = 1.8, 95% CI 1.2-2.9), for siblings (OR = 3.1, 95% CI 2.1-4.6) than parents (OR = 2.0, 95% CI 1.1-3.1), and when the index patient was diagnosed before the age of 55 years (OR = 4.1, 95% CI 2.1-8.0) with multiple tumors (OR = 9.5, 95% CI 4.4-20.6), with a distant stage (OR = 4.6, 95% CI 2.7-7.8), or with cancer of the right colon (OR = 3.0, 95% CI 2.0-4.4) or the rectum (OR = 3.0, 95% CI 1.8-4.8). The increase in risk was not affected by the relative's sex. Relatives of cases were not at increased risk for other common cancers. It is estimated that approximately 11.1% and 6.5% of colorectal cancers are attributable to a first degree family history of the disease for Japanese and Caucasians, respectively. These data and those of previous studies strongly suggest that individuals with a family history of colorectal cancer in a first degree relative are at increased risk for the disease and should receive regular diagnostic screening. Characteristics of the index case, such as age and stage at diagnosis, subsite and number of tumors, and race, as well as kinship, may be important in assessing the colorectal cancer risk of a relative.

Aged

Cloning and characterization of myr 6, an unconventional myosin of the dilute/myosin-V family.

We have isolated cDNAs encoding a second member of the dilute (myosin-V) unconventional myosin family in vertebrates, myr 6 (myosin from rat 6). Expression of myr 6 transcripts in the brain is much more limited than is the expression of dilute, with highest levels observed in choroid plexus and components of the limbic system. We have mapped the myr 6 locus to mouse chromosome 18 using an interspecific backcross. The 3' portion of the myr 6 cDNA sequence from rat is nearly identical to that of a previously published putative glutamic acid decarboxylase from mouse [Huang, W.M., Reed-Fourquet, L., Wu, E. & Wu, J.Y. (1990) Proc. Natl. Acad. Sci. USA 87, 8491-8495].

Amino Acid Sequence

Estimating relative risk functions in case-control studies using a nonparametric logistic regression.

The authors describe an approach to the analysis of case-control studies in which the exposure variables are continuous, i.e., quantitative variables, and one wishes neither to categorize levels of the exposure variable nor to assume a log-linear relation between level of exposure and disease risk. A dose-response association of an exposure variable with a disease outcome can be depicted by estimated relative risks at various exposure levels, and the functional relation between exposure dose and disease risk is here termed a relative risk function (RRF). A RRF takes values that are greater than zero: Values less than one imply lower risk; the value one implies no risk, and values greater than one imply increased risk, when compared with a reference value. The authors describe how a nonparametric logistic regression can be used to estimate and display these RRFs. Using data from a previously published case-control study of diet and colon cancer, RRFs for total energy, dietary fiber, and alcohol intakes are compared with the original results obtained from using categorized levels of exposure variables. For total energy and alcohol intakes, there were meaningful differences in study results based on the two analytic approaches. For energy, the nonparametric logistic regression detected a significant protective effect of low intakes, which was not found in the original analysis. For alcohol, the nonparametric logistic regression suggested that there were two underlying populations, non- or very light drinkers and moderate to heavy drinkers, with different relation of dose to disease risk. In contrast, the original analysis found a nonlinear increase in risk across intake categories and did not detect the complex, bimodal nature of the exposure distribution. These results demonstrate that nonparametric logistic regression can be a useful approach to displaying and interpreting results of case-control studies.

Adult

Kinesin-related proteins in the mammalian testes: candidate motors for meiosis and morphogenesis.

The kinesin superfamily of molecular motors comprises proteins that participate in a wide variety of motile events within the cell. Members of this family share a highly homologous head domain responsible for force generation attached to a divergent tail domain thought to couple the motor domain to its target cargo. Many kinesin-related proteins (KRPs) participate in spindle morphogenesis and chromosome movement in cell division. Genetic analysis of mitotic KRPs in yeast and Drosophila, as well as biochemical experiments in other species, have suggested models for the function of KRPs in cell division, including both mitosis and meiosis. Although many mitotic KRPs have been identified, the relationship between mitotic motors and meiotic function is not clearly understood. We have used sequence similarity between mitotic KRPs to identify candidates for meiotic and/or mitotic motors in a vertebrate. We have identified a group of kinesin-related proteins from rat testes (termed here testes KRP1 through KRP6) that includes new members of the bimC and KIF2 subfamilies as well as proteins that may define new kinesin subfamilies. Five of the six testes KRPs identified are expressed primarily in testes. Three of these are expressed in a region of the seminiferous epithelia (SE) rich in meiotically active cells. Further characterization of one of these KRPs, KRP2, showed it to be a promising candidate for a motor in meiosis: it is localized to a meiotically active region of the SE and is homologous to motor proteins associated with the mitotic apparatus. Testes-specific genes provide the necessary probes to investigate whether the motor proteins that function in mammalian meiosis overlap with those of mitosis and whether motor proteins exist with functions unique to meiosis. Our search for meiotic motors in a vertebrate testes has successfully identified proteins with properties consistent with those of meiotic motors in addition to uncovering proteins that may function in other unique motile events of the SE.

Amino Acid Sequence

Assessing familial aggregation of age at onset, by using estimating equations, with application to breast cancer.

In genetic research of chronic diseases, age-at-onset outcomes within families are often correlated. The nature of correlation of age-at-onset outcomes is indicative of common genetic and/or shared environmental risk factors among family members. Understanding patterns of such correlation may shed light on the disease etiology and, hence, is an important step to take prior to further searching for the responsible genes via segregation and linkage studies. Age-at-onset outcomes are different from those familiar quantitative or qualitative traits for which many statistical methods have been developed. In comparison with the quantitative traits, age-at-onset outcomes are often censored, i.e., instead of actual age-at-onset outcomes, only the current ages or ages at death are observed. They are also different from qualitative traits because of their continuity. Because of the complexity of correlated censored outcomes, few methods have yet been developed. A traditional approach is to impose a parametric joint distribution for the correlated age-at-onset outcomes, which has been criticized for requiring a stringent assumption about the entire distribution of age at onset. The purpose of this paper is to describe a method for assessing familial aggregation of correlated age-at-onset outcomes semiparametrically, by use of estimating equations. This method does not require any parametric assumption for modeling the age at onset. The estimates of parameters, including those quantifying the correlation within families, are consistent and have an asymptotic normal distribution that can be used to make inferences. To illustrate this new method, we analyzed two age-at-onset data sets that were obtained from studies conducted in the States of Washington and Hawaii, with the objective of quantifying the familial aggregations of age at onset of breast cancer.

Age of Onset

Regression analysis with missing covariate data using estimating equations.

In regression analysis, missing covariate data has been among the most common problems. Frequently, practitioners adopt the so-called complete-case analysis, i.e., performing the analysis on only a complete dataset after excluding records with missing covariates. Performing a complete-case analysis is convenient with existing statistical packages, but it may be inefficient since the observed outcomes and covariates on those records with missing covariates are not used. It can even give misleading statistical inference if missing is not completely at random. This paper introduces a joint estimating equation (JEE) for regression analysis in the presence of missing observations on one covariate, which may be thought of as a method in a general framework for the missing covariate data problem proposed by Robins, Rotnitzky, and Zhao (1994, Journal of the American Statistical Association 89, 846-866). A generalization of JEE to more than one such covariate is discussed. The JEE is generally applicable to estimating regression coefficients from a regression model, including linear and logistic regression. Provided that the missing covariate data is either missing completely at random or missing at random (in addition to mild regularity conditions), estimates of regression coefficients from the JEE are consistent and have an asymptotic normal distribution. Simulation results show that the asymptotic distribution of estimated coefficients performs well in finite samples. Also shown through the simulation study is that the validity of JEE estimates depends on the correct specification of the probability function that characterizes the missing mechanism, suggesting a need for further research on how to robustify the estimation from making this nuisance assumption. Finally, the JEE is illustrated with an application from a case-control study of diet and thyroid cancer.

Biometry

Estimation methods for the join distribution of repeated binary observations.

The joint distribution of repeated binary observations is multinomial, and can be specified using a representation first suggested by Bahadur (1961, in Studies in Item Analysis and Prediction,158-168. Stanford, California: Stanford University Press), and later by Cox (1972, Applied Statistics 21, 13-120). Using the Bahadur representation, the marginal probabilities of success can be related to a set of covariates using the logistic link function, or any other suitable link function. Besides the parameters of the marginal regression model, we may also have interest in the probability of success on any of the repeated measures. For example, in the Six Cities study, a longitudinal study of the health effects of air pollution, we have interest in both the marginal probability of a child wheezing at age t (t = 10, 11, 12), and the union probability of wheezing at any of the three ages. This "union" probability can be specified in terms of the joint probabilities and the second higher-order correlations. We discuss several methods of estimating the parameters of the Bahadur model.

Adult