Search PubMed⌕ Search

Biomedical subjects

Li Hsu

Publications and source records attributed to Li Hsu.

At least 19 recordsLinked to original sources

MyGeneRisk Colon: A Web-Based Tool for Personalized Colorectal Cancer Risk Prediction Based on Genetics and Lifestyle.

Colorectal cancer (CRC) is a leading cause of cancer-related death, with incidence rising substantially among individuals under 50 years of age. Polygenic risk scores (PRS) hold promise for identifying high-risk individuals; when combined with lifestyle factors, they substantially improve prediction accuracy compared with models based on lifestyle factors alone. However, few clinical tools currently exist that facilitate this integrated, PRS-enhanced risk assessment. To bridge this gap, we developed MyGeneRisk Colo n, a publicly accessible web portal that delivers individualized CRC risk prediction by incorporating genetic, demographic, family history, and lifestyle factors. This paper details the development of the underlying risk prediction model, the portal's architecture and data security, our reporting framework, and engagement with a community advisory panel. Designed as a user-friendly platform, MyGeneRisk Colon aims to effectively communicate personalized CRC risk profiles and educate users and healthcare providers about prevention strategies.

Journal Article↗

Genetic risk factors modulate the association between physical activity and colorectal cancer.

BACKGROUND: Physical activity (PA) is an established protective factor for colorectal cancer (CRC), but it is unclear if genetic variants modify this effect. To investigate this possibility, we conducted a genome-wide gene-PA interaction analysis. METHODS: Using logistic regression and two-step and joint tests, we analyzed interactions between common genetic variants across the genome and PA in relation to CRC risk. Self-reported PA levels were categorized as active (&#x2265; 8.75 MET-h/wk) vs. inactive (< 8.75 MET-h/wk) and as study- and sex-specific quartiles of activity. RESULTS: PA had an overall protective effect on CRC (OR [active vs. inactive] = 0.85; 95%CI = 0.81-0.90). The two-step GxE method identified an interaction between rs4779584, an intergenic variant near the GREM1 and SCG5 genes, and PA for CRC risk (p-interaction = 2.6&#xd7;10- 8). Stratification by genotype at this locus showed a significant reduction in CRC risk by 20% in active vs. inactive participants with the CC genotype (OR = 0.80; 95%CI = 0.75-0.85), but no significant PA-CRC association among CT or TT carriers. When PA was modeled as quartiles, the 1-d.f. GxE test identified that rs56906466, an intergenic variant near the KCNG1 gene, modified the association between PA and CRC (p-interaction = 3.5&#xd7;10- 8). Stratification at this locus showed that increase in PA (highest vs. lowest quartile) was associated with a lower CRC risk solely among TT carriers (OR = 0.77; 95%CI = 0.72-0.82). CONCLUSIONS: In summary, we identified two genetic variants that modified the association between PA and CRC risk. One of them, related to GREM1 and SCG5, suggests that the bone morphogenetic protein (BMP)-related, inflammatory, and/or insulin signaling pathways may be associated with the protective influence of PA on colorectal carcinogenesis.

GWAS↗

The contributions of normal variation and genetic background to mammalian gene expression.

BACKGROUND: Qualitative and quantitative variability in gene expression represents the substrate for external conditions to exert selective pressures for natural selection. Current technologies allow for some forms of genetic variation, such as DNA mutations and polymorphisms, to be determined accurately on a comprehensive scale. Other components of variability, such as stochastic events in cellular transcriptional and translational processes, are less well characterized. Although potentially important, the relative contributions of genomic versus epigenetic and stochastic factors to variation in gene expression have not been quantified in mammalian species. RESULTS: In this study we compared microarray-based measures of hepatic transcript abundance levels within and between five different strains of Mus musculus. Within each strain 23% to 44% of all genes exhibited statistically significant differences in expression between genetically identical individuals (positive false discovery rate of 10%). Genes functionally associated with cell growth, cytokine activity, amine metabolism, and ubiquitination were enriched in this group. Genetic divergence between individuals of different strains also contributed to transcript abundance level differences, but to a lesser extent than intra-strain variation, with approximately 3% of all genes exhibiting inter-strain expression differences. CONCLUSION: These results indicate that although DNA sequence fixes boundaries for gene expression variability, there remain considerable latitudes of expression within these genome-defined limits that have the potential to influence phenotypes. The extent of normal or expected natural variability in gene expression may provide an additional level of phenotypic opportunity for natural selection.

Animals↗

Methods to test for association between a disease and a multi-allelic marker applied to a candidate region.

We report the analysis results of the Genetic Analysis Workshop 14 simulated microsatellite marker dataset, using replicate 50 from the Danacaa population. We applied several methods for association analysis of multi-allelic markers to case-control data to study the association between Kofendrerd Personality Disorder and multi-allelic markers in a candidate region previously identified by the linkage analysis. Evidence for association was found for marker D03S0127 (p < 0.01). The analyses were done without any prior knowledge of the answers.

Alleles↗

Modeling the effect of an associated single-nucleotide polymorphism in linkage studies.

For linkage analysis in affected sibling pairs, we propose a regression model to incorporate information from a disease-associated single-nucleotide polymorphism located under the linkage peak. This model can be used to study if the associated single-nucleotide polymorphism marker partly explains the original linkage peak. Two sources of information are used for performing this task, namely the genotypes of the parents and the genotypes of the siblings. We applied the methods to three significantly disease-associated single-nucleotide polymorphisms and five microsatellite markers at the end of chromosome 3 of replicate 1 of Aipotu population. Two out of five of the microsatellite markers showed a LOD score higher than 3. The question to be answered was whether one of the single-nucleotide polymorphisms partly explains these high LOD scores. We did not have the answers when we analyzed the data.

Chromosomes, Human, Pair 3↗

Locally weighted transmission/disequilibrium test for genetic association analysis.

The transmission/disequilibrium test statistic has been used for assessing genetic association in affected-parent trios. In the presence of multiple tightly linked marker loci where local dependency may exist, haplotypes are reconstructed statistically to estimate the joint effects of these markers. In this manuscript, we propose an alternative to the haplotype approach by taking a weighted average of multiple loci, where the weight is proportional to the product of (1-2X recombination fraction) and the linkage disequilibrium between markers. As an illustration, we applied the method to the simulated Aipotu data.

Genome-Wide Association Study↗

Multivariate survival analysis for case-control family data.

Multivariate survival data arise from case-control family studies in which the ages at disease onset for family members may be correlated. In this paper, we consider a multivariate survival model with the marginal hazard function following the proportional hazards model. We use a frailty-based approach in the spirit of Glidden and Self (1999) to account for the correlation of ages at onset among family members. Specifically, we first estimate the baseline hazard function nonparametrically by the innovation theorem, and then obtain maximum pseudolikelihood estimators for the regression and correlation parameters plugging in the baseline hazard function estimator. We establish a connection with a previously proposed generalized estimating equation-based approach. Simulation studies and an analysis of case-control family data of breast cancer illustrate the methodology's practical utility.

Age of Onset↗

Performance of the log-linear approach to case-parent triad data for assessing maternal genetic associations with offspring disease: type I error, power, and bias.

Maternal genetic variation may serve as a biomarker in studies aimed at clarifying fetal determinants of infant or adult disease. The log-linear approach to case-parent triad data (LCPT) can be used to investigate maternal genetic polymorphisms in relation to offspring disease risk, but LCPT operating characteristics have been reported for only a limited range of situations. The authors performed a simulation study to investigate the performance of the LCPT for assessing maternal associations with offspring disease risk over a wide range of scenarios with varying sample sizes (n), high-risk allele frequencies (f ), and modes of inheritance, all of which greatly affect the expected number of triads in informative categories. For most f values less than 0.5, the LCPT approach with 200 triads allowed for approximately 80% power to detect valid, unbiased maternal relative risks of 2 when inheritance was log-additive or dominant. When inheritance was recessive, this was true for most f 's greater than 0.35. Outside of this range, however, power and bias depended greatly on the mode of inheritance, f, and n. On the basis of these findings, epidemiologists may consider the LCPT a useful approach for assessing maternal relative risks unless one expects a very rare or fairly common maternal allele to increase offspring disease risk.

Bias↗

Denoising array-based comparative genomic hybridization data using wavelets.

Array-based comparative genomic hybridization (array-CGH) provides a high-throughput, high-resolution method to measure relative changes in DNA copy number simultaneously at thousands of genomic loci. Typically, these measurements are reported and displayed linearly on chromosome maps, and gains and losses are detected as deviations from normal diploid cells. We propose that one may consider denoising the data to uncover the true copy number changes before drawing inferences on the patterns of aberrations in the samples. Nonparametric techniques are particularly suitable for data denoising as they do not impose a parametric model in finding structures in the data. In this paper, we employ wavelets to denoise the data as wavelets have sound theoretical properties and a fast computational algorithm, and are particularly well suited for handling the abrupt changes seen in array-CGH data. A simulation study shows that denoising data prior to testing can achieve greater power in detecting the aberrant spot than using the raw data without denoising. Finally, we illustrate the method on two array-CGH data sets.

Breast Neoplasms↗

Assessing maternal genetic associations: a comparison of the log-linear approach to case-parent triad data and a case-control approach.

BACKGROUND: In utero exposures, including maternal phenotypes, are potential risk factors for both early-onset and adult-onset diseases. Two alternative study designs use maternal genotypes at polymorphic loci as biomarkers of an offspring's in utero exposure: (1) a traditional case-control study with logistic regression analysis, in which cases, controls, and mothers of both types of subjects are genotyped; and (2) a case-parent triad study with log-linear analysis, in which cases and both parents are genotyped. METHODS: We used computer simulations to compare the operating characteristics of the log-linear approach to case-parent triad data and the case-control approach for assessing relative risks (RRs) associated with maternal genotypes. RESULTS: For high-risk allele frequencies (chromosomal prevalence; f) between 0.20 and 0.75, both methods allowed for valid, unbiased estimates of maternal RRs. The case-parent triad approach, however, had 43% greater power, on average, than the case-control approach with an equal number of genotypes, and 13% greater power with an equal number of cases. For example, under dominant inheritance, to detect 2-fold maternal RRs with 200 (or 150) cases when allele prevalence is between 0.15 and 0.40, the case-parent triad and equal-genotype case-control designs had, on average, 87% and 62% power, respectively. As f approached 0 or 1, the power of both methods decreased sharply. DISCUSSION: The greater efficiency of case-parent triads may be due to the inclusion of paternal genotype information, which allows for independent tests of disease association with maternal or offspring genotypes. These results highlight one potential advantage of case-parent triad data in assessing maternal genetics as risk factors for offspring disease. We discuss these findings and other considerations between the 2 methodological approaches.

Case-Control Studies↗

Risk of testicular germ cell cancer in relation to variation in maternal and offspring cytochrome p450 genes involved in catechol estrogen metabolism.

The incidence of testicular germ cell carcinoma (TGCC) is highest among men ages 20 to 44 years. Exposure to relatively high circulating maternal estrogen levels during pregnancy has long been suspected as being a risk factor for TGCC. Catechol (hydroxylated) estrogens have carcinogenic potential, thought to arise from reactive catechol intermediates with enhanced capability of forming mutation-inducing DNA adducts. Polymorphisms in maternal or offspring genes encoding estrogen-metabolizing enzymes may influence prenatal catechol estrogen levels and could therefore be biomarkers of TGCC risk. We conducted a population-based, case-parent triad study to evaluate TGCC risk in relation to maternal and/or offspring polymorphisms in CYP1A2, CYP1B1, CYP3A4, and CYP3A5. We identified 18- to 44-year-old men diagnosed with invasive TGCC from 1999 to 2004 through a population-based cancer registry in Washington State and recruited cases and their parents (110 case-parent triads, 50 case-parent dyads). Maternal or offspring carriage of CYP1A2 -163A was associated with reduced risk of TGCC [maternal heterozygote relative risk (RR), 0.6; 95% confidence interval (95% CI), 0.2-1.7; offspring heterozygote RR, 0.7; 95% CI, 0.3-1.5)]. Maternal CYP1B1 (48)Gly homozygosity was associated with a 2.7-fold increased risk of TGCC (95% CI, 0.9-7.9), with little evidence that Leu(432)Val or Asn(453)Ser genotypes were related to risk. Men were also at increased risk of TGCC if they carried the CYP3A4 -392G (RR, 7.0; 95% CI, 1.6-31) or CYP3A5 6986G (RR, 2.4; 95% CI, 1.1-5.6) alleles. These results support the hypothesis that maternal and/or offspring catechol estrogen activity may influence sons' risk of TGCC.

Adolescent↗

Array comparative genomic hybridization analysis of genomic alterations in breast cancer subtypes.

In this study, we performed high-resolution array comparative genomic hybridization with an array of 4153 bacterial artificial chromosome clones to assess copy number changes in 44 archival breast cancers. The tumors were flow sorted to exclude non-tumor DNA and increase our ability to detect gene copy number changes. In these tumors, losses were more frequent than gains, and gains in 1q and loss in 16q were the most frequent alterations. We compared gene copy number changes in the tumors based on histologic subtype and estrogen receptor (ER) status, i.e., ER-negative infiltrating ductal carcinoma, ER-positive infiltrating ductal carcinoma, and ER-positive infiltrating lobular carcinoma. We observed a consistent association between loss in regions of 5q and ER-negative infiltrating ductal carcinoma, as well as more frequent loss in 4p16, 8p23, 8p21, 10q25, and 17p11.2 in ER-negative infiltrating ductal carcinoma compared with ER-positive infiltrating ductal carcinoma (adjusted P values < or = 0.05). We also observed high-level amplifications in ER-negative infiltrating ductal carcinoma in regions of 8q24 and 17q12 encompassing the c-myc and c-erbB-2 genes and apparent homozygous deletions in 3p21, 5q33, 8p23, 8p21, 9q34, 16q24, and 19q13. ER-positive infiltrating ductal carcinoma showed a higher frequency of gain in 16p13 and loss in 16q21 than ER-negative infiltrating ductal carcinoma. Correlation analysis highlighted regions of change commonly seen together in ER-negative infiltrating ductal carcinoma. ER-positive infiltrating lobular carcinoma differed from ER-positive infiltrating ductal carcinoma in the frequency of gain in 1q and loss in 11q and showed high-level amplifications in 1q32, 8p23, 11q13, and 11q14. These results indicate that array comparative genomic hybridization can identify significant differences in the genomic alterations between subtypes of breast cancer.

Adult↗

Nonparametric correction for covariate measurement error in a stratified Cox model.

Stratified Cox regression models with large number of strata and small stratum size are useful in many settings, including matched case-control family studies. In the presence of measurement error in covariates and a large number of strata, we show that extensions of existing methods fail either to reduce the bias or to correct the bias under nonsymmetric distributions of the true covariate or the error term. We propose a nonparametric correction method for the estimation of regression coefficients, and show that the estimators are asymptotically consistent for the true parameters. Small sample properties are evaluated in a simulation study. The method is illustrated with an analysis of Framingham data.

Adult↗

Partially supervised learning using an EM-boosting algorithm.

Training data in a supervised learning problem consist of the class label and its potential predictors for a set of observations. Constructing effective classifiers from training data is the goal of supervised learning. In biomedical sciences and other scientific applications, class labels may be subject to errors. We consider a setting where there are two classes but observations with labels corresponding to one of the classes may in fact be mislabeled. The application concerns the use of protein mass-spectrometry data to discriminate between serum samples from cancer and noncancer patients. The patients in the training set are classified on the basis of tissue biopsy. Although biopsy is 100% specific in the sense that a tissue that shows itself to have malignant cells is certainly cancer, it is less than 100% sensitive. Reference gold standards that are subject to this special type of misclassification due to imperfect diagnosis certainty arise in many fields. We consider the development of a supervised learning algorithm under these conditions and refer to it as partially supervised learning. Boosting is a supervised learning algorithm geared toward high-dimensional predictor data, such as those generated in protein mass-spectrometry. We propose a modification of the boosting algorithm for partially supervised learning. The proposal is to view the true class membership of the samples that are labeled with the error-prone class label as missing data, and apply an algorithm related to the EM algorithm for minimization of a loss function. To assess the usefulness of the proposed method, we artificially mislabeled a subset of samples and applied the original and EM-modified boosting (EM-Boost) algorithms for comparison. Notable improvements in misclassification rates are observed with EM-Boost.

Algorithms↗

Semiparametric estimation of marginal hazard function from case-control family studies.

Estimating marginal hazard function from the correlated failure time data arising from case-control family studies is complicated by noncohort study design and risk heterogeneity due to unmeasured, shared risk factors among the family members. Accounting for both factors in this article, we propose a two-stage estimation procedure. At the first stage, we estimate the dependence parameter in the distribution for the risk heterogeneity without obtaining the marginal distribution first or simultaneously. Assuming that the dependence parameter is known, at the second stage we estimate the marginal hazard function by iterating between estimation of the risk heterogeneity (frailty) for each family and maximization of the partial likelihood function with an offset to account for the risk heterogeneity. We also propose an iterative procedure to improve the efficiency of the dependence parameter estimate. The simulation study shows that both methods perform well under finite sample sizes. We illustrate the method with a case-control family study of early onset breast cancer.

Adult↗

Familial aggregation of dyslexia phenotypes. II: paired correlated measures.

Dyslexia is a common and complex behavioral disorder characterized by unexpected difficulty in learning to read. Psychometric measures used to assess dyslexia often evaluate overlapping processes or abilities. To identify subphenotypes amenable to model-based linkage analyses, we have used careful language phenotyping, familial aggregation analyses of single phenotype measures, and segregation analyses. In the current study, to identify covariates to use in future segregation analyses we examined six pairs of related measures selected from among the most promising candidates in the initial aggregation analyses whose aggregation patterns were most consistent with a genetic basis. For these reciprocal aggregation analyses each measure is evaluated with the paired measure as the covariate to obtain information about the interdependence of the paired measures on shared genetic factors. Six pairs of measures were evaluated: 1) accuracy and efficiency of phonological decoding; 2) phonological nonword memory and written spelling; 3) phonological decoding accuracy and written spelling; 4) inattention ratings and rapid automatized naming for switching letters and numerals (RAS); 5) inattention ratings and oral reading rate; and 6) RAS and oral reading rate. Results of these analyses provide evidence that there may be a genetic contribution to efficiency of phonological decoding in addition to the genetic contribution it shares with accuracy of phonological decoding, a genetic contribution to phonological nonword memory in addition to the genetic contribution it shares with written spelling, a genetic contribution to written spelling in addition to the genetic contribution it shares with accuracy of phonological decoding, and a genetic contribution to inattention ratings in addition to the genetic contribution it shares with either RAS or oral reading rate.

Child↗

Some further results on incorporating risk factor information in assessing the dependence between paired failure times arising from case-control family studies: an application to prostate cancer.

In a typical case-control family study, detailed risk factor information is often collected on cases and controls, but not on their relatives for reasons of cost and logistical difficulty in locating the relatives. The impact of missing risk factor information for relatives on estimation of the strength of dependence between the disease risk of pairs of relatives is largely unknown. In this paper, we extend our earlier work on estimating the dependence of ages at onset between paired relatives from case-control family data to include covariates on cases and controls, and possibly relatives. Using population-based case-control families as our basic data structure, we study the effect of missing covariates for relatives and/or cases and controls on the bias of certain dependence parameter estimators via a simulation study. Finally we illustrate various analyses using a case-control family study of early onset prostate cancer.

Age of Onset↗

Single nucleotide polymorphism array analysis of flow-sorted epithelial cells from frozen versus fixed tissues for whole genome analysis of allelic loss in breast cancer.

Analysis of allelic loss in archival tumor specimens is constrained by quality and quantity of tissue and by technical limitations on the number of chromosomal sites that can be efficiently evaluated in conventional analyses using polymorphic microsatellite markers. Newly developed array-based assays have the potential to yield genome-wide data from small amounts of tissue but have not been validated for use with routinely processed specimens. We used the Affymetrix HuSNP assay, composed of 1494 single nucleotide polymorphism sites, to compare allelic loss results obtained from both formalin-fixed and frozen breast tissue samples. Tumor cells were separated from normal epithelia and nonepithelial cells by dissection and bivariate cytokeratin/DNA flow sorting; normal breast cells from the same patient served as constitutive normal. Allele results from the HuSNP array averaged 96% reproducibility between duplicates and were concordant between the fixed and frozen normal samples. We also analyzed DNA from the same samples after whole-genome amplification (primer extension preamplification). Although overall signal intensities were lower, the genotype data from the primer extension preamplification material was concordant with genomic DNA data from the same samples. Results from genomic normal tissue DNA averaged informative single nucleotide polymorphism at 379 (25%) loci genome-wide. Although data points were clustered and some segments of chromosomes were not informative, our data indicated that the Affymetrix HuSNP assay could provide an efficient and valid genome-wide analysis of allelic imbalance in routinely processed and whole genome-amplified pathology specimens.

Breast Neoplasms↗