Search PubMed⌕ Search

Biomedical subjects

C Kooperberg

Publications and source records attributed to C Kooperberg.

10 recordsLinked to original sources

Sequence analysis using logic regression.

Logic Regression is a new adaptive regression methodology that attempts to construct predictors as Boolean combinations of (binary) covariates. In this paper we use this algorithm to deal with single-nucleotide polymorphism (SNP) sequence data. The predictors that are found are interpretable as risk factors of the disease. Significance of these risk factors is assessed using techniques like cross-validation, permutation tests, and independent test sets. These model selection techniques remain valid when data is dependent, as is the case for the family data used here. In our analysis of the Genetic Analysis Workshop 12 data we identify the exact locations of mutations on gene 1 and gene 6 and a number of mutations on gene 2 that are associated with the affected status, without selecting any false positives.

Algorithms↗

Widespread collaboration of Isw2 and Sin3-Rpd3 chromatin remodeling complexes in transcriptional repression.

The yeast Isw2 chromatin remodeling complex functions in parallel with the Sin3-Rpd3 histone deacetylase complex to repress early meiotic genes upon recruitment by Ume6p. For many of these genes, the effect of an isw2 mutation is partially masked by a functional Sin3-Rpd3 complex. To identify the full range of genes repressed or activated by these factors and uncover hidden targets of Isw2-dependent regulation, we performed full genome expression analyses using cDNA microarrays. We find that the Isw2 complex functions mainly in repression of transcription in a parallel pathway with the Sin3-Rpd3 complex. In addition to Ume6 target genes, we find that many Ume6-independent genes are derepressed in mutants lacking functional Isw2 and Sin3-Rpd3 complexes. Conversely, we find that ume6 mutants, but not isw2 sin3 or isw2 rpd3 double mutants, have reduced fidelity of mitotic chromosome segregation, suggesting that one or more functions of Ume6p are independent of Sin3-Rpd3 and Isw2 complexes. Chromatin structure analyses of two nonmeiotic genes reveals increased DNase I sensitivity within their regulatory regions in an isw2 mutant, as seen previously for one meiotic locus. These data suggest that the Isw2 complex functions at Ume6-dependent and -independent loci to create DNase I-inaccessible chromatin structure by regulating the positioning or placement of nucleosomes.

Adenosine Triphosphatases↗

Correlates of serum lycopene in older women.

Experimental and epidemiological evidence suggests that lycopene, a predominant carotenoid found in human serum, may reduce the risk of certain cancers. We examined the association of dietary, physiological, and other factors with serum lycopene concentrations in a subsample of 946 postmenopausal women participating in the Women's Health Initiative. Pearson partial correlation coefficients and linear regression coefficients were calculated after adjustment for age, ethnicity, and serum low-density-lipoprotein (LDL) cholesterol. Serum lycopene was correlated with serum LDL cholesterol (r = 0.23) and dietary lycopene (r = 0.17, both p < 0.001). Individual food items found to be correlated with serum lycopene after adjustment included fresh tomatoes or tomato juice (r = 0.11), cooked tomatoes, tomato sauce, or salsa (r = 0.17), and spaghetti with meat sauce (r = 0.19, all p < 0.01). Age and body mass index were negatively associated with serum lycopene levels (both p < 0.001). Serum lycopene levels were highest in the summer and highest for those living in the northeastern United States. If we postulate that high serum lycopene levels reduce cancer risk, it becomes apparent that we have limited ability to detect this association from studies of lycopene intake. An understanding of factors associated with serum lycopene levels can be useful for the interpretation of studies of dietary lycopene and disease risk.

Aged↗

Linear regression for bivariate censored data via multiple imputation.

Bivariate survival data arise, for example, in twin studies and studies of both eyes or ears of the same individual. Often it is of interest to regress the survival times on a set of predictors. In this paper we extend Wei and Tanner's multiple imputation approach for linear regression with univariate censored data to bivariate censored data. We formulate a class of censored bivariate linear regression methods by iterating between the following two steps: 1. the data is augmented by imputing survival times for censored observations; 2. a linear model is fit to the imputed complete data. We consider three different methods to implement these two steps. In particular, the marginal (independence) approach ignores the possible correlation between two survival times when estimating the regression coefficient. To improve the efficiency, we propose two methods that account for the correlation between the survival times. First, we improve the efficiency by using generalized least squares regression in step 2. Second, instead of generating data from an estimate of the marginal distribution we generate data from a bivariate log-spline density estimate in step 1. Through simulation studies we find that the performance of the two methods that take the dependence into account is close and that they are both more efficient than the marginal approach. The methods are applied to a data set from an otitis media clinical trial.

Anti-Bacterial Agents↗

Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins.

We describe the development of a scoring function based on the decomposition P(structure/sequence) proportional to P(sequence/structure) *P(structure), which outperforms previous scoring functions in correctly identifying native-like protein structures in large ensembles of compact decoys. The first term captures sequence-dependent features of protein structures, such as the burial of hydrophobic residues in the core, the second term, universal sequence-independent features, such as the assembly of beta-strands into beta-sheets. The efficacies of a wide variety of sequence-dependent and sequence-independent features of protein structures for recognizing native-like structures were systematically evaluated using ensembles of approximately 30,000 compact conformations with fixed secondary structure for each of 17 small protein domains. The best results were obtained using a core scoring function with P(sequence/structure) parameterized similarly to our previous work (Simons et al., J Mol Biol 1997;268:209-225] and P(structure) focused on secondary structure packing preferences; while several additional features had some discriminatory power on their own, they did not provide any additional discriminatory power when combined with the core scoring function. Our results, on both the training set and the independent decoy set of Park and Levitt (J Mol Biol 1996;258:367-392), suggest that this scoring function should contribute to the prediction of tertiary structure from knowledge of sequence and secondary structure.

Models, Statistical↗

Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.

We explore the ability of a simple simulated annealing procedure to assemble native-like structures from fragments of unrelated protein structures with similar local sequences using Bayesian scoring functions. Environment and residue pair specific contributions to the scoring functions appear as the first two terms in a series expansion for the residue probability distributions in the protein database; the decoupling of the distance and environment dependencies of the distributions resolves the major problems with current database-derived scoring functions noted by Thomas and Dill. The simulated annealing procedure rapidly and frequently generates native-like structures for small helical proteins and better than random structures for small beta sheet containing proteins. Most of the simulated structures have native-like solvent accessibility and secondary structure patterns, and thus ensembles of these structures provide a particularly challenging set of decoys for evaluating scoring functions. We investigate the effects of multiple sequence information and different types of conformational constraints on the overall performance of the method, and the ability of a variety of recently developed scoring functions to recognize the native-like conformations in the ensembles of simulated structures.

Bayes Theorem↗

Hazard regression with interval-censored data.

In a recent paper, Kooperberg, Stone, and Truong (1995a) introduced hazard regression (HARE), in which linear splines and their tensor products are used to estimate the conditional log-hazard function based on possibly censored, positive response data and one or more covariates. Model selection is carried out in an adaptive fashion using maximum likelihood estimation of the unknown coefficients, Rao and Wald statistics to carry out stepwise addition and deletion of basis functions, and the Bayesian Information Criterion (BIC) to select the final model. In the present paper, the HARE methodology is extended to accommodate interval-censored data, time-dependent covariates, and cubic splines. The presence of interval-censored data means that the log-likelihood function may no longer be concave, presenting additional numerical challenges. The extended methodology is applied to a data set containing both interval-censoring and time-dependent covariates. The new software will be available in a future release of S-Plus.

Acquired Immunodeficiency Syndrome↗

Statistical modeling to predict elective surgery time. Comparison with a computer scheduling system and surgeon-provided estimates.

BACKGROUND: Accurate estimation of operating times is a prerequisite for the efficient scheduling of the operating suite. The authors, in this study, sought to compare surgeons' time estimates for elective cases with those of commercial scheduling software, and to ascertain whether improvements could be made by regression modeling. METHODS: The study was conducted at the University of Washington Medical Center in three phases. Phase 1 retrospectively reviewed surgeons' time estimates and the scheduling system's estimates throughout 1 yr. In phase 2, data were collected prospectively from participating surgeons by means of a data entry form completed at the time of scheduling elective cases. Data included the procedure code, estimated operating time, estimated case difficulty, and potential factors that might affect the duration. In phase 3, identical data were collected from five selected surgeons by personal interview. RESULTS: In phase 1, 26 of 43 surgeons provided significantly better estimates than did the scheduling system (P < 0.01), and no surgeon was significantly worse, although the absolute errors were large (34% of 157 min average case length). In phase 2, modeling improved the accuracy of the surgeons' estimates by 11.5%, compared with the scheduling system. In phase 3, applying the model from phase 2 improved the accuracy of the surgeons' estimates by 18.2%. CONCLUSIONS: Surgeons provide more accurate time estimates than does the scheduling software as it is used in our institution. Regression modeling effects modest improvements in accuracy. Further improvements would be likely if the hospital information system could provide timely historical data and feedback to the surgeons.

Appointments and Schedules↗

Trees and splines in survival analysis.

During the past few years several nonparametric alternatives to the Cox proportional hazards model have appeared in the literature. These methods extend techniques that are well known from regression analysis to the analysis of censored survival data. In this paper we discuss methods based on (partition) trees and (polynomial) splines, analyse two datasets using both Survival Trees and HARE, and compare the strengths and weaknesses of the two methods. One of the strengths of HARE is that its model fitting procedure has an implicit check for proportionality of the underlying hazards model. It also provides an explicit model for the conditional hazards function, which makes it very convenient to obtain graphical summaries. On the other hand, the tree-based methods automatically partition a dataset into groups of cases that are similar in survival history. Results obtained by survival trees and HARE are often complementary. Trees and splines in survival analysis should provide the data analyst with two useful tools when analysing survival data.

Algorithms↗

Using logistic regression to estimate the adjusted attributable risk of low birthweight in an unmatched case-control study.

Other authors have shown how to estimate attributable risk based on stratification. In this paper, we show how to estimate adjusted attributable risks, standard errors, and confidence intervals from an unmatched case-control study that has population-based controls and uses the logistic regression model to estimate relative risk. We apply the method to data from a case-control study of low birthweight. The method is conceptually simple, has no assumptions beyond those of the logistic model, makes use of computer-intensive statistical techniques (the bootstrap), and extends to interactions. A Fortran computer program to carry out the computations is available from the authors upon request.

Case-Control Studies↗