Search PubMed⌕ Search

Biomedical subjects

Joseph G Ibrahim

Publications and source records attributed to Joseph G Ibrahim.

At least 19 recordsLinked to original sources

A spatially resolved genomic-molecular atlas of human white-matter microstructure.

Human white matter has been linked to inherited variation, circulating molecular state and brain disease, but these layers have rarely been mapped onto the same tract anatomy. Here we measured genetic effects along 6,090 atlas-aligned fiber pathways sampled at 609,000 locations in 72,185 UK Biobank participants, and integrated proteomic and metabolomic profiles within the same anatomical frame. Genetic effects were not whole-tract properties: each locus formed a spatial footprint along fiber trajectories, ranging from single locations to broad multi-tract patterns and reflecting regional polygenicity rather than tract heritability. This map identified 258, 186 and 298 previously unreported loci for fractional anisotropy, mean diffusivity and axial diffusivity; spatial patterns replicated in adults and 157 of 315 FA loci replicated in adolescence in ABCD. Mendelian randomization linked localized genetic effects to neurodegenerative and psychiatric traits, with Alzheimer's disease showing directional effects across 12 of 17 tracts. Multi-omic analyses identified 97 proteomic and 161 metabolomic associations, with the broadest signals from lipid metabolites including linoleic acid and phosphatidylcholines. The strongest lipid-metabolite and genetic signals converged in the corpus callosum, placing inherited variation, disease risk and systemic lipid metabolism on the same localized tract segments.

Journal Article↗

Pseudo-likelihood methods for longitudinal binary data with non-ignorable missing responses and covariates.

In this paper we consider longitudinal studies in which the outcome to be measured over time is binary, and the covariates of interest are categorical. In longitudinal studies it is common for the outcomes and any time-varying covariates to be missing due to missed study visits, resulting in non-monotone patterns of missingness. Moreover, the reasons for missed visits may be related to the specific values of the response and/or covariates that should have been obtained, i.e. missingness is non-ignorable. With non-monotone non-ignorable missing response and covariate data, a full likelihood approach is quite complicated, and maximum likelihood estimation can be computationally prohibitive when there are many occasions of follow-up. Furthermore, the full likelihood must be correctly specified to obtain consistent parameter estimates. We propose a pseudo-likelihood method for jointly estimating the covariate effects on the marginal probabilities of the outcomes and the parameters of the missing data mechanism. The pseudo-likelihood requires specification of the marginal distributions of the missingness indicator, outcome, and possibly missing covariates at each occasions, but avoids making assumptions about the joint distribution of the data at two or more occasions. Thus, the proposed method can be considered semi-parametric. The proposed method is an extension of the pseudo-likelihood approach in Troxel et al. to handle binary responses and possibly missing time-varying covariates. The method is illustrated using data from the Six Cities study, a longitudinal study of the health effects of air pollution.

Air Pollutants↗

Estimation in regression models for longitudinal binary data with outcome-dependent follow-up.

In many observational studies, individuals are measured repeatedly over time, although not necessarily at a set of pre-specified occasions. Instead, individuals may be measured at irregular intervals, with those having a history of poorer health outcomes being measured with somewhat greater frequency and regularity. In this paper, we consider likelihood-based estimation of the regression parameters in marginal models for longitudinal binary data when the follow-up times are not fixed by design, but can depend on previous outcomes. In particular, we consider assumptions regarding the follow-up time process that result in the likelihood function separating into two components: one for the follow-up time process, the other for the outcome measurement process. The practical implication of this separation is that the follow-up time process can be ignored when making likelihood-based inferences about the marginal regression model parameters. That is, maximum likelihood (ML) estimation of the regression parameters relating the probability of success at a given time to covariates does not require that a model for the distribution of follow-up times be specified. However, to obtain consistent parameter estimates, the multinomial distribution for the vector of repeated binary outcomes must be correctly specified. In general, ML estimation requires specification of all higher-order moments and the likelihood for a marginal model can be intractable except in cases where the number of repeated measurements is relatively small. To circumvent these difficulties, we propose a pseudolikelihood for estimation of the marginal model parameters. The pseudolikelihood uses a linear approximation for the conditional distribution of the response at any occasion, given the history of previous responses. The appeal of this approximation is that the conditional distributions are functions of the first two moments of the binary responses only. When the follow-up times depend only on the previous outcome, the pseudolikelihood requires correct specification of the conditional distribution of the current outcome given the outcome at the previous occasion only. Results from a simulation study and a study of asymptotic bias are presented. Finally, we illustrate the main results using data from a longitudinal observational study that explored the cardiotoxic effects of doxorubicin chemotherapy for the treatment of acute lymphoblastic leukemia in children.

Adolescent↗

Semiparametric models for missing covariate and response data in regression models.

We consider a class of semiparametric models for the covariate distribution and missing data mechanism for missing covariate and/or response data for general classes of regression models including generalized linear models and generalized linear mixed models. Ignorable and nonignorable missing covariate and/or response data are considered. The proposed semiparametric model can be viewed as a sensitivity analysis for model misspecification of the missing covariate distribution and/or missing data mechanism. The semiparametric model consists of a generalized additive model (GAM) for the covariate distribution and/or missing data mechanism. Penalized regression splines are used to express the GAMs as a generalized linear mixed effects model, in which the variance of the corresponding random effects provides an intuitive index for choosing between the semiparametric and parametric model. Maximum likelihood estimates are then obtained via the EM algorithm. Simulations are given to demonstrate the methodology, and a real data set from a melanoma cancer clinical trial is analyzed using the proposed methods.

Algorithms↗

Joint models for multivariate longitudinal and multivariate survival data.

Joint modeling of longitudinal and survival data is becoming increasingly essential in most cancer and AIDS clinical trials. We propose a likelihood approach to extend both longitudinal and survival components to be multidimensional. A multivariate mixed effects model is presented to explicitly capture two different sources of dependence among longitudinal measures over time as well as dependence between different variables. For the survival component of the joint model, we introduce a shared frailty, which is assumed to have a positive stable distribution, to induce correlation between failure times. The proposed marginal univariate survival model, which accommodates both zero and nonzero cure fractions for the time to event, is then applied to each marginal survival function. The proposed multivariate survival model has a proportional hazards structure for the population hazard, conditionally as well as marginally, when the baseline covariates are specified through a specific mechanism. In addition, the model is capable of dealing with survival functions with different cure rate structures. The methodology is specifically applied to the International Breast Cancer Study Group (IBCSG) trial to investigate the relationship between quality of life, disease-free survival, and overall survival.

Acquired Immunodeficiency Syndrome↗

A class of Bayesian shared gamma frailty models with multivariate failure time data.

For multivariate failure time data, we propose a new class of shared gamma frailty models by imposing the Box-Cox transformation on the hazard function, and the product of the baseline hazard and the frailty. This novel class of models allows for a very broad range of shapes and relationships between the hazard and baseline hazard functions. It includes the well-known Cox gamma frailty model and a new additive gamma frailty model as two special cases. Due to the nonnegative hazard constraint, this shared gamma frailty model is computationally challenging in the Bayesian paradigm. The joint priors are constructed through a conditional-marginal specification, in which the conditional distribution is univariate, and it absorbs the nonlinear parameter constraints. The marginal part of the prior specification is free of constraints. The prior distributions allow us to easily compute the full conditionals needed for Gibbs sampling, while incorporating the constraints. This class of shared gamma frailty models is illustrated with a real dataset.

Adolescent↗

A flexible B-spline model for multiple longitudinal biomarkers and survival.

Often when jointly modeling longitudinal and survival data, we are interested in a multivariate longitudinal measure that may not fit well by linear models. To overcome this problem, we propose a joint longitudinal and survival model that has a nonparametric model for the longitudinal markers. We use cubic B-splines to specify the longitudinal model and a proportional hazards model to link the longitudinal measures to the hazard. To fit the model, we use a Markov chain Monte Carlo algorithm. We select the number of knots for the cubic B-spline model using the Conditional Predictive Ordinate (CPO) and the Deviance Information Criterion (DIC). The method and model selection approach are validated in a simulation. We apply this method to examine the link between viral load, CD4 count, and time to event in data from an AIDS clinical trial. The cubic B-spline model provides a good fit to the longitudinal data that could not be obtained with simple parametric models.

Acquired Immunodeficiency Syndrome↗

Wavelet thresholding with bayesian false discovery rate control.

The false discovery rate (FDR) procedure has become a popular method for handling multiplicity in high-dimensional data. The definition of FDR has a natural Bayesian interpretation; it is the expected proportion of null hypotheses mistakenly rejected given a measure of evidence for their truth. In this article, we propose controlling the positive FDR using a Bayesian approach where the rejection rule is based on the posterior probabilities of the null hypotheses. Correspondence between Bayesian and frequentist measures of evidence in hypothesis testing has been studied in several contexts. Here we extend the comparison to multiple testing with control of the FDR and illustrate the procedure with an application to wavelet thresholding. The problem consists of recovering signal from noisy measurements. This involves extracting wavelet coefficients that result from true signal and can be formulated as a multiple hypotheses-testing problem. We use simulated examples to compare the performance of our approach to the Benjamini and Hochberg (1995, Journal of the Royal Statistical Society, Series B57, 289-300) procedure. We also illustrate the method with nuclear magnetic resonance spectral data from human brain.

Bayes Theorem↗

Bayesian error-in-variable survival model for the analysis of GeneChip arrays.

DNA microarrays in conjunction with statistical models may help gain a deeper understanding of the molecular basis for specific diseases. An intense area of research is concerned with the identification of genes related to particular phenotypes. The technology, however, is subject to various sources of error that may lead to expression readings that are substantially different from the true transcript levels. Few methods for microarray data analysis have accounted for measurement error in a substantial way and that is the purpose of this investigation. We describe a Bayesian error-in-variable model for the analysis of microarray data from a clinical study of patients with acute lymphoblastic leukemia. We focus in particular on the problem of identifying genes whose expression patterns are associated with duration of remission. This is a question of great practical interest since relapse is a major concern in the treatment of this disease. We explore the effects of ignoring the uncertainty in the expression estimates on the selection and ranking of genes.

Bayes Theorem↗

A general class of Bayesian survival models with zero and nonzero cure fractions.

We propose a new class of survival models which naturally links a family of proper and improper population survival functions. The models resulting in improper survival functions are often referred to as cure rate models. This class of regression models is formulated through the Box-Cox transformation on the population hazard function and a proper density function. By adding an extra transformation parameter into the cure rate model, we are able to generate models with a zero cure rate, thus leading to a proper population survival function. A graphical illustration of the behavior and the influence of the transformation parameter on the regression model is provided. We consider a Bayesian approach which is motivated by the complexity of the model. Prior specification needs to accommodate parameter constraints due to the non-negativity of the survival function. Moreover, the likelihood function involves a complicated integral on the survival function, which may not have an analytical closed form, and thus makes the implementation of Gibbs sampling more difficult. We propose an efficient Markov chain Monte Carlo computational scheme based on Gaussian quadrature. The proposed method is illustrated with an example involving a melanoma clinical trial.

Adult↗

Bayesian analysis for generalized linear models with nonignorably missing covariates.

We propose Bayesian methods for estimating parameters in generalized linear models (GLMs) with nonignorably missing covariate data. We show that when improper uniform priors are used for the regression coefficients, phi, of the multinomial selection model for the missing data mechanism, the resulting joint posterior will always be improper if (i) all missing covariates are discrete and an intercept is included in the selection model for the missing data mechanism, or (ii) at least one of the covariates is continuous and unbounded. This impropriety will result regardless of whether proper or improper priors are specified for the regression parameters, beta, of the GLM or the parameters, alpha, of the covariate distribution. To overcome this problem, we propose a novel class of proper priors for the regression coefficients, phi, in the selection model for the missing data mechanism. These priors are robust and computationally attractive in the sense that inferences about beta are not sensitive to the choice of the hyperparameters of the prior for phi and they facilitate a Gibbs sampling scheme that leads to accelerated convergence. In addition, we extend the model assessment criterion of Chen, Dey, and Ibrahim (2004a, Biometrika 91, 45-63), called the weighted L measure, to GLMs and missing data problems as well as extend the deviance information criterion (DIC) of Spiegelhalter et al. (2002, Journal of the Royal Statistical Society B 64, 583-639) for assessing whether the missing data mechanism is ignorable or nonignorable. A novel Markov chain Monte Carlo sampling algorithm is also developed for carrying out posterior computation. Several simulations are given to investigate the performance of the proposed Bayesian criteria as well as the sensitivity of the prior specification. Real datasets from a melanoma cancer clinical trial and a liver cancer study are presented to further illustrate the proposed methods.

Bayes Theorem↗

A semiparametric mixture model for analyzing clustered competing risks data.

A very general class of multivariate life distributions is considered for analyzing failure time clustered data that are subject to censoring and multiple modes of failure. Conditional on cluster-specific quantities, the joint distribution of the failure time and event indicator can be expressed as a mixture of the distribution of time to failure due to a certain type (or specific cause), and the failure type distribution. We assume here the marginal probabilities of various failure types are logistic functions of some covariates. The cluster-specific quantities are subject to some unknown distribution that causes frailty. The unknown frailty distribution is modeled nonparametrically using a Dirichlet process. In such a semiparametric setup, a hybrid method of estimation is proposed based on the i.i.d. Weighted Chinese Restaurant algorithm that helps us generate observations from the predictive distribution of the frailty. The Monte Carlo ECM algorithm plays a vital role for obtaining the estimates of the parameters that assess the extent of the effects of the causal factors for failures of a certain type. A simulation study is conducted to study the consistency of our methodology. The proposed methodology is used to analyze a real data set on HIV infection of a cohort of female prostitutes in Senegal.

AIDS Vaccines↗

REC, Drosophila MCM8, drives formation of meiotic crossovers.

Crossovers ensure the accurate segregation of homologous chromosomes from one another during meiosis. Here, we describe the identity and function of the Drosophila melanogaster gene recombination defective (rec), which is required for most meiotic crossing over. We show that rec encodes a member of the mini-chromosome maintenance (MCM) protein family. Six MCM proteins (MCM2-7) are essential for DNA replication and are found in all eukaryotes. REC is the Drosophila ortholog of the recently identified seventh member of this family, MCM8. Our phylogenetic analysis reveals the existence of yet another family member, MCM9, and shows that MCM8 and MCM9 arose early in eukaryotic evolution, though one or both have been lost in multiple eukaryotic lineages. Drosophila has lost MCM9 but retained MCM8, represented by REC. We used genetic and molecular methods to study the function of REC in meiotic recombination. Epistasis experiments suggest that REC acts after the Rad51 ortholog SPN-A but before the endonuclease MEI-9. Although crossovers are reduced by 95% in rec mutants, the frequency of noncrossover gene conversion is significantly increased. Interestingly, gene conversion tracts in rec mutants are about half the length of tracts in wild-type flies. To account for these phenotypes, we propose that REC facilitates repair synthesis during meiotic recombination. In the absence of REC, synthesis does not proceed far enough to allow formation of an intermediate that can give rise to crossovers, and recombination proceeds via synthesis-dependent strand annealing to generate only noncrossover products.

Animals↗

A bayesian hierarchical model for the analysis of Affymetrix arrays.

An area of active research in DNA microarray analysis focuses on identifying differentially expressed genes between normal and malignant tissues. The analysis is complicated by the presence of several unreliable expression readings. Here, we illustrate a methodology where the expression estimates are modeled as censored data and discriminating genes are selected using ANOVA-based criteria.

Analysis of Variance↗

A Bayesian semiparametric joint hierarchical model for longitudinal and survival data.

This article proposes a new semiparametric Bayesian hierarchical model for the joint modeling of longitudinal and survival data. We relax the distributional assumptions for the longitudinal model using Dirichlet process priors on the parameters defining the longitudinal model. The resulting posterior distribution of the longitudinal parameters is free of parametric constraints, resulting in more robust estimates. This type of approach is becoming increasingly essential in many applications, such as HIV and cancer vaccine trials, where patients' responses are highly diverse and may not be easily modeled with known distributions. An example will be presented from a clinical trial of a cancer vaccine where the survival outcome is time to recurrence of a tumor. Immunologic measures believed to be predictive of tumor recurrence were taken repeatedly during follow-up. We will present an analysis of this data using our new semiparametric Bayesian hierarchical joint modeling methodology to determine the association of these longitudinal immunologic measures with time to tumor recurrence.

Bayes Theorem↗

Identification of differentially expressed genes in high-density oligonucleotide arrays accounting for the quantification limits of the technology.

In DNA microarray analysis, there is often interest in isolating a few genes that best discriminate between tissue types. This is especially important in cancer, where different clinicopathologic groups are known to vary in their outcomes and response to therapy. The identification of a small subset of gene expression patterns distinctive for tumor subtypes can help design treatment strategies and improve diagnosis. Toward this goal, we propose a methodology for the analysis of high-density oligonucleotide arrays. The gene expression measures are modeled as censored data to account for the quantification limits of the technology, and two gene selection criteria based on contrasts from an analysis of covariance (ANCOVA) model are presented. The model is formulated in a hierarchical Bayesian framework, which in addition to making the fit of the model straightforward and computationally efficient, allows us to borrow strength across genes. The elicitation of hierarchical priors, as well as issues related to parameter identifiability and posterior propriety, are discussed in detail. We examine the performance of our proposed method on simulated data, then present a detailed case study of an endometrial cancer dataset.

Biometry↗

Bayesian approaches to joint cure-rate and longitudinal models with applications to cancer vaccine trials.

Complex issues arise when investigating the association between longitudinal immunologic measures and time to an event, such as time to relapse, in cancer vaccine trials. Unlike many clinical trials, we may encounter patients who are cured and no longer susceptible to the time-to-event endpoint. If there are cured patients in the population, there is a plateau in the survival function, S(t), after sufficient follow-up. If we want to determine the association between the longitudinal measure and the time-to-event in the presence of cure, existing methods for jointly modeling longitudinal and survival data would be inappropriate, since they do not account for the plateau in the survival function. The nature of the longitudinal data in cancer vaccine trials is also unique, as many patients may not exhibit an immune response to vaccination at varying time points throughout the trial. We present a new joint model for longitudinal and survival data that accounts both for the possibility that a subject is cured and for the unique nature of the longitudinal data. An example is presented from a cancer vaccine clinical trial.

Bayes Theorem↗

Maximum likelihood methods for nonignorable missing responses and covariates in random effects models.

This article analyzes quality of life (QOL) data from an Eastern Cooperative Oncology Group (ECOG) melanoma trial that compared treatment with ganglioside vaccination to treatment with high-dose interferon. The analysis of this data set is challenging due to several difficulties, namely, nonignorable missing longitudinal responses and baseline covariates. Hence, we propose a selection model for estimating parameters in the normal random effects model with nonignorable missing responses and covariates. Parameters are estimated via maximum likelihood using the Gibbs sampler and a Monte Carlo expectation maximization (EM) algorithm. Standard errors are calculated using the bootstrap. The method allows for nonmonotone patterns of missing data in both the response variable and the covariates. We model the missing data mechanism and the missing covariate distribution via a sequence of one-dimensional conditional distributions, allowing the missing covariates to be either categorical or continuous, as well as time-varying. We apply the proposed approach to the ECOG quality-of-life data and conduct a small simulation study evaluating the performance of the maximum likelihood estimates. Our results indicate that a patient treated with the vaccine has a higher QOL score on average at a given time point than a patient treated with high-dose interferon.

Analysis of Variance↗