Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Analysing repeated measurements data: a practical comparison of methods.

A variety of methods are available for analysing repeated measurements data where the outcome is continuous. However, there is little information on how established methods, such as summary statistics and repeated measures analysis of variance (RMAOV), compare in practice with methods that have become available to applied statisticians more recently, such as marginal models (based on generalized estimating equation methodology) and multilevel models (that is, hierarchical random effects models). The aim of this paper is to exemplify the use of these methods, and directly compare their results by application to a clinical trial data set. The focus is on practical aspects rather than technical issues. The data considered were taken from a clinical trial of treatments for asthma in 240 children, in which a baseline and four post-randomization measurements of outcomes were taken. The simplicity of the method of summary statistics using the post-randomization mean of observations provided a useful initial analysis. However, fixed time effects or treatment-time interactions cannot be included in such an analysis, and choice of appropriate weighting when there is substantial missing data is problematic. RMAOV, marginal models and multilevel models generally provided similar estimates and standard errors for the treatment effects, although in one example with a relatively complex variance structure the marginal model produced less efficient estimates. Two advantages of multilevel models are that they provide direct estimates of variance components which are often of interest in their own right, and that they can be naturally extended to handle multivariate outcomes.

Adolescent↗

Latent-variable models for longitudinal data with bivariate ordinal outcomes.

We use the concept of latent variables to derive the joint distribution of bivariate ordinal outcomes, and then extend the model to allow for longitudinal data. Specifically, we relate the observed ordinal outcomes using threshold values to a bivariate latent variable, which is then modelled as a linear mixed model. Random effects terms are used to tie all together repeated observations from the same subject. The cross-sectional association between the two outcomes is modelled through the correlation coefficient of the bivariate latent variable, conditional on random effects. Assuming conditional independence given random effects, the marginal likelihood, under the missing data at random assumption, is approximated using an adaptive Gaussian quadrature for numerical integration. The model provides fixed effects parameters that are subject-specific, but retain the population-averaged interpretation when properly scaled. This is particularly well suited for the situation in which population comparisons and individual level contrasts are of equal importance. Data from a psychiatric trial, the Fluvoxamine (an antidepressant drug) study, are used to illustrate the methodology.

Antidepressive Agents, Second-Generation↗

Estimating paid and unpaid hours of personal assistance services in activities of daily living provided to adults living at home.

OBJECTIVE: To estimate the total hours of paid and unpaid personal assistance of daily living provided to adults living at home in the United States using nationally representative household survey data. DATA SOURCES: The Disability Followback Survey of the National Health Interview Survey on Disability (NHIS-D) conducted from 1994 to 1997. DATA COLLECTION/EXTRACTION METHODS: Data were obtained on persons receiving help with up to 5 ADLs and 10 IADLs, for up to 4 helpers, including the activities they helped with, whether the helper was paid or not, and the number of hours of help provided in the two weeks prior to the survey. The sample consists of 8,471 household-resident adults ages 18 and older receiving help with personal assistance. About 22 percent of the sample has missing data on hours, which we impute by multiple regression models using demographic, ADL, and IADL variables. FINDINGS: We estimate that 13.2 million noninstitutionalized adults receive an average of 31.4 hours per week of personal assistance in ADLs and IADLs per week, with 3.2 million people receiving an average of 17.6 hours of paid help and 11.7 million receiving an average of 30.7 hours of unpaid help. More persons ages 18-64 received help than those ages 65 and older (6.9 versus 6.2 million), but working-age recipients had fewer hours (27.4 versus 35.9) per week, due in part to less severe levels of disability. CONCLUSIONS: Personal assistance provided to adults with disabilities amounts to 21.5 billion hours of help per year, with an economic value in 1996 approaching $200 billion. Only 16 percent of this total is paid, representing $32 billion in home health services spent annually. This study, the first to estimate hours of assistance for both working-age and older adults, documents that older persons are more likely to receive paid personal assistance, while working-age people rely to a greater extent on unpaid help. This study begins to articulate the division of labor in the provision of personal assistance. Estimates of paid and unpaid hours of help by number of ADLs should inform policy concerning eligibility boundaries in long term care.

Activities of Daily Living↗

Blood pressure, LDL cholesterol, and intima-media thickness: a test of the "response to injury" hypothesis of atherosclerosis.

The "response to injury" hypothesis is a plausible model of the development of atherosclerosis supported by observations from animal models. The present study uses epidemiological data to investigate the hypothesis that wall damage due to hypertension is a precursor of low density lipoprotein cholesterol (LDL-C)-mediated atherosclerosis. The Los Angeles Atherosclerosis Study is following a cohort of 576 participants who were aged 40 to 60 years and were free of symptomatic cardiovascular disease at recruitment. Common carotid artery intima-media thickness (IMT) was assessed by B-mode ultrasonography. After exclusion for nonfasting blood draw and other missing data, 511 subjects were available for analysis. IMT was regressed on LDL-C within tertiles of systolic blood pressure (SBP): low (93 to 122 mm Hg), middle (123 to 132 mm Hg), and high (133 to 175 mm Hg). Covariates were age, sex, body height, body mass index, ethnicity, smoking status, diabetes, and pharmacological treatment for hypertension or hypercholesterolemia. IMT was significantly related to LDL-C in the high SBP group (beta=0.025+/-0.008, where beta values are IMT [mm]/LDL-C [mmol/L]; P=0.002) but not in the middle (beta=-0.006+/-0.008, P=0.39) or low (beta=-0.004+/-0.009, P=0.64) SBP group. The slope in the high SBP group was significantly greater than in the middle (P=0.004) or low (P=0.014) SBP group. Results were similar for women and men, and after the exclusion of diabetics and persons using antihypertensive or lipid-lowering medications. Elevated LDL-C was associated with increased IMT in the upper tertile of SBP but not in the lower tertiles. These findings are consistent with the hypothesis that wall injury due to elevated SBP increases the susceptibility of the artery wall to LDL-C-mediated atherogenesis.

Adult↗

Testing efficacy with detection controlled estimation: an application to telemedicine.

Detection controlled estimation (DCE) is a powerful new econometric estimator in the family of missing data estimators. By collecting measures from a variety of inspectors or inspection technologies, DCE is able to make inferences about the entire population, even when that population is not directly observed. Using this innovative method, we were able to assess whether telemedicine technology could be substituted for in-person visits when providing maintenance care for patients with hypertension. Our findings indicate that there is no support for the proposition that telemedicine is less effective than in-person visits for determining whether patients have high blood pressure. Indeed, our results imply that telemedicine misses 7% fewer cases of high blood pressure than in-person visits do. The results of this study indicate that DCE may be an effective tool for use in cost-effectiveness or cost-benefit analysis in health care.

Aged↗

Accounting for linkage disequilibrium among markers in linkage analysis: impact of haplotype frequency estimation and molecular haplotypes for a gene in a candidate region for Alzheimer's disease.

OBJECTIVES: Linkage disequilibrium (LD) between closely spaced SNPs can be accommodated in linkage analysis by specifying the multi-SNP haplotype frequencies, if known. Phased haplotypes in candidate regions can provide gold standard haplotype frequency estimates, and may be of inherent interest as markers. We evaluated the effects of different methods of haplotype frequency estimation, and the use of marker phase information, on linkage analysis of a multi-SNP cluster in a candidate region for Alzheimer's disease (AD). METHODS: We performed parametric linkage analysis of a five-SNP cluster in extended pedigrees to compare the use of: (1) haplotype frequencies estimated by molecular phase determination, maximum likelihood estimation, or by assuming linkage equilibrium (LE); (2) AD families or controls as the frequency source; and (3) unphased or molecularly phased SNP data. RESULTS: There was moderate to strong pairwise LD among the five SNPs. Falsely assuming LE substantially inflated the LOD score, but the method of haplotype frequency estimation and particular sample used made little difference provided that LD was accommodated. Use of phased haplotypes produced a modest increase in the LOD score over unphased SNPs. CONCLUSIONS: Ignoring LD between markers can lead to substantially inflated evidence for linkage in LOD score analysis of extended pedigrees with missing data. Use of marker phase information in linkage analysis may be important in disease studies where the costs of family recruitment and phenotyping greatly exceed the costs of phase determination.

Adult↗

Data reconciliation, structure analysis and simulation of waste flows: case study Vienna.

The management of complex waste flow systems requires a systematic approach for the handling of data, for obtaining a consistent picture of the system under consideration, and for simulating various policy scenarios and evaluating material control strategies. In this paper the implementation of a useful methodology is presented, which has been developed in previous works and is further enhanced for modelling, identifying, analysing and simulating material flow systems for which at most one measurement per flow is available for a single balancing period. The methodology enables the analyst to cope with missing data and uncertainty in the measurements. A data reconciliation procedure is used to minimise the uncertainty concerning flows by exploiting the redundancies created by restricting the available data to fulfil the available structural information. Statistical tests are introduced to enable the user to check the compatibility of the data with the a priori information. The origins analysis and destination analysis tools allow for a deeper insight into the system structure. Policy scenarios can be treated using the simulation tools. The waste flow system of the city of Vienna has been chosen to demonstrate step-by-step the procedure for building a reliable model and the effective application of the above mentioned methods and tools. Current and future research focuses on models balancing different interrelated quantities simultaneously and on incorporating stock accumulation and depletion behaviour.

Austria↗

Bayesian analysis of gene expression levels: statistical quantification of relative mRNA level across multiple strains or treatments.

BACKGROUND: Methods of microarray analysis that suit experimentalists using the technology are vital. Many methodologies discard the quantitative results inherent in cDNA microarray comparisons or cannot be flexibly applied to multifactorial experimental design. Here we present a flexible, quantitative Bayesian framework. This framework can be used to analyze normalized microarray data acquired by any replicated experimental design in which any number of treatments, genotypes, or developmental states are studied using a continuous chain of comparisons. RESULTS: We apply this method to Saccharomyces cerevisiae microarray datasets on the transcriptional response to ethanol shock, to SNF2 and SWI1 deletion in rich and minimal media, and to wild-type and zap1 expression in media with high, medium, and low levels of zinc. The method is highly robust to missing data, and yields estimates of the magnitude of expression differences and experimental error variances on a per-gene basis. It reveals genes of interest that are differentially expressed at below the twofold level, genes with high 'fold-change' that are not statistically significantly different, and genes differentially regulated in quantitatively unanticipated ways. CONCLUSIONS: Anyone with replicated normalized cDNA microarray ratio datasets can use the freely available MacOS and Windows software, which yields increased biological insight by taking advantage of replication to discern important changes in expression level both above and below a twofold threshold. Not only does the method have utility at the moment, but also, within the Bayesian framework, there will be considerable opportunity for future development.

Bayes Theorem↗

Differential diagnosis of jaundice: applicability of the Copenhagen Pocket Chart proved in Stockholm patients.

This paper shows that an algorithm for differential diagnosis of jaundice developed in Denmark has been successfully transferred for use in a Swedish hospital. The algorithm, which is based on data from nearly 1000 patients, utilises 21 items of information from the medical history, physical examination and blood chemistry. The algorithm recognises four diagnostic groups: benign obstructive jaundice, malignant obstructive jaundice, acute non-obstructive jaundice, and chronic non-obstructive jaundice. To each item of information, a score is attached reflecting its weight of evidence. Summing the scores for the symptoms and signs that are present leads to a probabilistic statement about the diagnosis. Because of missing data in the Swedish patient material, three of the items were excluded from the original algorithm. Corrections were made for differences in the distribution of diseases. In reclassification of 985 Danish patients the modified algorithm's "best bid", i.e. the diagnosis given the highest probability, was correct in 78% of cases. More important, 93% of the cases given a "confident" diagnosis (probability greater than 0.80) were correct. The corresponding figures when the algorithm was applied to Swedish patients were 76% and 93%, respectively. In both series the predicted probabilities were matched by a corresponding proportion of actual diagnostic hits. It is concluded that the algorithm leads to reliable estimates of diagnostic probabilities in jaundice and that the algorithm seems to work well in Sweden also.

Adult↗

Characteristics of replicated single-nucleotide polymorphism genotypes from COGA: Affymetrix and Center for Inherited Disease Research.

Genetic Analysis Workshop 14 provided re-genotyped single-nucleotide polymorphism (SNP) data. Specifically, both Center for Inherited Disease Research (CIDR) and Affymetrix genotyped the same 11,560 SNPs from the Affymetrix GeneChip Mapping 10K Array marker set on the same 184 individuals from the Collaborative Study on the Genetics of Alcoholism database. While the inconsistency rate between CIDR and Affymetrix (two different genotypes for the same subject) was low (0.2%), the non-replication rate (two different genotypes for the same subject or one identified genotype and one missing genotype) was substantial (9.5%). The missing data could be from no-call regions, which is inconsistent with recent recommendations about the use of no-call regions in association tests. In addition, no-call regions would suggest that the actual inconsistency rate is higher than reported. A high inconsistency rate has significant impact on power in related hypothesis tests. In addition, the data are consistent with assumptions made in a recently proposed likelihood ratio test of association for re-genotyped data.

Alcoholism↗

The implementation of prompted retinal screening for diabetic eye disease by accredited optometrists in an inner-city district of North London: a quality of care study.

Diabetic retinopathy remains the most common cause of blindness in people of working age but the provision of high quality eye screening for diabetic patients is still erratic in many health districts in the UK. National consensus guidelines recommend comprehensive population coverage, high sensitivity (>80%), high specificity (>95%), agreed clinical criteria, referral procedures and centralized data collection to facilitate audit. This study looks at the effectiveness of implementing a prompted recall programme for retinal screening in an inner-city district of North London. The scheme uses trained, accredited optometrists to screen patients with diabetes who are looked after in the community by their general practitioner. During the first 17 months of the scheme, 63 optometrists attended training and gained accreditation. Of the 666 patients recruited, 645 were scheduled for screening and 536 (83%) attended. Fourteen per cent of patients screened were found to have background retinopathy and 2.3% sight-threatening eye disease. In two audits, carried out 15 months apart in a random sample of GP practices, the incidence of recorded dilated fundoscopy increased from 48% at baseline to 56%, an increase of 8% (95% CIs 2%-14%). For referable eye disease, the sensitivity of this screening technique was 100%, the specificity 94% (95% CIs 90%-98%), the positive predictive value 79% (95% CIs 72%-86%) and the negative predictive value 100%. The administrative cost per case screened was Pound Sterling 12.60 (excluding clinical costs and any additional optometry payment).

Diabetic Retinopathy↗

Depression, depressive symptoms and mortality in persons aged 65 and over living in the community: a systematic review of the literature.

BACKGROUND: No recent attempt has been made to synthesize information on mortality and depression despite the theoretical and practical interest in the topic. Our objective was to estimate in the older population the influence on mortality of depression and depressive symptoms. METHODS: Data sources were: Medline, Embase, personal files and colleagues' records. Studies were considered if they included a majority of persons aged > or = 65 years at baseline either drawn from a total community sample or drawn from a random sample from the community. Samples from healthcare facilities were excluded. Effect sizes were extracted from the papers; if they were not included in the published papers, effect sizes were calculated if possible. No attempt was made to contact authors for missing data. RESULTS: We found 21 reports on 23 cohorts using depression diagnosis. For 15 of these, odds ratios were pooled using the Greenland method based on confidence intervals (CIs), giving an estimated odds ratio for mortality with depression of 1.73 (95% CI 1.53 to 1.95). A fixed effects meta-regression of these studies suggested that longer follow-up predicted smaller effect sizes (log odds ratios -0.096 per year (95% CI -0.179 to -0.014)). There is a weak suggestion of a reduced effect of depression on mortality for women. We were unable to pool effect sizes from the 17 studies using symptom totals and scales, or from eight studies of specific symptoms. CONCLUSIONS: The studies show that diagnosed depression in community-resident older people is associated with increased mortality. The picture for sex differences is still unclear.

Aged↗

Effect of a statewide neonatal resuscitation training program on Apgar scores among high-risk neonates in Illinois.

OBJECTIVE: The national Neonatal Resuscitation Program (NRP), started in 1987, provided training to hospital delivery room personnel to standardize knowledge and skills to reduce neonatal morbidity and mortality and increase successful resuscitation during the first few critical minutes after birth. The Apgar score continues to be used as the best established index of immediate postnatal health. The purpose of this study was to evaluate the impact of the NRP instruction in Illinois hospitals by examining Apgar scores among high-risk infants who are likely to benefit from the NRP. METHODS: A retrospective 3-time period cohort design was used (before the introduction of the NRP, 1985-1988; transition when NRP training occurred, 1989-1990; and after NRP training was completed at least once for some delivery room personnel in each Illinois hospital, 1991-1995). Illinois computerized birth certificate files on a selected group of 636 429 high-risk neonates provided information on Apgar scores and maternal characteristics. The American Academy of Pediatrics provided instructor lists to determine when NRP training started and when it was fully implemented in Illinois. Illinois Department of Public Health provided data to categorize hospitals into levels based on type and intensity of neonatal services (Level I, II, II+, III). High-risk neonates were defined as meeting 1 of the following criteria: maternal age <20 years old or >35 years old, birth weight <2500 g or >4000 g, presence of a maternal medical risk factor, and no prenatal care or prenatal care started after the first trimester. Several exclusion criteria were applied including the following: birth records with missing data, multiple birth or congenital anomaly, and hospital information that indicate no birth deliveries in 1 of the 11 study years or delivery outside of a hospital. One-minute and 5-minute Apgar scores were divided into categories for analysis (0-3, 4-6, 7-10). No change or a decrease in a low (0-6) 1-minute Apgar when compared with the 5-minute Apgar was a primary measure to evaluate effect of NRP resuscitation. Variables examined included the following: race/ethnicity, maternal age, level of education, presence of maternal medical risk factor, trimester started prenatal care, complications of labor and delivery, and a low birth weight. Analysis consisted of chi(2) tests, relative risk calculations, and logistic regression to reveal independent associations with no change in low 1-minute Apgar score or continued low (0-6) 5-minute Apgar. RESULTS: A total of 636 429 high-risk birth records was selected for detailed analyses out of 2 077 533 births in Illinois between 1985 and 1995 for 193 hospitals. The number of active NRP instructors in Illinois changed dramatically during the study period; for example, 1 to 6 between 1987 and 1988 to 1096 to 1242 between 1991 and 1995. The percentage of neonates reported to have low (<7) 1-minute Apgar score decreased in 1991 to 1995 overall and for each of 4 hospital levels. Overall and by hospital level, there was a statistically significant lower proportion of high-risk newborns who showed a decrease or no change in their 5-minute Apgar scores after the NRP instruction. After adjusting for several maternal characteristics, logistic regression analysis revealed that high-risk newborns with a low 1-minute Apgar were more likely to increase their 5-minute Apgar after the NRP instruction in 1991 to 1995. Additional analyses indicated that very low birth weight and low birth weight newborns benefited the most from NRP instruction. CONCLUSION: Although previous research has shown that the NRP instruction improves knowledge and skill among health care personnel in the delivery room, both short-term and long-term, there has been little evidence to demonstrate NRP impact on infant morbidity. Several strategies were used in this study to control for bias and to adjust for secular trends in decreased infant morbidity during the study period. This study demonstrated sufficient support for the hypothesis that a significant improvement occurred among neonates in their Apgar score after the NRP instruction in Illinois. Empirical support is provided for the clinical effectiveness of NRP instruction.

Apgar Score↗

Hospital policies and their influence on newborn body weight.

On the basis of data collected in a survey of the practices in maternity wards in Poland, we analysed newborn body weight differences between two groups of hospitals: one with the highest percentage of exclusive breastfeeding and other supportive practices, the other with the lowest. The aim of the study was to investigate whether hospital procedures can influence the newborn weight profile in the first days after birth. Healthy infants--normal birth, birthweight > or = 2500 g, discharged within 2-7 days, no missing data--were chosen for the analyses. The difference between discharge and birthweight, percentage of birthweight loss or gain on the day of discharge and percentage of the infants who at least regain birthweight on the day of discharge were compared between the two groups of hospitals. All analyses indicated the positive influence of exclusive breastfeeding and practices supporting it on infant body weight profile while in hospital.

Body Weight↗

Marginal regression models with a time to event outcome and discrete multiple source predictors.

Information from multiple informants is frequently used to assess psychopathology. We consider marginal regression models with multiple informants as discrete predictors and a time to event outcome. We fit these models to data from the Stirling County Study; specifically, the models predict mortality from self report of psychiatric disorders and also predict mortality from physician report of psychiatric disorders. Previously, Horton et al. found little relationship between self and physician reports of psychopathology, but that the relationship of self report of psychopathology with mortality was similar to that of physician report of psychopathology with mortality. Generalized estimating equations (GEE) have been used to fit marginal models with multiple informant covariates; here we develop a maximum likelihood (ML) approach and show how it relates to the GEE approach. In a simple setting using a saturated model, the ML approach can be constructed to provide estimates that match those found using GEE. We extend the ML technique to consider multiple informant predictors with missingness and compare the method to using inverse probability weighted (IPW) GEE. Our simulation study illustrates that IPW GEE loses little efficiency compared with ML in the presence of monotone missingness. Our example data has non-monotone missingness; in this case, ML offers a modest decrease in variance compared with IPW GEE, particularly for estimating covariates in the marginal models. In more general settings, e.g., categorical predictors and piecewise exponential models, the likelihood parameters from the ML technique do not have the same interpretation as the GEE. Thus, the GEE is recommended to fit marginal models for its flexibility, ease of interpretation and comparable efficiency to ML in the presence of missing data.

Biometry↗

Foundations for health status metrology: the stability of MOS SF-36 PF-10 calibrations across samples.

Interest in applying probabilistic conjoint measurement (PCM) models, such as those devised by the late Georg Rasch, to health status and quality of life data has grown significantly in the last few years. Applications have yet, however, to fully realize the opportunities for scientific generalization and practical convenience PCM offers. This article fleshes out the substance of some of these opportunities by comparing eight separate PCM calibrations of the SF-36 ten-item physical functioning scale (PF-10). The initial average correlation across the 28 pairs of calibrations is .84; after taking advantage of the PCM model's capacity to account for missing data by omitting from the comparisons items that vary due to sample idiosyncracies, the average correlation is .90. Opportunities for, and limitations on, generalization from PF-10 measures are explored.

Activities of Daily Living↗

Methods to adjust for bias and confounding in critical care health services research involving observational data.

Observational data are often used for research in critical care. Unlike randomized controlled trials, where randomization theoretically balances confounding factors, studies involving observational data pose the challenge of how to adjust appropriately for the bias and confounding that are inherent when comparing two or more groups of patients. This paper first highlights the potential sources of bias and confounding in critical care research and then reviews the statistical techniques available (matching, stratification, multivariable adjustment, propensity scores, and instrumental variables) to adjust for confounders. Finally, issues that need to be addressed when interpreting the results of observational studies, such as residual confounding, causality, and missing data, are discussed.

Bias↗

Mapping trait loci by use of inferred ancestral recombination graphs.

Large-scale association studies are being undertaken with the hope of uncovering the genetic determinants of complex disease. We describe a computationally efficient method for inferring genealogies from population genotype data and show how these genealogies can be used to fine map disease loci and interpret association signals. These genealogies take the form of the ancestral recombination graph (ARG). The ARG defines a genealogical tree for each locus, and, as one moves along the chromosome, the topologies of consecutive trees shift according to the impact of historical recombination events. There are two stages to our analysis. First, we infer plausible ARGs, using a heuristic algorithm, which can handle unphased and missing data and is fast enough to be applied to large-scale studies. Second, we test the genealogical tree at each locus for a clustering of the disease cases beneath a branch, suggesting that a causative mutation occurred on that branch. Since the true ARG is unknown, we average this analysis over an ensemble of inferred ARGs. We have characterized the performance of our method across a wide range of simulated disease models. Compared with simpler tests, our method gives increased accuracy in positioning untyped causative loci and can also be used to estimate the frequencies of untyped causative alleles. We have applied our method to Ueda et al.'s association study of CTLA4 and Graves disease, showing how it can be used to dissect the association signal, giving potentially interesting results of allelic heterogeneity and interaction. Similar approaches analyzing an ensemble of ARGs inferred using our method may be applicable to many other problems of inference from population genotype data.

Algorithms↗