Search PubMedSearch

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Suggestive genome-wide associations with inflammatory biomarkers in an admixed population, including a missense variant in the OR6K6 olfactory receptor gene associated with MCP-1.

BACKGROUND: Chronic low-grade inflammation drives cardiometabolic diseases and has a strong genetic basis. Most genome-wide association studies (GWAS) have focused on European populations, limiting knowledge of the genetic influences on inflammation in admixed populations such as those in Brazil. METHODS: This study is part of the cross-sectional ISA Capital Health Survey. It uses data from the 2015 ISA Nutrition cohort, which measured biochemical, genetic, anthropometric, and lifestyle factors in a probabilistic sample of São Paulo residents. Genomic DNA was extracted from 841 individuals. Genotyping was performed using the Axiom 2.0 Precision Medicine Research Array. After quality control and missing data exclusion, 244,338 SNPs from 638 individuals remained for GWAS-based association analysis with eight inflammatory biomarkers. Models were adjusted for sex, age, age2, overweight, and the first two principal components of ancestry. RESULTS: Most participants were male (53%) and not overweight (55%). The median age was 49, and 38% were older adults. In the genome-wide analysis of TNF-α, IL-10, IL-1β, monocyte chemoattractant protein-1 (MCP-1), and adiponectin, 12 SNPs were significantly associated, most of which were intronic. Notably, one signal mapped to the missense variant rs16841009 in the olfactory receptor gene OR6K6. This variant was associated with MCP-1, suggesting a possible involvement in inflammatory responses. CONCLUSIONS: We identified new SNPs linked to inflammatory biomarkers in a highly admixed Brazilian population, including a missense variant in an olfactory receptor gene linked to MCP-1. This association may be biologically important for inflammation and could affect the risk of cardiometabolic diseases.

Humans

An indicator of adverse pregnancy outcome in France: not receiving maternity benefits.

STUDY OBJECTIVE: The aim was to compare the social characteristics, the pregnancy outcome, and the antenatal care of women in France who did not receive maternity benefits to women who did. These benefits (860 FF, approx 86 pounds per month) are given to every pregnant woman, starting in the second trimester. Payments are made on the condition that at least three antenatal visits are made, the first being before the end of the first trimester. DESIGN: The study involved a random sample of women who were interviewed after delivery during their stay in hospital. Data on pregnancy outcome were collected from medical records. SETTING: The study was carried out in four public maternity units in different regions of France. PARTICIPANTS: 1692 women were included in the analysis (86.8% of the selected sample). Of 257 exclusions, 40 had multiple pregnancies, 189 had missing data, and 28 did not answer the question concerning maternity benefits. MEASUREMENTS AND MAIN RESULTS: 4.3% of the women did not receive any maternity benefits. These women lived in poorer social conditions than the women who received the benefits. They had a higher preterm delivery rate, after controlling for risk factors in a logistic regression. Women without maternity benefits were characterised by a lower level of care, yet the majority began their antenatal care during the first trimester or had more than six visits. CONCLUSIONS: Not receiving maternity benefits during pregnancy is an index of an underprivileged situation and a risk factor for pregnancy outcome.

Age Factors

A cautionary note on the use of autoregressive models in analysis of longitudinal data.

Rosner et al. presented a simple, easily implemented modelling method for retaining time order relationships in analyses of longitudinal data when successive measures are correlated. Evaluation of time order is particularly useful in epidemiologic studies concerned with exposure to potentially toxic substances and subsequent outcome, but may also have use in more traditional growth studies that relate intake to subsequent development. The analysis allows for unequally spaced measures and missing data. The estimation method permits varying numbers of observations per subject and, with measures equally spaced, one can fit the model with use of ordinary least squares regression software. We report on a potential false association that can result when both exposure and outcome are related to time. We illustrate this problem with a small scale simulation and example. We also note a more serious problem with Rosner's approach in interpreting parameters. Although the model may be useful for prediction, parameters depend on the autocorrelation and are not readily interpretable. We recommend alternative modelling strategies be used when autocorrelation of errors is suspected.

Age Factors

Running average analysis of clinical trial ambulatory blood pressure data.

A method is presented for analyzing ambulatory blood pressure monitoring (ABPM) time series data obtained from well-controlled clinical trials. The method uses running averages based on fixed time-of-day intervals (rather than a fixed number of neighboring measurements). These "interval running averages" effectively estimate average blood pressure during the specified time intervals, adjusting for unequal spacing between measurements, embedded missing data, varying measurement times-of-day, and doses of study medication taken during ABP monitoring. Blood pressure changes from baseline may be computed using the interval running averages in order to separate treatment effects from patients' normal daily blood pressure cycles. To ensure valid estimation of treatment effects over time, study medication dosing times should be rigorously controlled in the trial design and conduct. Interval running average curves may be presented graphically, and from them summary statistics may be computed for purposes of statistical analysis. By allowing for the inherent complications of ABP data collection, the effect of antihypertensive treatment in well-controlled clinical trials can be discerned.

Ambulatory Care

Cardiometabolic Multimorbidity Increases the Risk of Hip Fracture: A Longitudinal Cohort Study Based on CHARLS.

BACKGROUND: Cardiometabolic Multimorbidity (CMM) is defined as the co-occurrence of two or more conditions among heart disease, diabetes mellitus, stroke, and hypertension. Previous studies have shown associations between cardiometabolic diseases and fragility fractures; however, the relationship between CMM and hip fractures remains unclear in the Chinese population. This study therefore aims to investigate this association in a Chinese cohort to inform fracture prevention strategies. METHODS: This prospective cohort study used data from the China Health and Retirement Longitudinal Study (CHARLS) collected from 2011 to 2020. Participants from the 2011 baseline survey cohort were initially included. Subsequently, individuals were sequentially excluded if they were under 45 years of age, had incomplete baseline CMM information, had a history of hip fracture, lost to follow-up, or had missing data on confounders. Kaplan-Meier survival analysis, Cox proportional hazards regression, subgroup analyses, and sensitivity analyses were performed to evaluate the association between CMM and the risk of hip fracture. RESULTS: A total of 6314 participants aged 45 years and older were included, of whom 544 had CMM. Over a 9-year follow-up period, 287 incident hip fractures (4.55%) were identified. Among these, 36 participants had been diagnosed with CMM at baseline, whereas 251 had not. The incidence of hip fracture was significantly higher in participants with CMM than in those without CMM (13% vs. 8%, p = 0.015). After full adjustment for confounders, multivariable Cox regression showed that CMM was associated with a 70% increased risk of hip fracture (HR = 1.70, 95% CI: 1.318-2.47; p = 0.005). Subgroup analyses indicated that age and history of falls were significant effect modifiers. The association between CMM and hip fracture was more pronounced in participants under 60 years old (P for interaction = 0.048) and those with a history of falls (P for interaction = 0.014). CONCLUSION: These findings suggest that CMM increases the risk of hip fracture, particularly among relatively younger individuals and those with a history of falls.

Humans

Validity and reliability of a high school drug use questionnaire among Mexican students.

This study was carried out to ascertain the validity and reliability of data generated through an internationally developed self-administered questionnaire. Data were collected from two student populations in which prevalence of drug use was known (high and low prevalence rates) and information obtained from the questionnaire checked through personal interview. The questionnaire covered both demographic characteristics and drug use items. A total of 474 students were tested and 335 students (70.7 per cent) retested. Although percentages of missing data and inconsistent responses were high, the questionnaire appeared valid and reliable when analysed on a group basis. It was less reliable on an individual basis. In conclusion, it is considered that with minor modifications the questionnaire can be used in the populations studied with enough confidence in validity and reliability to allow comparisons.

Adolescent

Practice databases and their uses in clinical research.

A few large clinical information databases have been established within larger medical information systems. Although they are smaller than claims databases, these clinical databases offer several advantages: accurate and timely data, rich clinical detail, and continuous parameters (for example, vital signs and laboratory results). However, the nature of the data vary considerably, which affects the kinds of secondary analyses that can be performed. These databases have been used to investigate clinical epidemiology, risk assessment, post-marketing surveillance of drugs, practice variation, resource use, quality assurance, and decision analysis. In addition, practice databases can be used to identify subjects for prospective studies. Further methodologic developments are necessary to deal with the prevalent problems of missing data and various forms of bias if such databases are to grow and contribute valuable clinical information.

Clinical Medicine

A note on data monitoring, incomplete data and curtailed testing.

A question often asked in a fixed sample size clinical trial is whether the final outcome has been determined before the data has been completely collected. Recent papers in this journal and elsewhere have considered this question from the viewpoint of curtailed testing. In this paper, a related concept, "the probability of reversal" of a test procedure, is described and associated least upper bounds are given that are pertinent to the issue. The results are also of relevance in assessing, at the end of the trial, whether missing data are of consequence in the test of significance.

Clinical Trials as Topic

Linked cross-sectional study for evaluating the effect on the microfilarial load of the onchocerciasis control programme.

When analysing incomplete longitudinal data (ie data which includes complete longitudinal results plus cross-sectional results plus partially longitudinal results) one can increase the sensitivity of the data by taking the correlation structure of the data into account. For the described data on onchocerciasis, the adapted method of Rao and Rao, called the Linked Cross-Sectional method, which estimates the mean microfilarial load for each passage and hence the differences between passages, taking into account the correlation between the same individuals measured at different times is such a method that reduces significantly the standard deviations of the mean responses and hence increases their sensitivity. The method takes into account the missing data nature of the results. It can be simply generalized to any number of passages.

Cross-Sectional Studies

[Reduction of cerebral hemorrhage and respiratory distress syndrome in premature infants by avoiding perinatal asphyxia].

Intra- and periventricular haemorrhage (IVH/PVH) and, under certain conditions, the respiratory distress syndrome (RDS) seem to be typical sequelae of perinatal asphyxia in preterm born infants. Therefore, an association of IVH/PVH and RDS can be expected. We have retrospectively analyzed the data of 118 premature infants born between 1982 and 1986, weighing between 750 and 1499 g. 11 of these had experienced a severe IVH/PVH and a severe RDS at the same time, whereas 75 infants did not develop either of those. (2 of the 118 showed a severe IVH/PVH without evidence of severe RDS whereas 29 developed severe RDS without signs of serious IVH/PVH. 1 could not be evaluated due to missing data). This association of severe intracerebral haemorrhage and severe respiratory distress syndrome was statistically significant (p less than 0.005). The number of severe IVH/PVH has decreased during 1984-1986 in comparison to 1982/83 (4/76 vs. 9/42; p less than 0.05); the incidence of severe RDS has slightly declined. Comparing the perinatal conditions we found that the infants of the years 1984-1986 were more rarely delivered after an interval exceeding 24 h after premature rupture of the membranes (p less than 0.05), were more often delivered by caesarean section (p less than 0.005), and were nearly always primarily cared for by an experienced paediatrician (p less than 0.01). There were no significant differences between these two groups as far as dexamethasone-prophylaxis, mean birth weight, percentage of small-for-gestational-age infants and mean Apgar scores were concerned.(ABSTRACT TRUNCATED AT 250 WORDS)

Asphyxia Neonatorum

Differential diagnosis of jaundice: applicability of the Copenhagen Pocket Chart proved in Stockholm patients.

This paper shows that an algorithm for differential diagnosis of jaundice developed in Denmark has been successfully transferred for use in a Swedish hospital. The algorithm, which is based on data from nearly 1000 patients, utilises 21 items of information from the medical history, physical examination and blood chemistry. The algorithm recognises four diagnostic groups: benign obstructive jaundice, malignant obstructive jaundice, acute non-obstructive jaundice, and chronic non-obstructive jaundice. To each item of information, a score is attached reflecting its weight of evidence. Summing the scores for the symptoms and signs that are present leads to a probabilistic statement about the diagnosis. Because of missing data in the Swedish patient material, three of the items were excluded from the original algorithm. Corrections were made for differences in the distribution of diseases. In reclassification of 985 Danish patients the modified algorithm's "best bid", i.e. the diagnosis given the highest probability, was correct in 78% of cases. More important, 93% of the cases given a "confident" diagnosis (probability greater than 0.80) were correct. The corresponding figures when the algorithm was applied to Swedish patients were 76% and 93%, respectively. In both series the predicted probabilities were matched by a corresponding proportion of actual diagnostic hits. It is concluded that the algorithm leads to reliable estimates of diagnostic probabilities in jaundice and that the algorithm seems to work well in Sweden also.

Adult

Analyzing laboratory marker changes in AIDS clinical trials.

Repeated measurements of laboratory markers of immunologic or disease status, such as CD4 lymphocyte counts and HIV p24 antigen levels, can be important end points in comparative clinical trials. In this report, we consider comparison of treatment groups with respect to such markers, focusing on a distribution-free approach in which each participant's data are characterized by a single summary statistic. The summary statistics examined are (a) the slope of the least-squares regression of the marker, (b) the average of the last r measurements, and (c) the difference between the averages of the last r and the first s measurements. Under various models of marker time trends, these methods are compared with regard to statistical power. It is found that the slope is usually more efficient than the other two types of summaries. Adaptations for missing data are discussed and illustrated in an analysis of CD4 counts from a recent AIDS clinical trial.

Acquired Immunodeficiency Syndrome

A maximum likelihood method for estimating genome length using genetic linkage data.

The genetic length of a genome, in units of Morgans or centimorgans, is a fundamental characteristic of an organism. We propose a maximum likelihood method for estimating this quantity from counts of recombinants and nonrecombinants between marker locus pairs studied from a backcross linkage experiment, assuming no interference and equal chromosome lengths. This method allows the calculation of the standard deviation of the estimate and a confidence interval containing the estimate. Computer simulations have been performed to evaluate and compare the accuracy of the maximum likelihood method and a previously suggested method-of-moments estimator. Specifically, we have investigated the effects of the number of meioses, the number of marker loci, and variation in the genetic lengths of individual chromosomes on the estimate. The effect of missing data, obtained when the results of two separate linkage studies with a fraction of marker loci in common are pooled, is also investigated. The maximum likelihood estimator, in contrast to the method-of-moments estimator, is relatively insensitive to violation of the assumptions made during analysis and is the method of choice. The various methods are compared by application to partial linkage data from Xiphophorus.

Animals

BACTID: a microcomputer implementation of a PASCAL program for bacterial identification based on Bayesean probability.

A computer program (BACTID) is described which enables the identification of bacteria based on a priori data and Bayesean probability testing. The program is not limited to a specific format, has a short execution time, can be easily applied to a variety of situations, and can be run on almost any microcomputer system operating under either 8-bit CP/M or 16-bit MS-DOS or PC-DOS. Additionally, BACTID is not limited to one type of computer (hardware independent); is not limited by size of the computer's random access memory (RAM independent); can recognize various database matrices (format independent); is able to compensate for missing data; and allows for various methods of data entry. The efficacy of the program was checked against a commercially available test system and a 99.34% agreement was obtained. Also, the execution time for a 46 x 21 data matrix was as little as 3.5 seconds. These results show that microcomputer identification programs not only are viable alternatives to code-book registers, but also offer flexibility which is not found in commercial systems.

Bacteria

Accounting for the correlation between fellow eyes in regression analysis.

Regression techniques that appropriately use all available eyes have infrequently been applied in the ophthalmologic literature, despite advances both in the development of statistical models and in the availability of computer software to fit these models. We considered the general linear model and polychotomous logistic regression approaches of Rosner and the estimating equation approach of Liang and Zeger, applied to both linear and logistic regression. Methods were illustrated with the use of two real data sets: (1) impairment of visual acuity in patients with retinitis pigmentosa and (2) overall visual field impairment in elderly patients evaluated for glaucoma. We discuss the interpretation of coefficients from these models and the advantages of these approaches compared with alternative approaches, such as treating individuals rather than eyes as the unit of analysis, separate regression analyses of right and left eyes, or utilization of ordinary regression techniques without accounting for the correlation between fellow eyes. Specific advantages include enhanced statistical power, more interpretable regression coefficients, greater precision of estimation, and less sensitivity to missing data for some eyes. We concluded that these models should be used more frequently in ophthalmologic research, and we provide guidelines for choosing between alternative models.

Adolescent

[Prognostic factors in operated non-small cell cancer of the lung. Study from a randomized therapeutic trial].

This article presents the results of a prognostic study of primary resected lung cancer (non-small cell). The data result from a randomised clinical trial of immunotherapy with a non-specific adjuvant; the follow-up was between four to seven years. Thirty-five clinical, biological and anatomo-pathological parameters were gathered at the time of inclusion in the trial. The response criteria used were survival without recurrence and total survival. A multivariate analysis using the Cox's model was carried out for each criterion. At the reference date of the 1st April 1985, 125 relapses and 132 deaths were counted amongst 219 patients; there was only one patient lost to follow-up and only 39 missing data were observed. The negative therapeutic results of the immunotherapy used were confirmed by this new intermediate analysis. The rate of survival without recurrence at 5 years was 43% and the overall survival at five years was 42%. The use of Cox's model to show the prognostic information at the 5% level for survival without recurrence could be summarised by five factors: main staging (the prognostic factor), leucocytosis, the cutaneous reaction to proteus, Karnofsky index and presence of physical signs. For stages I and II the outcome was identical and no factor was predictive at the 5% level. For stage III the cutaneous reaction to proteus and leucocytosis were prognostic. For overall survival, the prognostic information at the 5% level could be summarised by five factors: staging (main prognostic factor), leucocytosis, Karnofsky index, presence of physical signs and lymphocytosis. For stages I and II whose outcome was identical only Karnofsky index and lymphocytosis were predictive at the 5% level.(ABSTRACT TRUNCATED AT 250 WORDS)

Carcinoma, Non-Small-Cell Lung

Intentionally incomplete longitudinal designs: I. Methodology and comparison of some full span designs.

Longitudinal designs are important in medical research and in many other disciplines. Complete longitudinal studies, in which each subject is evaluated at each measurement occasion, are often very expensive and motivate a search for more efficient designs. Recently developed statistical methods foster the use of intentionally incomplete longitudinal designs that have the potential to be more efficient than complete designs. Mixed models provide appropriate data analysis tools. Fixed effect hypotheses can be tested via a recently developed test statistic, FH. An accurate approximation of the statistic's small sample non-central distribution makes power computations feasible. After reviewing some longitudinal design terminology and mixed model notation, this paper summarizes the computation of FH and approximate power from its non-central distribution. These methods are applied to obtain a large number of intentionally incomplete full-span designs that are more powerful and/or less costly alternatives to a complete design. The source of the greater efficiency of incomplete designs and potential fragility of incomplete designs to randomly missing data are discussed.

Longitudinal Studies

Interpretation of data on dietary intake.

Although this discussion has focused on the interpretation of dietary data, assuming that it is representative of actual and usual intake and that the nutrient analysis based on it involved the use of up-to-date food composition tables, the readers should be sensitive to other potential sources of error or bias in obtaining information on food and/or nutrient intake. These include errors due to irregularity of food consumption, under- or overreporting of intake, errors in reporting either the amount or the description of the food consumed, recording errors on the part of the interviewer or coding errors on the part of the coder, limitations in the tables of food composition due to missing data for certain nutrients in certain foods or to biologic variability in the same foods from different sources or in those marketed under different conditions, imputed values, the unknown composition of formulated foods or foods prepared from home or commercial recipes, differences in bioavailability of nutrients as a function of the diet, or the use of abridged tables of food composition. In spite of the many unresolved issues relating to dietary standards and the interpretation of dietary intake data, we are still able to make a reasonable assessment of dietary adequacy of groups and individuals with our current system, which is viewed as a unique federal resource. It is hoped that the eventual passage of the National Nutrition Monitoring and Related Research Act will provide both the impetus and the resources to permit us to develop a more sophisticated system for assessing both food intake and nutritional status.

Data Interpretation, Statistical