Search PubMedSearch

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Poor agreement of occupational data between a hospital-based cancer registry and interview.

With occupation recognized as a risk factor for various cancers, collecting occupation and industry data by a number of vital registries, including cancer registries, has developed. Registries may be data sources for cancer etiology research and occupational disease surveillance, despite concerns that their data are fragmentary and may lack validity. To improve completeness and validity of occupational information in a hospital-based cancer registry, this study compared information obtained through abstracting medical records for the registry with information obtained through lung-cancer patient interviews. Employing the kappa statistic, agreement was generally poor, largely due to data missing in the medical record. Data quality of hospital-based cancer registries can be improved by employing trained cancer registrars to elicit occupational histories from patients.

Data Collection

Time-series analysis--cosinor analysis: a special case.

Cosinor analysis provides an accessible means of evaluating and estimating the parameter of a cyclic phenomenon. Cosinor analysis does not require that the data be equal intervals without missing data. Cosinor analysis does require that the data can reasonably be considered to take the form of a deterministic cycle with a known period.

Humans

Validation of malaria surveillance case reports: implications for studies of malaria risk.

STUDY OBJECTIVE: The aim of the study was to investigate the quality of national malaria surveillance reports in the United Kingdom. DESIGN: Persons with malaria reported to the Malaria Reference Laboratory (MRL) in 1987 were contacted by post to verify existing records with respect to key variables. The MRL data set was then analysed for inaccuracies. SETTING: The study was confined to UK residents. PARTICIPANTS: 602 persons with malaria in 1987 responded (53%). MEASUREMENTS AND MAIN RESULTS: Review of case reports showed few missing data except for duration of residence in the UK, detailed chemoprophylactic regimens, and compliance. There were more missing surveillance data in reports of ethnic minority groups, principally in dates of travel (p = 0.008) and chemoprophylaxis use (p less than 0.0001). Patient recall in the survey was at variance with the surveillance reports in dates of travel and onset of infection, chemoprophylaxis use, and in compliance. Surveillance reports overestimated the number of days between leaving a malarious area and onset of symptoms (by 9 d for P falciparum and by 24 d for P vivax), and underestimated the delay between onset and diagnosis of P falciparum by 3 d. Over 50% of patients who had recalled the use of chloroquine, proguanil, pyrimethamine/dapsone, and pyrimethamine had not been recorded as having taken these drugs on the surveillance reports. Reported compliance also differed between the two data sets. CONCLUSIONS: It is recommended that research units test the quality of their surveillance data before embarking on analytical studies used to generate health policy guidelines.

Adult

Repeated measures designs in behavioral toxicology: application to chronic marijuana smoke exposure.

This paper discusses the application of repeated measures methods in the statistical analysis of an experiment in behavioral toxicology. The chronic marijuana smoke exposure study conducted at the National Center for Toxicological Research is used for an example of the types of problems that one encounters in analyzing these types of studies. In particular, the standard univariate analysis most frequently used for repeated measures analyses has some very restrictive assumptions on the form of the covariance matrices. These assumptions are not met in the example discussed and are rarely met in many other problems. Other possible models for analyzing repeated measures when these assumptions are not met are presented and discussed. Other problems specific to the chronic marijuana smoke exposure study that may occur in similar type studies are presented. These include pooling the experimental units into groups with comparable baselines, choosing a function of the measures to be analyzed, dealing with a large data set with many observation times and missing data, unequal group sizes and different designs for different subsets of the experimental animals. The standard univariate repeated measures analysis was chosen to analyze the data even though the violations of the covariance assumptions may lead to finding differences that do not exist (Type I or false-positive errors), since the other methods presented also had covariance assumptions that were not met or had low power. Use of Bonferroni-type multiple comparisons on the single degree of freedom contrasts of interest hopefully reduced the chances of these false-positive results.

Analysis of Variance

The impact of changing methods of data collection on the reliability of self-reported drug use of adolescents.

The purpose of this study is to determine the impact of different modes of data collection on the reliability of self-reported drug use of adolescents in a panel study. Adolescents were assigned to four groups based upon the ways they chose to respond to the survey instruments: 1) mailed questionnaires in both years, 2) survey interview in one year and mailed questionnaire in the next year, 3) mailed questionnaire in one year and survey interview in the following year, and 4) survey interview in both years. The quality of the self-reported data was examined in terms of return rates, missing data, internal consistency, and consistency of reported information over time. No significant differences were found between groups, suggesting that the mode of data collection does not affect the reliability of adolescents' self-reports of substance use.

Adolescent

Applications of multiple imputation to the analysis of censored regression data.

The first part of the article reviews the Data Augmentation algorithm and presents two approximations to the Data Augmentation algorithm for the analysis of missing-data problems: the Poor Man's Data Augmentation algorithm and the Asymptotic Data Augmentation algorithm. These two algorithms are then implemented in the context of censored regression data to obtain semiparametric methodology. The performances of the censored regression algorithms are examined in a simulation study. It is found, up to the precision of the study, that the bias of both the Poor Man's and Asymptotic Data Augmentation estimators, as well as the Buckley-James estimator, does not appear to differ from zero. However, with regard to mean squared error, over a wide range of settings examined in this simulation study, the two Data Augmentation estimators have a smaller mean squared error than does the Buckley-James estimator. In addition, associated with the two Data Augmentation estimators is a natural device for estimating the standard error of the estimated regression parameters. It is shown how this device can be used to estimate the standard error of either Data Augmentation estimate of any parameter (e.g., the correlation coefficient) associated with the model. In the simulation study, the estimated standard error of the Asymptotic Data Augmentation estimate of the regression parameter is found to be congruent with the Monte Carlo standard deviation of the corresponding parameter estimate. The algorithms are illustrated using the updated Stanford heart transplant data set.

Algorithms

Ensuring data quality in a multicenter clinical trial: remote site data entry, central coordination and feedback.

In an ongoing multicenter clinical trial, "Treatment Strategies in Schizophrenia," the five participating sites have the capacity to perform a variety of tasks or study functions independently. These tasks include (a) verification of diagnostic eligibility through the use of computerized decision algorithms; (b) assignment of patients to treatment based on prognostic indicators using a computerized randomization algorithm; (c) entry of data into a microcomputer using a clinical trial data management system that performs simple range and missing data item checks; and (d) regular transfer of all data to the central coordinating team. The clinical trial data management system employed allows for both independent site functioning and assurance of consistency across sites. The integration of a variety of software outside the main data management system provides the central coordinators with the tools to monitor critical data as it is collected, as well as the capacity to assess the flow, quality, and uniformity of the ongoing trial.

Clinical Trials as Topic

Nicardipine and propranolol in the treatment of essential hypertension.

Two hundred thirty-four patients with supine diastolic blood pressure of between 95 and 114 mm Hg were enrolled into a double-blind, randomized, parallel, multicenter trial. The patients were randomized to either nicardipine 30 mg tid, propranolol 40 mg tid, or nicardipine 30 mg tid and propranolol 40 mg tid for six weeks. Two hundred six patients yielded data for analyses. Of the 28 not included, seven had missing data, whereas the remaining 21 were excluded because they either failed to meet inclusion criteria or were noncompliant at endpoint. Both nicardipine and propranolol as monotherapies and in combination achieved statistically significant, (P less than .01), supine diastolic blood pressure reduction relative to baseline. The combination of nicardipine and propranolol showed a greater reduction in supine diastolic and systolic measurements than either of the monotherapies. Nicardipine produced greater blood pressure reductions one hour after dosing, whereas the propranolol treatment tended to produce slightly greater blood pressure decreases eight hours after dose. The combination always resulted in the greatest blood pressure reduction, independent of time after dose. Adverse experiences were reported by 26% of patients in the nicardipine-treated group, most often transient vasodilatory effects, by 17% of the propranolol-treated patients, and by 18% of the combination-treated group. This study demonstrated at the doses studied that nicardipine alone produced equivalent blood pressure reductions to those obtained by propranolol alone, but that the combination of these two drugs produced greater reductions in blood pressures than either of the monotherapies.

Adult

Microcomputer application of Bayesean probability testing for the identification of bacteria.

A computer program (BACTID) is described which facilitates the identification of bacteria based on a priori data and Bayesean probability testing. The program is not limited to a specific format, has a short execution time, can be easily applied to a variety of situations, and can be run on almost any microcomputer system operating under either 8-bit CP/M or 16-bit MS-DOS/PC-DOS. Additionally, BACTID (1) is not limited to one type of computer (hardware independent), (2) is not limited by size of the computer's random access (RAM independent), (3) can recognize various data bases matrices (format independent), (4) is able to compensate for missing data and (5) allows for various methods of data entry. The efficacy of the program was checked against a commercially available test system and a 99.34% agreement was obtained. Also, the execution time for a 46 x 21 element data matrix was as little as 3.5 s. These results show that microcomputer identification programs are not only viable alternatives to code book registers, but also offer flexibility which is not found in commercial systems.

Bacteria

A comparison of response rate, data quality, and cost in the collection of data on sexual history and personal behaviors. Mail survey approaches and in-person interview.

The authors examined differences in rate of response, data quality, and cost between mail approaches and in-person interview in the collection of data on sexual history and personal behaviors. A sample of women from a midwestern United States university (n = 342) was identified from health service medical records as having been seen for a sexually transmitted disease (cases) or a contraceptive visit (controls) during the latter half of 1985. The women were randomly assigned to one of three data collection strategies. A total of 268 subjects (78%) participated. Results indicated no differences in validity by method of data collection or by case-control status but there were significant differences in completeness, cost, and response rates. In-person interviews resulted in more complete data than mail approaches, although all instruments had low proportions of missing data (0.001-0.006). Response rate differences were not found when data collection methodologies were compared (75-82%) but were found in case-control analyses. Cases were consistently less likely to participate and significantly less likely to respond by mail (p less than 0.05). The cost of the in-person interview was approximately four times that of the mail survey for the data collection. Implications of the case-control response rate difference suggest that mail methodologies, although low in cost, may introduce sampling bias in studies of sexually transmitted diseases.

Adolescent

A method for assessing patterns of familial resemblance in complex human pedigrees, with an application to the nevus-count data in Utah kindreds.

An analytic method is described for estimating phenotypic correlations between pairs of members of specific relationships in pedigrees. In estimating correlations, this new method allows simultaneous adjustment for available covariates such as age, gender, environmental factors, and variables reflecting ascertainment mode, through mean- and variance-regression models. The estimated correlations and regression coefficients corresponding to covariates are consistent and asymptotically normally distributed. Differing from a full-likelihood approach, this new method does not require the assumption of a particular joint distribution of phenotypes from a pedigree, such as the multivariate normal distribution, but instead only requires correct specification of mean- and variance-regression models. Within this framework, missing data, if they are missing completely at random, can be ignored without biasing estimates. The method is illustrated by an application using nevus-count data from 28 Utah kinships. The results from the analysis are that covariate-adjusted nevus counts are correlated between parents and children (correlation .22; P less than .001) and between siblings (correlation .32; P less than .001), while the correlation of -.04 between husband and wife is not significantly different (P = .31) from 0. This result is consistent with a genetic etiology of nevus count.

Dysplastic Nevus Syndrome

Closed-form estimates for missing counts in two-way contingency tables.

One method for analyzing contingency tables with missing observations is to model the missing-data mechanism using log-linear models. Previous methods for obtaining estimates (of missing counts and parameters) have required an iterative algorithm. In many cases, however, one can obtain estimates by use of a simple algebraic formula. We illustrate the method with data on smoking and birth weight.

Algorithms

The use of lung function tests in identifying factors that affect lung growth and aging.

Lung function tests are used both clinically, in assessing disease, and epidemiologically, in identifying those factors which influence the growth and aging process of the lungs. The user must beware of several common pitfalls in the use of these tests, however. First, the commonly used tests of lung function can only identify patterns of dysfunction, not specific pathologic processes. Second, these tests are subject to many sources of intra- and inter-subject variability, making it difficult to dissect out the signal (for example, the rate of lung aging in adults) from the noise which may greatly exceed the signal. Finally, the analysis of longitudinal pulmonary function data is complicated by the alinearity of the growth and aging process, missing data, variable follow-up times, information censoring mechanisms and covariate processes, and problems in defining abnormality.

Adolescent

A longitudinal study of the reporting of emotional and somatic symptoms during and after pregnancy.

One hundred and eight pregnant women, most of their husbands and a comparison group of non-expectant parents were recruited for a long-term study which involved responding to a 55-item Symptom Checklist (SCL) and the Beck Depression Inventory three times during pregnancy and once during the first postpartum month. Responses to the SCL were factor analysed, and the four groups were then compared on their factor scores as well as their scores on the Beck Depression Inventory (BDI) using discriminant analysis and trend analysis. The discriminant analyses were done twice: once using all the data provided by all subjects and once using only subjects with no missing data. At each measurement period, the pregnant women were distinguished from the other groups by a different factor of the SCL: at 3-5 months, it was 'Feeling Sick'; at 6-8 months, it was 'Feeling Overweight'; at 9 months, it was 'Feeling Overweight/Physical Stress'; and at postpartum, it was 'Physical Stress'. Also, trend analysis showed a significant tendency for the scores of pregnant women on the SCL 'Negative Emotional State', factor and on the BDI to increase over time, in contrast to those of the other groups.

Adaptation, Psychological

The use of a computer in the organization of a large-scale co-operative controlled clinical trial of the treatment of pulmonary tuberculosis.

The organization of controlled clinical trials requires well-defined clerical procedures which, with large-scale studies, are so monotonous that it is difficult to attract and keep staff of a sufficiently high calibre. In a large-scale trial of pulmonary tuberculosis, a computer is being used to undertake much of the routine clerical work, including (1) the preparation of an appointments diary, which also specifies the exact requirements for the trial, and (2) the storage of all the data on magnetic tape so that periodic checks can be made for missing data and interim and final analyses can be produced rapidly. These arrangements reduce the time spent by clerical and statistical staff to a minimum. Although a strict evaluation of the effect of introducing the computer has not yet been made, this approach does appear to be sufficiently promising to warrant further investigations of a similar type.

Antitubercular Agents

Cytogenetic effects of inhaled ozone in man.

Peripheral blood samples were collected from 30 normal male volunteers before and at intervals after inhaling 0.4 ppm ozone for 4 h. Data from 4 of the subjects were excluded from the analysis because of missing data points. The blood samples were cultured for 48 h, slides made and stained with a uniform Giemsa stain, and 100 metaphase spreads per subject per treatment scored for chromosome aberrations. Cells with suspected aberrations were photographed, destained, restained with a banding procedure and rephotographed to identify the specific chromosomes and regions involved. Pre-exposure, immediate post-exposure, 3 days post-exposure, 2 weeks post-exposure and 4 weeks post-exposure means for the percentage of cells with 46 chromosomes were 93.0, 93.6, 91.7, 94.5 and 94.2, respectively; in the same order, the mean number of cells with chromatid and/or chromosome breaks per order, the mean number of cells with chromatid and/or chromosome breaks per 100 cells was 0.96, 0.85, 1.00, 0.88 and 0.81 respectively, and for chromatid and/or chromosome gaps per 100 cells: 1.35, 0.96, 1.35, 0.81 and 0.77, respectively. The means for each of these parameters as well as the mean frequencies of complex aberrations are not statistically significantly different between blood sampling times. The distribution of aberrations by chromosome and light and dark bands is not significantly influenced by ozone exposure. These data indicate no apparent detectable human cytogenetic effect due to exposure to ozone under the conditions of this experiment.

Adult

Creation and validation of medical data.

The problems of creating and validating medical data are discussed. These difficulties relate to the structuring of data and its definiton. The necessity for clarity of thought, coherence and logic in data collection is stressed. Difficulty in medical and lay staff training shows that this is a continuing problem. Data validation must be the major province of senior as well as junior doctors. It is an absoulte essential continuing process and without valid data most analyses are not worthwhile because of high error rates and/or too much missing data.

Humans

Reconstruction algorithm for incomplete projections in the framework of linear operators in normed linear spaces.

Based on the linearity of the Radon transform and the convolution-backprojection reconstruction algorithm, a new linear-vector space notation is introduced that is of general use in computed tomography (CT). Using this notation, a consistency condition for the completion of incomplete projection data is described. This consistency condition leads to singular or ill-conditioned systems of linear equations for the unknown projection data. Using regularization methods, an algorithm for the consistent projection completion is presented that can exploit symmetries of the missing data region. The performance of the new algorithm is documented with simulated and actualy measured CT-projection data. The algorithm quantitatively improves CT reconstructions with realistic amounts of data and noise and can be used for the completion of arbitrary regions of missing projections.

Fourier Analysis