Missing data: nurses with their patients.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The statistical analysis of longitudinal quality of life data in the presence of missing data is discussed. In cancer trials missing data are generated due to the fact that patients die, drop out, or are censored. These missing data are problematic in the monitoring of the quality of life during the trial. However, by means of assuming that the cause of the missing data lies in the observed history of the patients and not in their unobserved future, the missing data are ignorable. Consequently, all available data can be used to estimate quality of life change patterns with time. The computations that are required are illustrated with real quality of life data and three commonly used computer packages for statistical analysis.
We discuss the analysis of growth curve data with missing or incomplete information. The approach is to fit subject-specific models and then to carry out an analysis in terms of the estimated parameters. This achieves reduction of data and eliminates the need for special considerations for subjects with missing data. Although there is no perfect substitute for complete data, our approach provides a way to handle missing data using a straightforward application of well-known statistical methodology.
The effects of missing values for a confounding variable are investigated in the setting of case-control studies in which, for simplicity, the effect of one binary risk factor and one categoric confounding variable on disease risk is under investigation. Some ad hoc techniques with which to deal with missing values are examined under different assumptions about the missing-data mechanism. Examples are given to illustrate that the magnitude of the bias that is introduced by applying an inadequate procedure can be large under circumstances that occur frequently in empiric research. This is true even for so-called complete case analysis, i.e., when only data on subjects with complete information are used. Appropriate bias corrections are derived. Making use of data on those subjects who are neglected in complete case analysis by creating an additional category always results in biased estimation. An alternative is to allocate these subjects to the cells of the contingency table in an appropriate manner. This approach yields consistent estimates if the data are missing at random. Choosing an appropriate method for dealing with missing values always requires some knowledge of why the data are missing. This suggests that investigators should carry out validation studies to understand whether the missing values occur randomly across the study population or occur more frequently in specific subgroups.
Recent methodological development in phylogenetic inference has focused predominantly on molecular data. However, renewed interest in other data types, particularly morphological data, has followed from the increased recognition of the power of total evidence and tip-dating approaches, including fossil data, for inference of time-scaled trees and rates of evolution. However, attention has largely focused on the improvement of models of morphological evolution and other analytical tools with much less discussion about data acquisition itself. Here we review past and current practice for describing and collecting morphological data for phylogenetic inference. We present a systematic review of 164 phylogenetic analyses conducted over the last 35 years and focused on a diverse group of extinct arthropods: trilobites. Trends in increasing matrix size, data type, and coding strategy are evident. Where present, polymorphic characters have been predominantly derived from discretized continuous characters, although increasingly practitioners are utilizing alternative approaches for the treatment of quantitative characters. Not surprisingly, traditional indices that describe character consistency are highly correlated with matrix size but show surprising variation at different taxonomic scales. More recent attempts to describe data quality using information theory imply that characters can have high information content even if data are missing for many tips, providing support against the exclusion of characters because of missing data. In consideration of this, as well as advances in the study of developmental biology and variational complexity, we identify several avenues for increasing the quality and quantity of morphological data going forward.
This paper describes the problem of informative censoring in longitudinal studies where the primary outcome is rate of change in a continuous variable. Standard approaches based on the linear random effects model are valid only when the data are missing in a non-ignorable fashion. Informative censoring, which is a special type of non-ignorably missing data, occurs when the probability of early termination is related to an individual subject's true rate of change. When present, informative censoring causes bias in standard likelihood-based analyses, as well as in weighted averages of individual least-squares slopes. This paper reviews several methods proposed by others for analysis of informatively censored longitudinal data, and outlines a new approach based on a log-normal survival model. Maximum likelihood estimates may be obtained via the EM algorithm. Advantages of this approach are that it allows general unbalanced data caused by staggered entry and unequally-timed visits, it utilizes all available data, including data from patients with only a single measurement, and it provides a unified method for estimating all model parameters. Issues related to study design when informative censoring may occur are also discussed.
Autoregressive time series model-based spectral estimates of heart period sequences can provide a parsimonious and visually attractive representation of the dynamics of interbeat intervals. While a corollary to Wold's decomposition theorem implies that the discrete Fourier periodogram spectral estimate and the autoregressive spectral estimate converge asymptotically, there are practical differences between the two approaches when applied to short blocks of data. Autoregressive spectra can achieve good frequency resolution and excellent statistical stability on short segments of heart period data of sinus origin. However, the order of the autoregressive model (number of free parameters to be estimated) must be explicitly chosen, a decision that influences the trade-off of frequency resolution with statistical stability. Akaike's Information Criterion (AIC), an information-theoretic rule for picking the optimum order, is sensitive to the aggregate amount of data in the analysis. Thus, the best model order for estimating the spectrum of a 4-minute segment of data will generally be lower than the best order for estimating an hourly spectrum based on averaging 15 4-minute spectra. A major advantage of the autoregressive model approach to spectral analysis is the ease with which it can be extended to handle messy data frequently seen in heart rate variability studies. A number of autoregressive-based robust-resistant techniques are available for the analysis of heart period sequences that contain a high volume of nonsinus and other unusual beats intervals. A theoretically satisfying framework is also available for spectral analysis of unevenly sampled data and missing data.
Missing data in sample surveys is virtually unavoidable, whether it is an entire unit that is missing or only an item for a responding unit. Compensation for unit nonresponse is usually made through the assignments of weights to responding units; for item nonresponse, the compensation often is by an imputation procedure. This paper reviews the extent of missing data in a large federal survey, the National Medical Care Utilization and Expenditure Survey, and the imputation procedures used to compensate for item missing data. The effects of imputation on several types of estimates from the survey are examined. In addition, several methods for analyzing survey data with imputed values are reviewed, and recommendations about preferred strategies are made for selected circumstances.
The ID3 algorithm for inductive learning was tested using preclassified material for patients suspected to have a thyroid illness. Classification followed a rule-based expert system for the diagnosis of thyroid function. Thus, the knowledge to be learned was limited to the rules existing in the knowledge base of that expert system. The learning capability of the ID3 algorithm was tested with an unselected learning material (with some inherent missing data) and with a selected learning material (no missing data). The selected learning material was a subgroup which formed a part of the unselected learning material. When the number of learning cases was increased, the accuracy of the program improved. When the learning material was large enough, an increase in the learning material did not improve the results further. A better learning result was achieved with the selected learning material not including missing data as compared to unselected learning material. With this material we demonstrate a weakness in the ID3 algorithm: it can not find available information from good example cases if we add poor examples to the data.
Analysts must deal frequently with missing data in multivariate analysis. In such cases, estimating the covariance maxtrix V of the dependent variables usually involves initial estimation and iterative adjustment of imputed missing data values, and/or smoothing of an estimate V which is not necessarily positive semi-definite. This paper presents an alternative procedure for computing estimates of relevant multivariate parameters in situations where missing data occur at random and with small probability. MISCAT is a computer program which computes multivariate ratio estimates of the means and a corresponding positive semi-definite estimate of the covariance matrix. It is an extension of GENCAT, which is a program for the generalizaed least squares analysis of categorical data. Thus, one advantage of dealing with missing data in this manner is that variation among the ratio estimates may be conveniently analyzed within MISCAT using asymptotic regression methodology, provided that sample sizes are sufficiently large. An example is given to illustrate such analysis for longitudinal data from a multicenter clinical trial.
A new multivariate statistical quality control method has been developed. It is an extension of the method developed by Kume, which is able to find abnormal values in multivariate biochemical data of a clinical laboratory. The present method makes use of the difference between two sets of data measured from the samples of the same patient obtained on different days. The Mahalanobis' distance between two samples can be calculated from the difference of their observations. If the Mahalanobis' distance of the two data is larger than the critical value decided in advance, the reliability of the measurement is doubtful. The characteristic of the present method is that it can apply to data with missing values by estimating them from measured data. Some numerical examples are shown to demonstrate the availability of the method.
The purpose of this report is to document the procedures used in the 1988 National Survey of Family Growth (NSFG) to select the sample, weight the data to produce national estimates, impute missing data, and estimate sampling errors. Therefore, this report necessarily contains a great deal of technical detail. For readers who do not need this level of detail, this summary briefly describes the procedures used. The National Survey of Family Growth is conducted every few years by the National Center for Health Statistics (NCHS), a part of the U.S. Department of Health and Human Services. The purpose of the survey is to collect and publish data from a national sample of women on childbearing, factors affecting childbearing (such as contraception, sterilization, and infertility), and related aspects of maternal and infant health. Interviewing for Cycle IV of the survey was done in 1988 by Westat, Inc., under a contract with NCHS. Personal interviews were conducted between January and August of 1988 with a national sample of 8,450 women in the civilian noninstitutionalized population of the United States. Interviews were conducted in person by trained female interviewers and lasted an average of 70 minutes. The interview focused on the woman's pregnancies, if any; her use of contraception; her ability to bear children (fecundity and infertility); her use of medical services for family planning, infertility, and prenatal care; her marriage and cohabitation history, if any; and a wide range of demographic and economic characteristics. This report describes some of the main methodological aspects of the survey, including the sample design, weighting, sampling errors, and imputation of missing data. These topics will be described briefly and less technically in this summary. Each topic is discussed in more detail in the rest of the report.
The impact of the clinical database system SISCOPE on medical services was evaluated and objective data compiled on the quality of information recording and reporting using a fully structured data entry system compared to traditional free text reporting. 1565 upper endoscopy reports produced with SISCOPE over a period of 12 months were assessed for completeness and compared to 152 and 208 free text reports done 4 months before and 1 month after the study period, respectively. Data on four common gastrointestinal findings (esophageal varices, ulcers, polyps and tumors) were evaluated. Physicians' compliance with the new system was good, as reflected by a constant level of quality of reporting over time, although a very slight decline in the ratio of computer generated reports to the total number of examinations was noted. Structured reports had an 18% missing data rate and contained 60% more relevant information than free text reports, which had a 48% missing data rate. No educational effect of the system was seen as missing data rates returned to pre-computerization levels just one month after the end of the study. It is concluded that menu-driven structured data entry systems result in production of far superior reports as compared to free text systems, probably due to their reminder effect.
Physiologists often wish to compare the effects of several different treatments on a continuous variable of interest, which requires an analysis of variance. Analysis of variance, as presented in most statistics texts, generally requires that there be no missing data and often that each sample group be the same size. Unfortunately, this requirement is rarely satisfied, and investigators are confronted with the problem of how to analyze data that do not strictly fit the traditional analysis of variance paradigm. One can avoid these pitfalls by recasting the analysis of variance as a multiple linear regression problem. When there are no missing data, the results of a traditional analysis of variance and the corresponding multiple regression problem are identical; when the sample sizes are unequal or there are missing data, one can use a regression formulation to analyze data that cannot be easily handled in a traditional analysis of variance paradigm and thus overcome a practical computational limitation of traditional analysis of variance. In addition to overcoming practical limitations of traditional analysis of variance, the multiple linear regression approach is more efficient because in one run of a statistics routine, not only is the analysis of variance done but also one obtains estimates of the size of the treatment effects (as opposed to just an indication of whether such effects are present or not), and many of the pairwise multiple comparisons are done (they are equivalent to t tests for significance of the regression parameter estimates). Finally, interaction between the different treatment factors is easier to interpret than it is in traditional analysis of variance.
BACKGROUND: Improvements in neonatal and paediatric care in recent decades have increased the survival of children with non-progressive neurological impairment. Respiratory disease in children with neurological impairment is common, with symptoms difficult to manage and lower respiratory tract infection occurring frequently. To reduce these, prophylactic antibiotics are being increasingly used, but the type, duration and dose of antibiotics can vary considerably, and there is limited evidence about their effectiveness in children and young people. A joint United Kingdom and Australia multicentre, randomised, double-blind, placebo-controlled trial comparing 52 weeks of azithromycin to placebo in children and young people with neurological impairment at risk of lower respiratory tract infection (PARROT) was planned to address this gap. PARROT was a multicentre, parallel group, blinded, pragmatic randomised controlled trial of 52-week duration with a planned sample size of 500 (250 in each arm) participants with neurological impairment. The primary outcome was the proportion of children and young people hospitalised with lower respiratory tract infection over the 52-week period. RESULTS: In total, 90 children and young people (62 in Australia, 28 in the United Kingdom) aged 3-17 years, with a diagnosed non-progressive, non-neuromuscular neurological impairment, who had persistent respiratory symptoms were randomised (1 : 1) to receive azithromycin or placebo. Baseline demographic and clinical characteristics were relatively well balanced across the two treatment groups and countries. Overall, mean (standard deviation) age was 9.2 (4.4) years, with 64% of participants having cerebral palsy, 67% being non-ambulant and 54% being totally tube-fed. At baseline, mean (standard deviation) numbers of hospital admissions with lower respiratory tract infection in the preceding year were 1.8 (2.0)/year, and general practitioner attendances 3.3 (3.0)/year. The PARROT trial was closed early to recruitment due to challenges arising from the COVID-19 pandemic. Sixty-five (72%) participants (azithromycin n = 30, placebo n = 35) completed 52 weeks of treatment and were not withdrawn early from the trial. Regarding the primary outcome, 11 (36.7%) in the azithromycin group were hospitalised with lower respiratory tract infection and 9 (25.7%) in the placebo group [absolute risk reduction 0.11 (95% confidence interval -0.12 to 0.33), relative risk 1.43 (95% confidence interval 0.68 to 2.97)]. Analysis of secondary outcome data was limited by the number of missing data, but parent-reported quality of life for young person and parent, sleep amount/quality for young person and parent, and respiratory symptoms were similar between groups and countries. LIMITATIONS: As PARROT was stopped early and was consequently underpowered, it is not possible to say whether azithromycin prophylaxis is any more effective than placebo in reducing the proportion of children admitted to hospital with lower respiratory tract infection after a 52-week period. CONCLUSIONS AND FUTURE WORK: Although we cannot comment on the effectiveness of prophylactic antibiotics in this context, we can draw some useful conclusions from this trial. Thus, the importance placed by families on hospitalisation and its prevalence in both treatment groups, even during the pandemic, would suggest that this is an appropriate primary outcome measure for future trials in this high-risk group of children and young people. Furthermore, the high attrition rate and large numbers of missing data, specifically for questionnaire-based outcomes at later follow-up points, should encourage researchers to be mindful of minimising trial burden to families for any future trials wherever possible. FUNDING: This synopsis presents independent research funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme as award number 16/17/01.
A wide variety in outcome criteria hinders comparison of results between smoking cessation studies. Three important methodological issues are discussed: analysis of data of participants who drop out of therapy, treatment of missing data, and repeated use of significance tests. These issues determine to a great extent the results of evaluation studies. In general, they are of interest to all researchers of addiction who study the effects of interventions. Several ways to decide on these issues and the consequences of these decisions are considered. Little consensus exists about the criterion for dropout. It is concluded that a dropout criterion is a burden rather than a help. A better criterion would be the number of sessions present. Few satisfying techniques exist to handle the problem of missing data. Evaluation studies need to set a priori standards to counter the increased risk of a Type I error, caused by the repeated use of significance tests. Reviewers need to be aware of the variety in data treatment before comparing results.
The range of tilt angles for which projected images of two-dimensionally periodic specimens can be obtained in electron microscopy is limited both by technical aspects, such as goniometer design, and by the more fundamental limitation of object thickness. The lack of a full set of projections causes a missing cone in the reciprocal space data for the object, which will give an anisotropic resolution in a three-dimensional reconstruction and may cause the quality to be impaired by spurious features. The problem is governed by a linear operator which maps the three-dimensional object onto the set of projections. The eigenvalue spectrum of this operator is determined by the range of tilt angles and the spatial extent of the object. If the object is spatially restricted, the eigenvalues are all positive, and it is in principle possible to retrieve experimentally unavailable structure data from those that are measured. However, with restricted angle data, some of the eigenvalues are extremely small, so the problem is 'ill-conditioned' or sensitive to small perturbations in the data, such as noise, and it is necessary to regularize the solution. We applied two methods of band-limited extrapolation and inference on electron microscope data. Alternating projections onto convex sets regularized by a regularization parameter and a least squares estimation regularized by the Shannon entropy functional yield similar results if a close object extent constraint is available. The criterion of maximum entropy, however, allows a relaxation of this constraint.
A positive association between compliance and clinical outcome has been observed in several randomized, controlled, clinical trials. This association, seen in the placebo-treated group as well as the active-treatment group, clarifies the possibility that data analyses incorporating estimates of protocol adherence are potentially biased. In the presence of non-compliance, or missing data from any cause, several statistical analyses may seem plausible, with none clearly superior to the others. These may include an analysis of all patients randomized, with imputed values for missing data, and an analysis restricted to protocol-adherent patients. The recommended approach is a conservative one that examines consistency among the plausible analyses. Using compliance data in trial conduct can also introduce bias into trial results by inducing differential treatment of compliers and non-compliers. This possibility arises, for instance, when adherence is affected by the randomized treatment. Non-compliance can have a substantial impact on statistical power and sample size requirements in a clinical trial. Under certain assumptions, required sample sizes are doubled with 30% non-compliance and tripled with 40% non-compliance.