Search PubMedSearch

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking

Protecting against nonrandomly missing data in longitudinal studies.

Nonrandomly missing data can pose serious problems in longitudinal studies. We generally have little knowledge about how missingness is related to the data values, and longitudinal studies are often far from complete. Two approaches that have been used to handle missing data--use of maximum likelihood with an ignorable mechanism and direct modeling of the missing data mechanism--have the disadvantage of not giving consistent estimates under important classes of nonrandom mechanisms. We introduce two protective estimators, that is, estimators that retain their consistency over a wide range of nonrandom mechanisms. We compare these protective estimators using longitudinal data from a mental health panel study. We also investigate their robustness to certain departures from normality.

Data Interpretation, Statistical

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans

A model-based approach to the imputation of missing data: home injury incidences.

Missing or incomplete data cases are a problem in all types of statistical analyses. In disease surveillance, this problem inhibits determining the actual incidence of a disease event and monitoring the disease occurrence. Several statistical techniques have been developed to impute values for incomplete data cases. We present a model-based approach to the imputation of missing data elements as applied to determining the incidence of home injury deaths.

Accidents, Home

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources.

MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Humans

An application of multivariate ratio methods for the analysis of a longitudinal clinical trial with missing data.

This paper presents an analysis of a longitudinal multi-center clinical trial with missing data. It illustrates the application, the appropriateness, and the limitations of a straightforward ratio estimation procedure for dealing with multivariate situations in which missing data occur at random and with small probability. The parameter estimates are computed via matrix operators such as those used for the generalized least squares analysis of catetorical data. Thus, the estimates may be conveniently analyzed by asymptotic regression methods within the same computer program which computes the estimates, provided that the sample size is sufficiently computer program which computes the estimates, provided that the sample size is sufficiently large.

Clinical Trials as Topic

The reporting and handling of missing data in genetic epidemiological studies of mental health in childhood and adolescence: A systematic review.

BACKGROUND: Genetic epidemiological analyses of child and adolescent mental health often use data from prospective longitudinal cohorts. Missingness due to selective attrition is therefore an important potential source of bias in such analyses. Informatively reporting on missingness and taking appropriate steps to handle it in analyses can mitigate this potential bias. Here, we aim to systematically assess how researchers report and address missingness in genetic epidemiological studies of child and adolescent mental health-related outcomes using cohort data. METHODS: We systematically searched the Ovid Medline database for studies published between August 2012 and August 2025, reporting polygenic score, genome-wide association, or Mendelian randomization analyses, of data on children or adolescents participating in cohort studies. We extracted information from eligible studies based on criteria adapted from the strengthening and reporting of observational studies in epidemiology (STROBE) guidelines. RESULTS: A total of 133 eligible studies were included, of which 125 (93.98%) reported the number of complete cases in all waves, while 84 (63.16%) detailed the amount of missingness on all key variables. Most studies used complete case analysis, while 39 studies explicitly reported applying other methods to handle missingness, with multiple imputation (n = 20, 15.04%) being the most common, followed by full information maximum likelihood 10 (8.1%). Only 18 studies (13.53%) reported an assumed missing mechanism along with the method used to address missingness. Full reporting of both the extent and handling of missingness at the item level was rare, occurring in only 5 (3.76%) and 15 (11.28%) studies, respectively, among the 123 studies that used multi-item instruments. CONCLUSION: Best practice recommendations for reporting on missing data handling emphasize the importance of detailing the proportion of missingness, types of mechanisms underpinning missingness, and details of approaches used. Based on this review, these recommendations for proper reporting of missing data are rarely followed in full.

children and adolescents

Missing data in longitudinal studies.

When observations are made repeatedly over time on the same experimental units, unbalanced patterns of observations are a common occurrence. This complication makes standard analyses more difficult or inappropriate to implement, means loss of efficiency, and may introduce bias into the results as well. Some possible approaches to dealing with missing data include complete case analyses, univariate analyses with adjustments for variance estimates, two-step analyses, and likelihood based approaches. Likelihood approaches can be further categorized as to whether or not an explicit model is introduced for the non-response mechanism. This paper will review the use of likelihood based analyses for longitudinal data with missing responses, both from the point of view of ease of implementation and appropriateness in view of the non-response mechanism. Models for both measured and dichotomous outcome data will be discussed. The appropriateness of some non-likelihood based analyses is briefly considered.

Algorithms

Hierarchical time-oriented approaches to missing data inference.

In practice clinical data are nearly always incomplete. When confronted with such data, a physician or investigator must make inferences about missing information. Possible strategies for inference include (1) interpolation, (2) extrapolation, (3) repeating the nearest value, (4) repeating the previous value, (5) patient-specific mean values, (6) patient-specific linear regression over time, (7) disease-specific mean values, (8) normal values, and (9) linear regression of correlated co-recorded variables. This study analyzes these strategies in a time-oriented data bank of patients with systemic lupus erythematosus, demonstrating that more accurate inferences of missing data are obtained when (1) strategies are tailored to the characteristics of the individual variable, (2) time-oriented strategies (e.g., interpolation) rather than non-time-oriented strategies (e.g., disease mean) are incorporated, (3) a ranked set of strategies is incorporated in a hierarchical stepwise fashion, and (4) the degree to which missing data are "nonrandomly" missing is assessed to allow estimation of bias. Interpolation is the best single technique with these data while linear regression of correlated co-recorded variables is a relatively weak technique. Inferences made by these hierarchical time-oriented approaches show significantly smaller mean differences from the actual values than do results from typical statistical package strategies.

Data Interpretation, Statistical

Missing data in two-way analysis of variance.

We previously described a regression approach to analysis of variance computations that permitted analysis of unbalanced designs and experiments with missing data in two-way (or higher) analyses of variance Am. J. Physiol. 255 (Regulatory Integrative Comp. Physiol. 24): R353-R367, 1988. That approach can only be used correctly under a set of narrow, and relatively uninteresting, circumstances. In fact, in the example we worked, we extended that approach beyond its intended scope and incorrectly computed F statistics for testing hypotheses about the main effects of strain of rat and nephron site in a study of renal Na(+)-K(+)-adenosinetriphosphatase. This paper presents the correct approach, which can be generalized to most situations likely to be encountered when two-way, or higher, analyses of variance are used.

Adenosine Triphosphatases