Search PubMedSearch

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Reconstruction algorithm for incomplete projections in the framework of linear operators in normed linear spaces.

Based on the linearity of the Radon transform and the convolution-backprojection reconstruction algorithm, a new linear-vector space notation is introduced that is of general use in computed tomography (CT). Using this notation, a consistency condition for the completion of incomplete projection data is described. This consistency condition leads to singular or ill-conditioned systems of linear equations for the unknown projection data. Using regularization methods, an algorithm for the consistent projection completion is presented that can exploit symmetries of the missing data region. The performance of the new algorithm is documented with simulated and actualy measured CT-projection data. The algorithm quantitatively improves CT reconstructions with realistic amounts of data and noise and can be used for the completion of arbitrary regions of missing projections.

Fourier Analysis

Descriptive study of prognostic factors influencing survival of compensated silicotic patients.

The objective of this study was to assess survival and the prognostic influencing survival of compensated silicotic patients. All workers compensated for silicosis in the Province of Quebec from 1938 to 1985 (n = 1,165) were included. Clinical data were those collected during the exam that led to a compensation decision. Due to missing data, a subcohort of 961 patients was used for multivariate analysis of clinical prognostic factors with the Cox proportional hazards model. The following factors made an independent contribution to survival: age at compensation, smoking, dyspnea, expectoration, abnormal breath sounds, radiographic appearance, and vital capacity. On the basis of the model, patients with small opacities alone on their chest radiograph and who did not have dyspnea, expectoration, or abnormal breath sounds had a survival similar to the average Quebec man; other patients had a poorer survival. We conclude that it is possible to identify at the time of compensation, silicotic patients who are likely to have a life expectancy similar to that of the general population. Symptoms and physical signs as well as radiographic and lung function abnormalities appear to be useful prognostic indicators in compensated silicotic patients.

Dyspnea

Gestational age reporting and preterm delivery.

This study examines recent trends in the reporting completeness and quality of gestational age estimates derived from the date of the last normal menses (DLNM) as reported in South Carolina vital records from 1974 to 1985. Noteworthy improvements in the completeness of reporting emerged during this period with a decline from 31.1 percent missing information in 1974 to 6.6 percent missing in 1985. Completeness of reporting and strategies for imputing values for missing data were analyzed for their impact on the calculation of the percentage of preterm live births. The results indicate that the underreporting of gestational age can lead to marked underestimation of the preterm percentage in a population and to misinterpretation of trends in these percentages. Based on the results of this analysis, it is recommended that preterm percentages be based on cases with DLNM gestational age values between 20 and 50 weeks. Since cases with missing or implausible gestational age data have a greater risk of a poor pregnancy outcome, these findings emphasize the importance of identifying both the completeness of data reporting and the use of imputation and deletion strategies when employing population-based DLNM data to calculate gestational age related indicators.

Adolescent

Investigating drug plasma levels and clinical response using random regression models.

The reanalysis of the Riesby dataset using a random regression model indicates a significant effect of DMI plasma measurements on HAM-D scores across the four timepoints of the study. This effect is especially strong when the HAM-D change from baseline score is used in place of the actual HAM-D score at the four timepoints. A significant effect of IMI was not found for either the actual HAM-D score or the HAM-D change from baseline score. An endogenous effect was marginally significant when the actual HAM-D score was used; however, this effect was not observed when the HAM-D change score was used as the dependent measure. There also was evidence of a marginally significant effect due to autocorrelation of the residuals, indicating that the residuals at a given timepoint were related to the residuals from previous timepoints according to a first-order autoregressive process. In contrast, when the repeated measures MANOVA was used to analyze these data, a significant effect of DMI was not observed, since subjects without complete data at all timepoints had to be dropped from the analysis. In longitudinal psychiatric studies where missing data are the rule (rather than the exception), the random regression approach provides an attractive alternative to the traditional methods of analyzing longitudinal data.

Humans

Data omitted from psychiatric consultation notes.

To assess how often psychiatric consultants omit written data from their consultation notes, the authors reviewed 78 initial consultation notes written by second-year psychiatric residents. Data considered essential for an adequate psychiatric evaluation were typically omitted. Categories that were observed to have the highest frequencies of missing data included family history of psychiatric illness (60.3%), history of substance abuse (44.9%), marital status (37.2%), previous psychotropic drug use (35.9%), previous psychiatric treatment (26.9%), and patient history of psychiatric illness (24.4%). The frequencies of omissions were significantly (p less than .001, except for the last item, p less than .01) higher than those from the consultation notes written by a second cohort of psychiatric residents who used a worksheet that listed data categories. The authors' findings argue for the use of worksheets delineating data categories to ensure that clinicians write adequate consultation notes.

Data Collection

A new computer program for detailed off-line analysis of swimming navigation in the Morris water maze.

The program TRACK-ANALYZER runs on AT-compatible microcomputers and performs off-line analysis of data recorded from Morris water-maze experiments by means of a video tracking system or digitizing tablet. Raw data must be available on disk as ASCII-files listing position coordinates sampled at a constant frequency. Automatic recognition and correction of artifacts and missing data is a key feature of the program, as well as the option to combine commands to user-defined macros. TRACK-ANALYZER offers maximal flexibility regarding experimental schedule, maze geometry and recording parameters and may also be used to analyze open-field activity. In addition to calculating basic parameters of the swim path, such as path length and time, the program counts crossings and hits of goal platform and four virtual reference annuli, calculates search times in five different maze fields, determines directionality, tortuosity and turning preferences of swimming behavior and allows the viewing of any number of trials simultaneously on screen. ASCII-formated data output may easily be exported to commercial statistics and graphics software.

Animals

The nominative technique: a new method of estimating heroin prevalence.

Over the years, nominative estimates of heroin prevalence have been consistently higher than self-reports of heroin use. During this time, nominative data have generally followed mainstream patterns of drug use: nominative estimates for young adults and for males are higher than nominative estimates for older persons, youth, and females; moreover, the recent downward trends in drug use have been replicated by the nominative heroin data. Thus, the overall picture presented by the nominative data--similar patterns but higher levels of prevalence--seems to support the validity of the new approach. Nevertheless, considerable caution should be exercised in interpreting nominative data. This is chiefly because a substantial minority of nominators cannot report the number of other close friends of the heroin user who also "know." While missing data has been handled by a conservative imputation rule, the fact that so many persons are unable to provide an answer to this key question casts doubt on the accuracy of the answers that were given. In fact, the nominative approach might tend to produce over-estimates, because of the potential for undercounts of the numbers of others who "know." Additional tests of validity should be performed, such as application of the nominative approach to nonsensitive behaviors or minimally sensitive behaviors, such as marijuana use or perhaps cocaine use. Certainly, the overall validity of the nominative heroin data would be supported if in future surveys new nominative heroin estimates for relatively unstigmatized forms of drug use proved to be similar to self-reported levels of use, thus pointing to the unique difference in estimates that might be observed for heroin. Finally, in interpreting the heroin estimates presented here, it should be remembered that both the nominative and self-report estimates refer to heroin use in the household population of the United States. Thus, many heroin addicts and other users who reside in various unconventional living arrangements would not be included in the counts presented here. Among the excluded groups are transients residing in rooming houses or "crashing" in the home of one "friend" after another or who are incarcerated in jails or confined to residential drug treatment centers. This is a caution for interpreting the estimates presented in this paper, not a criticism of the nominative technique itself.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

Reliability of the CES-D Scale in different ethnic contexts.

Reliability of the Center for Epidemiologic Studies Depression Scale, a 20-item symptom checklist, is examined using data from a sample of community respondents containing Anglos (254), Blacks (270), and Mexican Americans (181). Although the survey response rate was lower for Mexican Americans, quality of the data provided by this group was not significantly different from that for Anglos or Blacks. That is, there were no differences among these groups in terms of missing data or internal consistency reliabilty (as measured by Cronbach's alpha and Spearman-Brown split halves). Factor-analytic results also demonstrate the same general structure of responses among the three groups.

Black or African American

Data attrition in follow-up studies of alcoholics.

Alcoholics who responded to posttreatment follow-up evaluation showed more improvement than those wh did not respond. Thus nonresponders may bias follow-up research tht does not account for missing data. Several procedures are outlined to increase follow-up return.

Adult

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning

RECPAM: a computer program for recursive partition amalgamation for censored survival data and other situations frequently occurring in biostatistics. II. Applications to data on small cell carcinoma of the lung (SCCL).

The RECPAM methodology previously presented in part I (A. Ciampi et al., Comput. Methods Programs Biomed. 26 (1988) 239-256) is applied to the analysis of survival data on small cell carcinoma of the lung (SCCL). It is shown how RECPAM can help answer the following questions which occur frequently in the analysis of clinical data: Is it possible to find a classification of patients with a certain disease into distinct prognostic groups? Given a covariate of special interest, does it have an independent prognostic significance even after confounding is taken into account? Does the prognostic significance of a covariate of special interest vary across patient subgroups? For the SCCL data, a prognostic classification is obtained and the tumor marker LDH is treated as a variable of special interest. Many features of RECPAM are illustrated, including, among others, Forward and Backward (Pruning) Stopping Rules, treatment of missing data, and use of several dissimilarity measures.

Biomarkers, Tumor

Repeated-measures model for the investigation of temporal trends using longitudinal family studies: application to systolic blood pressure.

A contemporary path model for the analysis of familial resemblance is extended to incorporate repeated measurements on the entire pedigree over time, in order to assess age-related changes in familiality. The parameters of the model can be defined as arbitrary functions of the ages, age differences, or cohabitation times of the family members at the exact time of measurement. Tracking of the phenotypes is decomposed into a familial and a nonfamilial component, which varies with both the time span between measurements and the ages at measurement. Some of the family members may have data missing on one or more visits, and the visits may be unequally spaced both within and across families. The method incorporates all measurements available from all visits into a single model. The model is applied to longitudinal data on systolic blood pressure in 490 East Boston families measured two times at 3-year intervals. Evidence for some nonfamilial tracking is found. Additionally, significant temporal trends are demonstrated in the familiality as a function of age, t2(A), which appears to be near zero at birth, grow to a maximum of about 40% at around age 30, and then appears to monotonically decrease again. No evidence was found for temporal trends in marital resemblance or residual sibling environmental effects. This model provides an objective method of investigating developmental changes in the correlational structure of families over time using repeated-measures and of estimating continuous changes in familiality with age.

Age Factors

Comparison of questionnaire and diary methods in acute childhood respiratory illness surveillance.

We compared two prospective survey methods, an interviewer-administered questionnaire and a daily diary, used concurrently to record acute respiratory illness experience over a 2-yr period in 422 children 5 to 11 yr of age from East Boston, Massachusetts. Respondents contributed more months of data with the questionnaire than with the diary method. Respiratory symptom and illness rates, as determined for the first year by each of the methods, were compared for 277 children who had less than 4 months of missing data. Respondents from families with more children tended to report a lower total respiratory illness rate by the diary than by the questionnaire method (p = 0.006). Although upper respiratory illness rates did not differ by method, lower respiratory illnesses were reported more frequently (p = 0.0001) by questionnaire than by diary. In the group of 49 children who were identified as having had greater than one lower respiratory illness, 25% of the illnesses reported as having been lower respiratory by questionnaire were reported as having been another form of respiratory illness by diary. For this group the ratio of 3:1 of boys to girls for the diary as compared with 1.5:1 for the questionnaire suggests the presence of reporting bias and no comparability of methods. Standardization of an acute respiratory illness questionnaire would provide greater opportunity than use of diaries for synthesis of prospective data from different epidemiologic studies.

Acute Disease

Problems in setting up an executing large-scale psychiatric epidemiological studies.

This paper focuses on problems that can be encountered in conceptualizing, executing and writing up large-scale psychiatric epidemiological studies. It makes no attempt to cover fundamental issues of design and analysis, rather it centers on problems associated with projects of considerable size. In the conceptual area, it discusses the prerequisites to be considered before deciding to launch such a study. It notes the administrative and scientific uses of epidemiological studies and considers the strengths and weaknesses of large-scale studies to address those concerns. Issues in carrying out such studies are discussed including decisions about study design, sampling method and instrumentation. All are dependent on the central purpose of the study but trade-offs between feasibility and scientific rigor are always present. Data collection and analysis problems highlighted in large-scale studies are examined. They include the difficulty, in the former, of adequately motivating and supervising field personnel and, in the latter, of dealing with problems that accompany missing data and complicated sampling strategies. Potential problems in data access and use and writing up the results are seen as arising from the presence of a large investigative team with diverse interests. Lastly, the comparative worth of these studies is considered.

Data Collection

A longitudinal analysis of factors related to survival in old age.

Data from a longitudinal study of the elderly in rural North Wales are used in an exploratory study of the relationships between very broadly defined social circumstances and longevity. A statistical modeling approach is adopted and has some nonroutine features necessitated by missing data on dates of death. A variety of demographic, socioeconomic, social network, quality-of-life, dependence, and health variables are found, individually, to be related to survival. Multivariate analysis demonstrated that many of these relationships are spurious and, in particular, there is no prima facie evidence that survival is affected by social networks or quality-of-life factors. However, socioeconomic factors emerge as important for the old elderly.

Aged

Temporal trends of human immunodeficiency virus type 1 (HIV-1) infection among inmates entering a statewide prison system, 1985-1987.

Acquired immune deficiency syndrome (AIDS) became the leading cause of death among Maryland State prisoners in 1985. To identify the prevalence, risk factors, and temporal trends for infection with the human immunodeficiency virus type 1 (HIV-1) in the statewide prison system, excess sera were obtained from incoming male inmates during specified periods between April and June 1985, 1986, and 1987. Correctional medical personnel also provided demographic variables of age, race, offense category, sentence, jurisidiction, and an indicator of intravenous drug use. Once rendered anonymous, specimens were assayed for antibody to HIV-1, using ELISA and Western blot techniques. For data from April to June 1985, 1986, and 1987, the crude prevalence of anti-HIV-1 was 7.1, 7.7, and 7.0%, respectively. Although one-third of incoming inmates were identified as intravenous drug users (IVDUs), the drug use variable was missing for 70% of the 1985 sample, and 40% of the 1986 sample. Several strategies were used to examined temporal trends in the context of missing data. Univariate analyses suggested no substantial change over time for either HIV-1 seroprevalence or risk of infection among IVDUs.(ABSTRACT TRUNCATED AT 250 WORDS)

Acquired Immunodeficiency Syndrome

Familial aggregation in the presence of temporal trends.

Models for assessing temporal trends in familial aggregation are described for both cross-sectional and longitudinal family data. Simultaneous linear structural equations on latent variables are used to model the dependence among family members. The coefficients of the equations are assumed to be parametric functions of time, so that quite complex temporal trends in familial aggregations can be accommodated. Variable family sizes and missing data values pose no problem as the parameters of the models are estimated via maximum likelihood techniques. One of the models is applied to systolic blood pressure data in 542 Japanese-American nuclear families. The results indicate limited evidence for temporal variation in the genetic expression, but that there is substantial temporal variation in environmental influences, which appear to peak at middle age.

Age Factors

Suggestive genome-wide associations with inflammatory biomarkers in an admixed population, including a missense variant in the OR6K6 olfactory receptor gene associated with MCP-1.

BACKGROUND: Chronic low-grade inflammation drives cardiometabolic diseases and has a strong genetic basis. Most genome-wide association studies (GWAS) have focused on European populations, limiting knowledge of the genetic influences on inflammation in admixed populations such as those in Brazil. METHODS: This study is part of the cross-sectional ISA Capital Health Survey. It uses data from the 2015 ISA Nutrition cohort, which measured biochemical, genetic, anthropometric, and lifestyle factors in a probabilistic sample of São Paulo residents. Genomic DNA was extracted from 841 individuals. Genotyping was performed using the Axiom 2.0 Precision Medicine Research Array. After quality control and missing data exclusion, 244,338 SNPs from 638 individuals remained for GWAS-based association analysis with eight inflammatory biomarkers. Models were adjusted for sex, age, age2, overweight, and the first two principal components of ancestry. RESULTS: Most participants were male (53%) and not overweight (55%). The median age was 49, and 38% were older adults. In the genome-wide analysis of TNF-α, IL-10, IL-1β, monocyte chemoattractant protein-1 (MCP-1), and adiponectin, 12 SNPs were significantly associated, most of which were intronic. Notably, one signal mapped to the missense variant rs16841009 in the olfactory receptor gene OR6K6. This variant was associated with MCP-1, suggesting a possible involvement in inflammatory responses. CONCLUSIONS: We identified new SNPs linked to inflammatory biomarkers in a highly admixed Brazilian population, including a missense variant in an olfactory receptor gene linked to MCP-1. This association may be biologically important for inflammation and could affect the risk of cardiometabolic diseases.

Humans