Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Emergency medical transport of the elderly: a population-based study.

Patterns of utilization of emergency medical services transport (EMS) by the elderly are poorly understood. We determined population-based rates of EMS utilization by the elderly and characterized utilization patterns by age, gender, race, and reason for transport. This observational, population-based study was conducted in Forsyth County, NC, a semi-urban county served by one convalescent ambulance service and one EMS service. Using data on all 1990 EMS transports and the 1990 U.S. census data, age-, gender-, and race-specific transport rates for persons aged 60 or older were calculated. Reasons for transport and frequency of repeat users were established. After exclusion of transports because of an address outside the county, a nonhospital destination, a scheduled transport, or missing data, 4,688 transports (78% of total) remained for analysis. The overall rate of transport was 104/1,000 county residents. Transport rates increased for successively older five-year age groups, demonstrating a 5.7-fold stepwise increase from ages 60-65 to 85+ (51/1,000 to 291/1,000). There was no difference in mean age between patients who were frequent EMS users (more than three transports during the year) (n = 66) and other elderly transportees. Reasons for transport differed little between those 60 to 84 years of age and those 85 years of age and older with the exception of chest pain, cardiac arrest, and seizures, all of which were significantly more prevalent in the younger age group.(ABSTRACT TRUNCATED AT 250 WORDS)

Age Factors↗

Estimating variance components in natural populations using inferred relationships.

Until recently, the estimation of the heritability of a trait has required knowledge of the pedigree within a population. In natural populations such knowledge is often unknown. Two techniques have been developed which use marker information to estimate heritabilities without reference to the exact nature of the relationships: a regression-based estimator that regresses phenotypic similarity for a pair of individuals against an estimate of their relationship and a likelihood-based estimator that maximizes the probability of the genotypic and phenotypic data given a known population structure. Computer simulation was used to compare the behaviour of these estimators. Bias in estimates of heritability decreased with increasing marker information, decreasing simulated heritability, increasing relatedness and increasing sample size. The techniques displayed reasonable tolerance to the percentage of missing data. The regression-based technique shows least average bias, but largest variance over simulations. Likelihood-based techniques show larger average bias, but smaller variances over estimates. A modified form of the likelihood technique, requiring fewer initial assumptions about population parameters, is presented. The modified form shows less bias in its estimates of heritability than the likelihood technique originally proposed.

Analysis of Variance↗

Partially supervised learning using an EM-boosting algorithm.

Training data in a supervised learning problem consist of the class label and its potential predictors for a set of observations. Constructing effective classifiers from training data is the goal of supervised learning. In biomedical sciences and other scientific applications, class labels may be subject to errors. We consider a setting where there are two classes but observations with labels corresponding to one of the classes may in fact be mislabeled. The application concerns the use of protein mass-spectrometry data to discriminate between serum samples from cancer and noncancer patients. The patients in the training set are classified on the basis of tissue biopsy. Although biopsy is 100% specific in the sense that a tissue that shows itself to have malignant cells is certainly cancer, it is less than 100% sensitive. Reference gold standards that are subject to this special type of misclassification due to imperfect diagnosis certainty arise in many fields. We consider the development of a supervised learning algorithm under these conditions and refer to it as partially supervised learning. Boosting is a supervised learning algorithm geared toward high-dimensional predictor data, such as those generated in protein mass-spectrometry. We propose a modification of the boosting algorithm for partially supervised learning. The proposal is to view the true class membership of the samples that are labeled with the error-prone class label as missing data, and apply an algorithm related to the EM algorithm for minimization of a loss function. To assess the usefulness of the proposed method, we artificially mislabeled a subset of samples and applied the original and EM-modified boosting (EM-Boost) algorithms for comparison. Notable improvements in misclassification rates are observed with EM-Boost.

Algorithms↗

Associations between health-related quality of life and demographics and health risks. Results from Rhode Island's 2002 behavioral risk factor survey.

BACKGROUND: Health-Related Quality of Life (HRQOL) has received much attention in recent years. HRQOL indicators have been used to track population trends, identify health disparities, and monitor progress in achieving national health objectives for 2010. Prior studies have examined health risks and HRQOL at the national level as well as at the state level. This paper examines multiple indicators of HRQOL by demographic characteristics and selected health behaviors for Rhode Island adults. METHODS: Data from Rhode Island's 2002 Behavioral Risk Factor Surveillance System (BRFSS), a random digit dialled telephone survey, were used for this study. The state wide sample contained a total of 3,843 respondents ages 18 and older. Multiple Imputation (MI) was applied to handle missing data, and data were modelled for each of 10 HRQOL indicators using multivariable logistic regression. RESULTS: By examining HRQOL through a multivariable approach we identified the strongest predictors for multiple indicators of poor HRQOL as well as predictors for specific indicators of poor HRQOL. Predictors for multiple indicators of poor HRQOL were: disability, inability to work, unemployment, lower income, lack of exercise, asthma, and smoking (specifically associated with poor mental health). CONCLUSION: Using multiple measures of HRQOL can help to assess the burden of poor health in a population, identify subgroups with unmet HRQOL needs, inform the development of targeted interventions, and monitor changes in a population's HRQOL over time. Use of these HRQOL measures in longitudinal and intervention studies is needed to increase our understanding of the causal relationships between demographics, health risk behaviors, and HRQOL.

Adult↗

Dynamic three-dimensional undersampled data reconstruction employing temporal registration.

Dynamic 3D imaging is needed for many applications such as imaging of the heart, joints, and abdomen. For these, the contrast and resolution that magnetic resonance imaging (MRI) offers are desirable. Unfortunately, the long acquisition time of MRI limits its application. Several techniques have been proposed to shorten the scan time by undersampling the k-space. To recover the missing data they make assumptions about the object's motion, restricting it in space, spatial frequency, temporal frequency, or a combination of space and temporal frequency. These assumptions limit the applicability of each technique. In this work we propose a reconstruction technique based on a weaker complementary assumption that restricts the motion in time. The technique exploits the redundancy of information in the object domain by predicting time frames from frames where there is little motion. The proposed method is well suited for several applications, in particular for cardiac imaging, considering that the heart remains relatively still during an important fraction of the cardiac cycle, or joint imaging where the motion can easily be controlled. This paper presents the new technique and the results of applying it to knee and cardiac imaging. The results show that the new technique can effectively reconstruct dynamic images acquired with an undersampling factor of 5. The resulting images suffer from little temporal and spatial blurring, significantly better than a sliding window reconstruction. An important attraction of the technique is that it combines reconstruction and registration, thus providing not only the 3D images but also its motion quantification. The method can be adapted to non-Cartesian k-space trajectories and nonuniform undersampling patterns.

Algorithms↗

Model-based methodology for analyzing incomplete quality-of-life data and integrating them into the Q-TWiST framework.

BACKGROUND: The standard Q-TWiST approach defines a series of health states and weights each state's duration according to its quality of life (QOL) to calculate quality-adjusted lifetimes. However, a fixed weight may not adequately reflect time variations in QOL. METHODS: To account for measurements derived from irregular visits and informative missing data, the authors estimated the mean QOL profile using a mixed-effect growth curve model for the response, combined with a logistic regression model for the drop-out process. RESULTS: Using data from a clinical study of lymphoma patients, the authors demonstrated better readaptation to normal life for patients younger than 30. Sensitivity analyses and computer simulations demonstrated that modeling the drop-out probability as a function of the QOL measurements is necessary if conditioning by health state is not possible. CONCLUSION: Our model-based approach is useful to analyze studies with incomplete QOL data, especially when approximate QOL assessment by health state is not possible.

Adult↗

Familial aggregation in the presence of temporal trends.

Models for assessing temporal trends in familial aggregation are described for both cross-sectional and longitudinal family data. Simultaneous linear structural equations on latent variables are used to model the dependence among family members. The coefficients of the equations are assumed to be parametric functions of time, so that quite complex temporal trends in familial aggregations can be accommodated. Variable family sizes and missing data values pose no problem as the parameters of the models are estimated via maximum likelihood techniques. One of the models is applied to systolic blood pressure data in 542 Japanese-American nuclear families. The results indicate limited evidence for temporal variation in the genetic expression, but that there is substantial temporal variation in environmental influences, which appear to peak at middle age.

Age Factors↗

[Diagnostic agreement between emergency teams and hospital services].

OBJECTIVES: Over the last 10 years the Public Health Emergency Service of Andalusia (Spain) has been conducting a study into the diagnostic agreement among its teams (061 teams) and those of primary care and hospitals. Diagnostic agreement between these teams and hospital teams was evaluated. When discrepancies were found, an assessment was made of whether these corresponded to the emergency team, transfer resources or hospital. PATIENTS AND METHOD: A descriptive study was performed. Five hundred ten patients whose particulars were already known were randomly selected. The patients, who required transfer to a public hospital, received assistance from 061 teams in Malaga in 2001. Data were gathered on personal details, the assistance received, transfer, hospital and diagnosis or diagnoses. The maximum number of diagnoses permitted was three, coded in accordance with the CIE-9 CM classification. The Kappa index was used for comparisons. RESULTS: Ten cases were lost due to missing data. The mean number of diagnoses per patient was 1.48 for 061 teams and was 1.50 in hospital reports. The most common of diagnoses related to injuries and cardiovascular diseases (non-specific diagnoses accounted for approximately 20%). Fifty-nine percent of the patients had at least one diagnosis that coincided. We obtained kappa = 0.478 for a confidence level of 95% (the agreement rate was 73.9%). CONCLUSIONS: Overall agreement was moderate, with better results in the Advanced Coordination Team and conventional ambulance transfer due to the simplicity of the diagnoses. Results classified as "good" were achieved only in the Hospital Costa del Sol, which uses working guidelines similar to those of the Public Health Emergency Service. The percentage of inexact diagnoses was high. Proposals for improvement should range from revising the working methods used to applying new technologies.

Adolescent↗

A data-driven clustering method for time course gene expression data.

Gene expression over time is, biologically, a continuous process and can thus be represented by a continuous function, i.e. a curve. Individual genes often share similar expression patterns (functional forms). However, the shape of each function, the number of such functions, and the genes that share similar functional forms are typically unknown. Here we introduce an approach that allows direct discovery of related patterns of gene expression and their underlying functions (curves) from data without a priori specification of either cluster number or functional form. Smoothing spline clustering (SSC) models natural properties of gene expression over time, taking into account natural differences in gene expression within a cluster of similarly expressed genes, the effects of experimental measurement error, and missing data. Furthermore, SSC provides a visual summary of each cluster's gene expression function and goodness-of-fit by way of a 'mean curve' construct and its associated confidence bands. We apply this method to gene expression data over the life-cycle of Drosophila melanogaster and Caenorhabditis elegans to discover 17 and 16 unique patterns of gene expression in each species, respectively. New and previously described expression patterns in both species are discovered, the majority of which are biologically meaningful and exhibit statistically significant gene function enrichment. Software and source code implementing the algorithm, SSClust, is freely available (http://genemerge.bioteam.net/SSClust.html).

Algorithms↗

Dinoflagellate, Euglenid, or Cercomonad? The ultrastructure and molecular phylogenetic position of Protaspis grandis n. sp.

Protaspis is an enigmatic genus of marine phagotrophic biflagellates that have been tentatively classified with several different groups of eukaryotes, including dinoflagellates, euglenids, and cercomonads. This uncertainty led us to investigate the phylogenetic position of Protaspis grandis n. sp. with ultrastructural and small subunit (SSU) rDNA sequence data. Our results demonstrated that the cells were dorsoventrally flattened, shaped like elongated ovals with parallel lateral sides, 32.5-55.0 mum long and 20.0-35.0 mum wide. Moreover, two heterodynamic flagella emerged through funnels that were positioned subapically, each within a depression and separated by a distinctive protrusion. A complex multilayered wall surrounded the cell. Like dinoflagellates and euglenids, the nucleus contained permanently condensed chromosomes and a large nucleolus throughout the cell cycle. Pseudopodia containing numerous mitochondria with tubular cristae emerged from a ventral furrow through a longitudinal slit that was positioned posterior to the protrusion and flagellar apparatus. Batteries of extrusomes were present within the cytoplasm and had ejection sites through pores in the cell wall. The SSU rDNA phylogeny demonstrated a very close relationship between the benthic P. grandis n. sp. and the planktonic Cryothecomonas longipes. These ultrastructural and molecular phylogenetic data for Protaspis indicated that the current taxonomy of Protaspis and Crythecomonas is in need of re-evaluation. The composition and identity of Protaspis is reviewed and suggestions for future taxonomic changes are presented. Problems within the genus Cryothecomonas are highlighted as well, and the missing data needed to resolve ambiguities between the two genera are clarified.

Animals↗

Protocol for the 'e-Nudge trial': a randomised controlled trial of electronic feedback to reduce the cardiovascular risk of individuals in general practice [ISRCTN64828380].

BACKGROUND: Cardiovascular disease (including coronary heart disease and stroke) is a major cause of death and disability in the United Kingdom, and is to a large extent preventable, by lifestyle modification and drug therapy. The recent standardisation of electronic codes for cardiovascular risk variables through the United Kingdom's new General Practice contract provides an opportunity for the application of risk algorithms to identify high risk individuals. This randomised controlled trial will test the benefits of an automated system of alert messages and practice searches to identify those at highest risk of cardiovascular disease in primary care databases. DESIGN: Patients over 50 years old in practice databases will be randomised to the intervention group that will receive the alert messages and searches, and a control group who will continue to receive usual care. In addition to those at high estimated risk, potentially high risk patients will be identified who have insufficient data to allow a risk estimate to be made. Further groups identified will be those with possible undiagnosed diabetes, based either on elevated past recorded blood glucose measurements, or an absence of recent blood glucose measurement in those with established cardiovascular disease. OUTCOME MEASURES: The intervention will be applied for two years, and outcome data will be collected for a further year. The primary outcome measure will be the annual rate of cardiovascular events in the intervention and control arms of the study. Secondary measures include the proportion of patients at high estimated cardiovascular risk, the proportion of patients with missing data for a risk estimate, and the proportion with undefined diabetes status at the end of the trial.

Journal Article↗

Suggestive genome-wide associations with inflammatory biomarkers in an admixed population, including a missense variant in the OR6K6 olfactory receptor gene associated with MCP-1.

BACKGROUND: Chronic low-grade inflammation drives cardiometabolic diseases and has a strong genetic basis. Most genome-wide association studies (GWAS) have focused on European populations, limiting knowledge of the genetic influences on inflammation in admixed populations such as those in Brazil. METHODS: This study is part of the cross-sectional ISA Capital Health Survey. It uses data from the 2015 ISA Nutrition cohort, which measured biochemical, genetic, anthropometric, and lifestyle factors in a probabilistic sample of São Paulo residents. Genomic DNA was extracted from 841 individuals. Genotyping was performed using the Axiom 2.0 Precision Medicine Research Array. After quality control and missing data exclusion, 244,338 SNPs from 638 individuals remained for GWAS-based association analysis with eight inflammatory biomarkers. Models were adjusted for sex, age, age2, overweight, and the first two principal components of ancestry. RESULTS: Most participants were male (53%) and not overweight (55%). The median age was 49, and 38% were older adults. In the genome-wide analysis of TNF-α, IL-10, IL-1β, monocyte chemoattractant protein-1 (MCP-1), and adiponectin, 12 SNPs were significantly associated, most of which were intronic. Notably, one signal mapped to the missense variant rs16841009 in the olfactory receptor gene OR6K6. This variant was associated with MCP-1, suggesting a possible involvement in inflammatory responses. CONCLUSIONS: We identified new SNPs linked to inflammatory biomarkers in a highly admixed Brazilian population, including a missense variant in an olfactory receptor gene linked to MCP-1. This association may be biologically important for inflammation and could affect the risk of cardiometabolic diseases.

Humans↗

C3: A comprehensive physician activity and billing tool

Purpose: The Clinical Charge Capture system (C3) was developed at the University of Michigan to increase the efficiency and accuracy with which information about physician activity and billing is tracked in academic medical centers. Description: This Oracle-based, Visual Basic system integrates the operating room scheduling system, transcription database, clinical data repository, referring physician database, and IDX to allow physicians and staff to perform paperless and on-line standard tasks such as preauthorizing procedures; creating a bill which describes the charges for procedures performed along with their supporting diagnoses; identifying inpatient daily care and consult charges; dictating, editing, signing, and providing attestations for procedural and inpatient notes (menu-driven boilerplate notes are used for common procedures); submitting of charges on-line to IDX; and downloading of payment data from IDX. A messaging system between physicians and billing specialists allows questions to be posed regarding coding issues and options. Summary information about charges is presented and the status of the bill as it progresses through the internal review and billing process is demonstrated. Any missing data are flagged such that delivery of a bill is accurate, timely, and complete. Outpatient clinic visit charges are acquired on line using bar code technology with direct download of clinic charges to IDX. Generation of charges and referral letters may be performed immediately following the performance of a procedure or patient encounter or subsequently in the office. Resident activity is also tracked. Finally, search functions are provided which allow the program to serve as a clinical information research database. Results: The time to bill submission for operative procedures in fiscal year 1996 (Pre-C3) when compared to 1999 (Post-C3) decreased in each individual surgical division (See figure)as well as for the overall Department (Total: mean Pre-C3=40 days, mean Post-C3=8 days). The average bill was increased by 9% for each primary charge submitted. Conclusions: We conclude that this system has the potential to enhance the efficiency, accuracy, and organization of routine physician documentation, billing, and data collection activities.

Journal Article↗

Detection of disease genes by use of family data. I. Likelihood-based theory.

We present a class of likelihood-based score statistics that accommodate genotypes of both unrelated individuals and families, thereby combining the advantages of case-control and family-based designs. The likelihood extends the one proposed by Schaid and colleagues (Schaid and Sommer 1993, 1994; Schaid 1996; Schaid and Li 1997) to arbitrary family structures with arbitrary patterns of missing data and to dense sets of multiple markers. The score statistic comprises two component test statistics. The first component statistic, the nonfounder statistic, evaluates disequilibrium in the transmission of marker alleles from parents to offspring. This statistic, when applied to nuclear families, generalizes the transmission/disequilibrium test to arbitrary numbers of affected and unaffected siblings, with or without typed parents. The second component statistic, the founder statistic, compares observed or inferred marker genotypes in the family founders with those of controls or those of some reference population. The founder statistic generalizes the statistics commonly used for case-control data. The strengths of the approach include both the ability to assess, by comparison of nonfounder and founder statistics, the potential bias resulting from population stratification and the ability to accommodate arbitrary family structures, thus eliminating the need for many different ad hoc tests. A limitation of the approach is the potential power loss and/or bias resulting from inappropriate assumptions on the distribution of founder genotypes. The systematic likelihood-based framework provided here should be useful in the evaluation of both the relative merits of case-control and various family-based designs and the relative merits of different tests applied to the same design. It should also be useful for genotype-disease association studies done with the use of a dense set of multiple markers.

Alleles↗

Coordinating outcomes measurement in ataxia research: do some widely used generic rating scales tick the boxes?

The objective of this study was to examine the psychometric properties of four widely used generic health status measures in Friedreich's ataxia (FA), to determine their suitability as outcome measures. Fifty-six people with genetically confirmed FA completed the Barthel Index (BI), General Health Questionnaire (GHQ-12), EuroQol (EQ-5D), and Medical Outcomes Study 36-item Short Form Health Survey (SF-36) by means of postal survey. Six psychometric properties (data quality, scaling assumptions, acceptability, reliability, validity, and responsiveness) were examined. The response rate was 97%. In general, the psychometric properties of the four measures satisfied recommended criteria. However, closer examination highlighted limitations restricting their use for treatment trials. For example, the BI had high levels of missing data, EQ-5D had poor discriminant ability, and five SF-36 scales had high floor and/or ceiling effects. Most scale scores did not span the entire scale range, had means that differed notably from the scale mid-point, and had wide confidence intervals. Effect sizes (ES) were small for all four measures raising questions about their ability to detect clinically significant change. Results highlight the potential limitations of these four scales for evaluating health outcomes in FA and suggest the need for new disease-specific patient-based measures of its impact.

Activities of Daily Living↗

Modeling antitumor activity by using a non-linear mixed-effects model.

The response of solid tumors to antitumor treatment generally declines markedly with treatment time. Sometimes, a tumor regrows (rebounds) before the end of the treatment period. Studies of the patterns of tumor response to treatment are important, because they may provide useful information for clinical decision-making. We have investigated patterns of tumor response in mouse xenograft tumors by using data from a study conducted at St. Jude Children's Research Hospital. We applied a biexponential non-linear mixed-effects model to an analysis of changes in tumor volume over a given period of treatment. The model gives a good fit to the data, even for small sample sizes. We addressed the relation between the baseline tumor volumes and the decay rates of the first and second stages of the tumor's response to treatment, and we applied sensitive analysis to determine the effect of using different imputed values for missing data. We also proposed a novel approach to a comparison of the antitumor effects of three different treatments, and we used the data from a St. Jude study to demonstrate the potential of this comparison approach in cancer clinical decision-making.

Algorithms↗

An indicator of adverse pregnancy outcome in France: not receiving maternity benefits.

STUDY OBJECTIVE: The aim was to compare the social characteristics, the pregnancy outcome, and the antenatal care of women in France who did not receive maternity benefits to women who did. These benefits (860 FF, approx 86 pounds per month) are given to every pregnant woman, starting in the second trimester. Payments are made on the condition that at least three antenatal visits are made, the first being before the end of the first trimester. DESIGN: The study involved a random sample of women who were interviewed after delivery during their stay in hospital. Data on pregnancy outcome were collected from medical records. SETTING: The study was carried out in four public maternity units in different regions of France. PARTICIPANTS: 1692 women were included in the analysis (86.8% of the selected sample). Of 257 exclusions, 40 had multiple pregnancies, 189 had missing data, and 28 did not answer the question concerning maternity benefits. MEASUREMENTS AND MAIN RESULTS: 4.3% of the women did not receive any maternity benefits. These women lived in poorer social conditions than the women who received the benefits. They had a higher preterm delivery rate, after controlling for risk factors in a logistic regression. Women without maternity benefits were characterised by a lower level of care, yet the majority began their antenatal care during the first trimester or had more than six visits. CONCLUSIONS: Not receiving maternity benefits during pregnancy is an index of an underprivileged situation and a risk factor for pregnancy outcome.

Age Factors↗

The ethics of gene therapy and abortion: public opinion.

On the grounds that the public should be consulted in decisions concerning the legitimate scope of germ-line genetic therapy (GLGT), survey data on the ethics of GLGT were collected from a large (n = 1,403) representative national sample of Australians in 2002. The data show that opinion is quite divided over GLGT in the case of a 'death sentence' genetic defect: 36% would forbid it, 23% have mixed feelings and 41% would allow it. For less serious conditions there is more opposition to GLGT. Thus, 48% would forbid GLGT to remedy a minor physical defect and 52% would oppose GLGT to counteract a propensity to violence, but fully 73% would disallow GLGT for cosmetic reasons. The data also show that opposition to abortion is lower than opposition to GLGT in the case of a 'death sentence' genetic defect, but at about the same level as, or greater than, opposition to GLGT for less serious issues. The questions show good measurement properties, including low missing data rates, so they are likely to provide an accurate picture of the public's views on the ethics of GLGT. It is suggested that a system for monitoring public opinion on these issues be developed.

Abortion, Induced↗