Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Sources”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Pediatric farm injuries involving non-working children injured by a farm work hazard: five priorities for primary prevention.

OBJECTIVES: To describe pediatric farm injuries experienced by children who were not engaged in farm work, but were injured by a farm work hazard and to identify priorities for primary prevention. DESIGN: Secondary analysis of data from a novel evaluation of an injury control resource using a retrospective case series. DATA SOURCES: Fatal, hospitalized, and restricted activity farm injuries from Canada and the United States. SUBJECTS: Three hundred and seventy known non-work childhood injuries from a larger case series of 934 injury events covering the full spectrum of pediatric farm injuries. METHODS: Recurrent injury patterns were described by child demographics, external cause of injury, and associated child activities. Factors contributing to pediatric farm injury were described. New priorities for primary prevention were identified. RESULTS: The children involved were mainly resident members of farm families and 233/370 (63.0%) of the children were under the age of 7 years. Leading mechanisms of injury varied by data source but included: bystander and passenger runovers (fatalities); drowning (fatalities); machinery entanglements (hospitalizations); falls from heights (hospitalizations); and animal trauma (hospitalizations, restricted activity injuries). Common activities leading to injury included playing in the worksite (all data sources); being a bystander to or extra rider on farm machinery (all data sources); recreational horseback riding (restricted activity injuries). Five priorities for prevention programs are proposed. CONCLUSIONS: Substantial proportions of pediatric farm injuries are experienced by children who are not engaged in farm work. These injuries occur because farm children are often exposed to an occupational worksite with known hazards. Study findings could lead to more refined and focused pediatric farm injury prevention initiatives.

Adolescent↗

Terms used by nurses to describe patient problems: can SNOMED III represent nursing concepts in the patient record?

OBJECTIVE: To analyze the terms used by nurses in a variety of data sources and to test the feasibility of using SNOMED III to represent nursing terms. DESIGN: Prospective research design with manual matching of terms to the SNOMED III vocabulary. MEASUREMENTS: The terms used by nurses to describe patient problems during 485 episodes of care for 201 patients hospitalized for Pneumocystis carinii pneumonia were identified. Problems from four data sources (nurse interview, intershift report, nursing care plan, and nurse progress note/flowsheet) were classified based on the substantive area of the problem and on the terminology used to describe the problem. A test subset of the 25 most frequently used terms from the two written data sources (nursing care plan and nurse progress note/flowsheet) were manually matched to SNOMED III terms to test the feasibility of using that existing vocabulary to represent nursing terms. RESULTS: Nurses most frequently described patient problems as signs/symptoms in the verbal nurse interview and intershift report. In the written data sources, problems were recorded as North American Nursing Diagnosis Association (NANDA) terms and signs/symptoms with similar frequencies. Of the nursing terms in the test subset, 69% were represented using one or more SNOMED III terms.

Delivery of Health Care↗

Agreement between medical record data and patients' accounts of their medical history and treatment for dyspepsia.

We examined agreement between data abstracted from medical records and interview data for patients with dyspepsia admitted to hospital for endoscopy, to determine the extent to which health records could be used to validate self-reports of dyspepsia and the management of this condition. Results from the sample of 220 patients showed that there was poor agreement between data sources for information about duration of dyspepsia (k=0.34) and previous barium meal examination (k=0.34). Patients reported significantly longer dyspepsia histories (Wilcoxon sign test Z=4.13, p<0.0001) and significantly more barium meals (sign test Z=8.43, p<0.0001) than were documented in their records. There was also disagreement between data sources regarding the number of drugs taken before and after endoscopy (k=0.28 and k=0.31, respectively). Where there was disagreement for number of drugs there was no significant difference in the direction of the disagreement. There was moderate agreement regarding the name of pre-endoscopy medication (k=0.55) and substantial agreement for the name of medication used post-endoscopy (k=0.62). There was very poor agreement regarding diagnosis. The medical record was the gold standard for this information. Choice of data source, medical records or self-reports, will in many instances provide significantly different results and it is likely that this may also be true for other variables of interest to researchers. Thus in the case where no gold standards are available researchers need to consider carefully the implication of choice of data source on their results.

Dyspepsia↗

How good are US studies of HIV costs of care?

BACKGROUND: Valid, timely estimates of the costs of HIV care are needed by health planners and policy makers. OBJECTIVE: To perform a methodologic critique of published estimates of resource utilization and costs of HIV care. DATA SOURCES: MEDLINE database for 1990-1998. DATA SELECTION: Included articles focused on adults with a spectrum of HIV disease in which the authors developed their own resource use and cost data. Thirty one articles met these criteria. DATA EXTRACTION: Studies were compared based on: (1) utilization and cost estimates, in 1995 dollars; (2) study period; (3) research design; (4) sampling frame; (5) sample size and patient characteristics; (6) data sources and scope of services; and (7) methods used in the analysis. DATA SYNTHESIS: The most recent estimates pertain to the first half of 1995, before the use of protease inhibitor therapy. We found wide variations in the estimates and identified three major sources for this: (1) patient samples that were restricted to subgroups of the national HIV-infected population; (2) utilization data that were limited in scope (e.g., inpatient care only); and (3) invalid methods for estimating annual or lifetime costs, particularly in dealing with decedents. CONCLUSIONS: To accurately estimate resource use and costs for HIV care nationwide, a nationally representative probability sample of HIV-infected patients is required. Even in research that is not intended to provide national estimates, the scope of utilization data should be broadened and greater attention to methodologic issues in the analysis of annual and lifetime costs is needed.

Acquired Immunodeficiency Syndrome↗

Cost-of-illness of scleroderma: the case for rare diseases.

OBJECTIVE: To determine the societal costs of scleroderma (SSc), a rare chronic connective tissue disease that affects approximately 98,000 Americans. Lack of reliable national databases limit rare disease cost studies, and this study suggests methods of using multiple data sources to assess the costs of rare diseases. METHODS: Primary and secondary data sources were used to calculate direct and indirect costs of SSc, including discounted lifetime mortality and morbidity costs. A prevalence-based, human capital approach was used. Sensitivity analyses were used to vary parameters that are uncertain, such as prevalence, mortality, and labor costs. RESULTS: Annual direct and indirect costs of SSc in the United States are $1.5 billion. Morbidity represents the major cost burden, with costs of $819 million (56%) of total costs. The current value of lifetime earnings lost was $179 million (12%) or $300,000 per death. Direct costs were $462 million (32%) or $4,731 per person annually, indicating that costs are spread over the long disease duration. CONCLUSIONS: This study provides one model for the assessment of rare disease costs. Triangulation of data sources and sensitivity analyses are important for determining the costs of rare diseases. The high cost of SSc, despite its low prevalence, suggests that the burden of rare chronic diseases can be high. The high morbidity costs reflect the young age of onset of the disease as well as the need for treatments to decrease morbidity costs. Local shared databases and national surveys are needed to improve cost estimates of rare diseases.

Adolescent↗

Combining gene expression profiles and protein-protein interaction data to infer gene functions.

The ever-increasing flow of gene expression profiles and protein-protein interactions has catalyzed many computational approaches for inference of gene functions. Despite all the efforts, there is still room for improvement, for the information enriched in each biological data source has not been exploited to its fullness. A composite method is proposed for classifying unannotated genes based on expression data and protein-protein interaction (PPI) data, which extracts information from both data sources in novel ways. With the noise nature of expression data taken into consideration, importance is attached to the consensus expression patterns of gene classes instead of the actual expression profiles of individual genes, thus characterizing the composite method with enhanced robustness against microarray data variation. With regard to the PPI network, the traditional clear-cut binary attitude towards inter- and intra-functional interactions is abandoned, whereas a more objective perspective into the PPI network structure is formed through incorporating the varied function-function interaction probabilities into the algorithm. The composite method was implemented in two numerical experiments, where its improvement over single-data-source based methods was observed and the superiority of the novel data handling operations was discussed.

Algorithms↗

The large increase in incidence of Type I diabetes mellitus in Poland.

AIMS/HYPOTHESIS: A rising incidence of Type I (insulin-dependent) diabetes mellitus in different countries in Europe during the last decade has been recently reported. However, in the early 1990s, Poland was reported to have a stable low incidence of this disease. This study aimed to estimate the annual incidence of Type I diabetes in a north-eastern region of Poland (Białystok region) and investigate if it is associated with age, sex, urban rural differences and the season of disease onset. METHODS: A register of patients with Type I diabetes using two independent sets of data sources was established in 1994 as part of the EURODIAB TIGER programme. The primary data sources were paediatric and internal medicine divisions of the hospitals in the Białystok province and the secondary were outpatient diabetic clinics in the region. The degree of ascertainment was 98.5 % for the combinated data sources. RESULTS: We found a significant rising trend in the incidence of Type I diabetes in children under 15 years of age (in 1998 the incidence was approximately twice as high as in 1994). Increasing incidence rates were observed in the rural areas but not in urban populations. Seasonal variation in the incidence was also found, with a peak in winter and nadir in summer. CONCLUSIONS/INTERPRETATION: These results show that the north-eastern region of Poland is an area with a moderate rather than a low risk of Type I diabetes. Our observations confirm the important role of environmental and socio-economic factors or both in the pathogenesis of Type I diabetes.

Adolescent↗

Medical records as sources of data on cardiovascular disease events in persons with diabetes.

PURPOSE: The aim of this study is to evaluate medical records as a source of data on cardiovascular disease over a 20-year interval. METHODS: Participants in a population-based cohort of persons with Type 1 diabetes were asked whether they had been told by a doctor that they had several specific cardiovascular events. In addition, they were asked when and where they were hospitalized for myocardial infarction, stroke, surgical procedures, and for other conditions and procedures. The medical care institution was contacted to obtain copies of the relevant hospitalization. RESULTS: Overall, the confirmation of the self-reported events was 86.0% when medical records were obtained. Percent confirmed varied with the diagnosis. Reports of poor circulation in the lower extremities were confirmed in 42.6%, stroke was confirmed in 70%, and coronary bypass surgery was confirmed in 100% of cases. The success of obtaining medical records was greater for those events that were reported to have occurred more recently than those reported further in the past, especially when 10 or more years had elapsed. CONCLUSION: Medical record confirmation of reported cardiovascular events in persons with Type 1 diabetes was high for some events when medical records could be obtained but was lower for "poor circulation" to the legs and stroke possibly related to the lack of specificity of our questions, to incorrect attribution of symptoms by the respondent, or to inaccurate recall of a physician's examination. Medical record confirmation was better for more recent than past events. Therefore, when hard copy documentation is needed, it should be sought within 10 years of the event.

Cardiovascular Diseases↗

Efficacy of the dietary supplement S-adenosyl-L-methionine.

OBJECTIVE: To review existing published clinical evidence surrounding the dietary supplement SAMe (S-adenosyl-L-methionine). DATA SOURCES: The majority of information was obtained from primary published literature identified through MEDLINE search (1966-February 2001). Information was also obtained through secondary and tertiary sources when available. STUDY SELECTION AND DATA EXTRACTION: All articles identified from data sources were evaluated and all relevant information included in this review. DATA SYNTHESIS: The majority of clinical trial evidence surrounds the application of SAMe for various depressive disorders, osteoarthrits, and fibromyalgia. Sample sizes of these trials and the dose employed have varied considerably. Several reviews and at least two meta-analyses have examined the available evidence surrounding SAMe in the therapy of depression for trials completed prior to 1994 and concluded that SAMe was superior to placebo in treating depressive disorders and approximately as effective as standard tricyclic antidepressants. Much of this information exists in the form of isolated case reports or solitary clinical trials. SAMe appears to be well tolerated, with the majority of adverse effects presenting as mild to moderate gastrointestinal complaints. However, it is apparent that this agent is not without risk of more significant psychiatric and cardiovascular adverse events. Information documenting drug or food interactions with SAMe is very limited. CONCLUSIONS: Consumers should be instructed to avoid unmonitored consumption of this dietary supplement until sufficient discussion has taken place with their primary healthcare provider. Although there exists significant potential for therapeutic application of SAMe, its uncertain risk profile precludes definitive recommendation at this time. Healthcare providers and consumers should likely temper their enthusiasm for this dietary supplement until sufficient information becomes available.

Animals↗

Agreement between hospital records and maternal recall of mode of delivery: evidence from 12 391 deliveries in the UK Millennium Cohort Study.

OBJECTIVE: The objective of this study was to measure the agreement between hospital records and maternal reporting of mode of delivery in a representative UK sample. DESIGN: Population-based survey (Millennium Cohort Study). SETTING: UK. POPULATION: A total of 12,391 singleton infants born in 2000-2002. METHODS: Mothers were interviewed when infants were approximately 9 months old. Information was collected by interview on many obstetric and perinatal factors including mode of delivery. Record linkage to the mother's delivery hospital records was undertaken in those who gave consent (90%). A matching record was found for 83%. Maternal report and hospital records were compared using mode of delivery classified into three (normal, assisted and caesarean) and six groups. Factors associated with disagreement between the two data sources were identified. MAIN OUTCOME MEASURE: Proportion of records in which there was agreement between the two data sources. RESULTS: Agreement between maternal report and hospital records was at least 94% using six mode of delivery groups and 98% using three groups. Much of the disagreement (57-63%, depending on country) was between forceps and ventouse, and between planned and emergency caesarean. Disagreement was more common in women whose babies were first born and in women not born in the UK. CONCLUSION: Our study confirms that maternal reporting of mode of delivery is highly reliable. This is important for clinical staff caring for women and those conducting epidemiological studies. Additional data sources may be necessary to gather reliable data from ethnic minority women, particularly those born outside the UK, or to distinguish forceps from ventouse, or planned from emergency caesarean section.

Cohort Studies↗

Paired MEG data set source localization using recursively applied and projected (RAP) MUSIC.

An important class of experiments in functional brain mapping involves collecting pairs of data corresponding to separate "Task" and "Control" conditions. The data are then analyzed to determine what activity occurs during the Task experiment but not in the Control. Here we describe a new method for processing paired magnetoencephalographic (MEG) data sets using our recursively applied and projected multiple signal classification (RAP-MUSIC) algorithm. In this method the signal subspace of the Task data is projected against the orthogonal complement of the Control data signal subspace to obtain a subspace which describes spatial activity unique to the Task. A RAP-MUSIC localization search is then performed on this projected data to localize the sources which are active in the Task but not in the Control data. In addition to dipolar sources, effective blocking of more complex sources, e.g., multiple synchronously activated dipoles or synchronously activated distributed source activity, is possible since these topographies are well-described by the Control data signal subspace. Unlike previously published methods, the proposed method is shown to be effective in situations where the time series associated with Control and Task activity possess significant cross correlation. The method also allows for straightforward determination of the estimated time series of the localized target sources. A multiepoch MEG simulation and a phantom experiment are presented to demonstrate the ability of this method to successfully identify sources and their time series in the Task data.

Algorithms↗

Scaling an expert system data mart: more facilities in real-time.

Clinical Data Repositories are being rapidly adopted by large healthcare organizations as a method of centralizing and unifying clinical data currently stored in diverse and isolated information systems. Once stored in a clinical data repository, healthcare organizations seek to use this centralized data to store, analyze, interpret, and influence clinical care, quality and outcomes. A recent trend in the repository field has been the adoption of data marts--specialized subsets of enterprise-wide data taken from a larger repository designed specifically to answer highly focused questions. A data mart exploits the data stored in the repository, but can use unique structures or summary statistics generated specifically for an area of study. Thus, data marts benefit from the existence of a repository, are less general than a repository, but provide more effective and efficient support for an enterprise-wide data analysis task. In previous work, we described the use of batch processing for populating data marts directly from legacy systems. In this paper, we describe an architecture that uses both primary data sources and an evolving enterprise-wide clinical data repository to create real-time data sources for a clinical data mart to support highly specialized clinical expert systems.

Computer Systems↗

The accuracy of general practitioner records of smoking and alcohol use: comparison with patient questionnaires.

BACKGROUND: General practitioner (GP) records are increasingly being used as sources of information on potential confounders such as smoking use and alcohol intake in epidemiological studies. The aim of this study was to assess the accuracy of GP records on smoking use and alcohol intake compared with data from patient questionnaires. METHODS: Patients registered with 42 practices in Oxfordshire that agreed to take part in a post-marketing surveillance study of omeprazole were sent a postal questionnaire that included questions about alcohol and tobacco use. Two years later, data on these aspects of lifestyle were abstracted from the GP records. RESULTS: A total of 892 patients agreed to take part in the study; 804 (90 per cent) completed the postal questionnaire, and the records of 856 (96 per cent) were reviewed. Information on smoking and alcohol use was present in 74 per cent and 63 per cent of GP records, respectively. Agreement between the two data sources was moderate for both smoking (kappa = 0.50) and alcohol use (kappa = 0.52). With regard to smoking, the main discrepancy between the two data sources was that 46 per cent (94/206) of patients who reported themselves as exsmokers were recorded as being never smokers in the GP record. With regard to alcohol, there were no systematic differences between the two data sources. CONCLUSION: Data from GP records on smoking status and alcohol use are incomplete and subject to some misclassification. This is a source of potential failed adjustment for confounding, which should be considered in epidemiological studies that make use of these records.

Alcohol Drinking↗

Practical options for estimating cost of hospital inpatient stays.

Analysts often estimate the cost of hospital services by applying cost/charge (c/c) ratios from federal or state data sources to the charges provided on hospital discharge records. Recently, a number of sources of discharge data are not permitting the release of hospital identities. This study compares several sources of c/c data for use in the restricted environment. Accounting data from four state systems and from files of the federal Centers for Medicare and Medicaid Services (formerly HCFA) are employed. In one analysis hospitals are grouped by selected characteristics. C/c varies by state and characteristics. Some HCFA and state measures track each other closely. A wider analysis of hospital-specific data for 51 states offers a separate test and extension of the initial results. The study supports a practical policy option of releasing grouped c/c ratios attached to discharge records when identity must be masked. Key words: hospital cost, cost to charge ratios, privacy protections.

Accounting↗

The performance of different lookback periods and sources of information for Charlson comorbidity adjustment in Medicare claims.

BACKGROUND: The Charlson Score is a particularly popular form of comorbidity adjustment in claims data analysis. However, the effects of certain implementation decisions have not been empirically examined. OBJECTIVE: To determine the effects of alternative data sources and lookback periods on the performance of Charlson scores in the prediction of mortality following hospitalization. SUBJECTS: A representative sample of 1,387 elderly patients hospitalized in 1993, drawn from the Medicare Current Beneficiary Survey (MCBS). Three years of linked Medicare claims and survey instruments were available for all patients, as was 2-year mortality follow-up. STATISTICAL METHODS: Nested Cox regression and comparisons of areas under the Receiver Operating Characteristic (ROC) curve were used to evaluate ability to predict mortality. RESULTS: Compared with a 1-year lookback involving solely inpatient claims, statistically and empirically significant improvements in the prediction of mortality are obtained by incorporating alternative sources of data (particularly 2 years of inpatient data and 1 year of outpatient and auxiliary claims), but only if indices derived from distinct sources of data are entered into the regression distinctly. The area under the ROC curve for 1-year mortality predication increases from 0.702 to 0.741 (P = 0.002). Furthermore, these improvements in explanatory power obtained whether one also controls for Charlson scores based on self-reported health history and/or secondary diagnoses from the claim for the index hospitalization itself. Finally, claims-based comorbidity adjustment performs comparably to survey-derived adjustment, with areas under the ROC curve of 0.702 and 0.704, respectively. CONCLUSIONS: The widespread practice of comorbidity adjustment in pre-existing administrative data sources can be improved by taking more complete advantage of existing administrative data sources.

Aged↗

Finding leading indicators for disease outbreaks: filtering, cross-correlation, and caveats.

Bioterrorism and emerging infectious diseases such as influenza have spurred research into rapid outbreak detection. One primary thrust of this research has been to identify data sources that provide early indication of a disease outbreak by being leading indicators relative to other established data sources. Researchers tend to rely on the sample cross-correlation function (CCF) to quantify the association between two data sources. There has been, however, little consideration by medical informatics researchers of the influence of methodological choices on the ability of the CCF to identify a lead-lag relationship between time series. We draw on experience from the econometric and environmental health communities, and we use simulation to demonstrate that the sample CCF is highly prone to bias. Specifically, long-scale phenomena tend to overwhelm the CCF, obscuring phenomena at shorter wave lengths. Researchers seeking lead-lag relationships in surveillance data must therefore stipulate the scale length of the features of interest (e.g., short-scale spikes versus long-scale seasonal fluctuations) and then filter the data appropriately--to diminish the influence of other features, which may mask the features of interest. Otherwise, conclusions drawn from the sample CCF of bi-variate time-series data will inevitably be ambiguous and often altogether misleading.

Bioterrorism↗

Completeness and validity of cancer registration in a major public referral hospital in Saudi Arabia.

BACKGROUND: The 1994 Saudi National Cancer Registry (NCR), a population-based registry, showed a crude incidence rate (CIR) of 39/100,000 for all cancers in the Saudi population. The low CIR suggested possible under-reporting, especially during the early years of operation. This study as aimed at estimating the number of missed cases due to under-reporting, and to assess the validity of reported data from a major public referral hospital in Riyadh city. MATERIALS AND METHODS: We compared cancer cases from the three data sources: medical records (MR), which were the original source of NCR cases; pathology reports (PR); and death certificates (DC). We estimated the missing cancer cases using the capture-recapture method with log-linear models of Fienberg to correct for interdependency between the three data sources. To assess the validity of the data, we reabstracted records of about 8% (39/4760 of previously reported cases from the same hospital. RESULTS: A total of 811 cancer cases were reported through the three sources, i.e., MR 611 (75.3%), PR 639 (78.8%) and DC 204 (25.2%). After fitting a series of log-linear models to the three sources of data, the three sources were found to be statistically dependent. Capture-repcapture method indicated that 384 cases were missed, giving an estimation of 1195 cancer cases to be reported. Using these 1195 estimated cases; the estimated ascertainment rates were 51% for medical records, 53% for pathology reports, 17% for death certificates, and 68% for the aggregated registry. In the validity assessment, major disagreement between the abstracted and reabstracted data was found to be highest for stage of disease (44%), followed by histology code and behavior (25.6%). Minor disagreements were most common for date of diagnosis (36%) and grade (36%). Overall, agreements were highest for laterality (95%), followed by primary site codes (90%) and basis of diagnosis (85%). Agreement of tumor description variables (site, histology, behavior, and stage) was 57%. CONCLUSION: Cancer registration will require substantial improvements in both completeness of reporting and data quality at the hospital level. Use of multiple data sources and estimation of missed cases will help ensure completeness of case registration.

Journal Article↗

Exploratory analysis of climate data using source separation methods.

We present an example of exploratory data analysis of climate measurements using a recently developed denoising source separation (DSS) framework. We analyzed a combined dataset containing daily measurements of three variables: surface temperature, sea level pressure and precipitation around the globe, for a period of 56 years. Components exhibiting slow temporal behavior were extracted using DSS with linear denoising. The first component, most prominent in the interannual time scale, captured the well-known El Niño-Southern Oscillation (ENSO) phenomenon and the second component was close to the derivative of the first one. The slow components extracted in a wider frequency range were further rotated using a frequency-based separation criterion implemented by DSS with nonlinear denoising. The rotated sources give a meaningful representation of the slow climate variability as a combination of trends, interannual oscillations, the annual cycle and slowly changing seasonal variations. Again, components related to the ENSO phenomenon emerge very clearly among the found sources.

Atmospheric Pressure↗