Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

The use of routinely collected computer data for research in primary care: opportunities and challenges.

INTRODUCTION: Routinely collected primary care data has underpinned research that has helped define primary care as a specialty. In the early years of the discipline, data were collected manually, but digital data collection now makes large volumes of data readily available. Primary care informatics is emerging as an academic discipline for the scientific study of how to harness these data. This paper reviews how data are stored in primary care computer systems; current use of large primary care research databases; and, the opportunities and challenges for using routinely collected primary care data in research. OPPORTUNITIES: (1) Growing volumes of routinely recorded data. (2) Improving data quality. (3) Technological progress enabling large datasets to be processed. (4) The potential to link clinical data in family practice with other data including genetic databases. (5) An established body of know-how within the international health informatics community. CHALLENGES: (1) Research methods for working with large primary care datasets are limited. (2) How to infer meaning from data. (3) Pace of change in medicine and technology. (4) Integrating systems where there is often no reliable unique identifier and between health (person-based records) and social care (care-based records-e.g. child protection). (5) Achieving appropriate levels of information security, confidentiality, and privacy. CONCLUSION: Routinely collected primary care computer data, aggregated into large databases, is used for audit, quality improvement, health service planning, epidemiological study and research. However, gaps exist in the literature about how to find relevant data, select appropriate research methods and ensure that the correct inferences are drawn.

Biomedical Research↗

Integrated statistical analysis of cDNA microarray and NIR spectroscopic data applied to a hemp dataset.

Both cDNA microarray and spectroscopic data provide indirect information about the chemical compounds present in the biological tissue under consideration. In this paper simple univariate and bivariate measures are used to investigate correlations between both types of high dimensional analyses. A large dataset of 42 hemp samples on which 3456 cDNA clones and 351 NIR wavelengths have been measured, was analyzed using graphical representations. For this purpose we propose clustered correlation and clustered discrimination images. Large, tissue-related differences are seen to dominate the cDNA-NIR correlation structure but smaller, more difficult to detect, variety-related differences can be found at specific cDNA clone/NIR wavelength combinations.

Algorithms↗

Profound effect of normalization on detection of differentially expressed genes in oligonucleotide microarray data analysis.

BACKGROUND: Oligonucleotide microarrays measure the relative transcript abundance of thousands of mRNAs in parallel. A large number of procedures for normalization and detection of differentially expressed genes have been proposed. However, the relative impact of these methods on the detection of differentially expressed genes remains to be determined. RESULTS: We have employed four different normalization methods and all possible combinations with three different statistical algorithms for detection of differentially expressed genes on a prototype dataset. The number of genes detected as differentially expressed differs by a factor of about three. Analysis of lists of genes detected as differentially expressed, and rank correlation coefficients for probability of differential expression shows that a high concordance between different methods can only be achieved by using the same normalization procedure. CONCLUSIONS: Normalization has a profound influence of detection of differentially expressed genes. This influence is higher than that of three subsequent statistical analysis procedures examined. Algorithms incorporating more array-derived information than gene-expression values alone are urgently needed.

Animals↗

PPD - Proteome Profile Database.

With the complete sequencing of multiple genomes, there have been extensions in the methods of sequence analysis from single gene/protein-based to analyzing multiple genes and proteins simultaneously. Therefore, there is a demand of user-friendly software tools that will allow mining of these enormous datasets. PPD is a WWW-based database for comparative analysis of protein lengths in completely sequenced prokaryotic and eukaryotic genomes. PPD's core objective is to create protein classification tables based on the lengths of proteins by specifying a set of organisms and parameters. The interface can also generate information on changes in proteins of specific length distributions. This feature is of importance when the user's interest is focused on some evolutionarily related organisms or on organisms with similar or related tissue specificity or life-style. PPD is available at: PPD Home.

Animals↗

A novel approach for nontargeted data analysis for metabolomics. Large-scale profiling of tomato fruit volatiles.

To take full advantage of the power of functional genomics technologies and in particular those for metabolomics, both the analytical approach and the strategy chosen for data analysis need to be as unbiased and comprehensive as possible. Existing approaches to analyze metabolomic data still do not allow a fast and unbiased comparative analysis of the metabolic composition of the hundreds of genotypes that are often the target of modern investigations. We have now developed a novel strategy to analyze such metabolomic data. This approach consists of (1) full mass spectral alignment of gas chromatography (GC)-mass spectrometry (MS) metabolic profiles using the MetAlign software package, (2) followed by multivariate comparative analysis of metabolic phenotypes at the level of individual molecular fragments, and (3) multivariate mass spectral reconstruction, a method allowing metabolite discrimination, recognition, and identification. This approach has allowed a fast and unbiased comparative multivariate analysis of the volatile metabolite composition of ripe fruits of 94 tomato (Lycopersicon esculentum Mill.) genotypes, based on intensity patterns of >20,000 individual molecular fragments throughout 198 GC-MS datasets. Variation in metabolite composition, both between- and within-fruit types, was found and the discriminative metabolites were revealed. In the entire genotype set, a total of 322 different compounds could be distinguished using multivariate mass spectral reconstruction. A hierarchical cluster analysis of these metabolites resulted in clustering of structurally related metabolites derived from the same biochemical precursors. The approach chosen will further enhance the comprehensiveness of GC-MS-based metabolomics approaches and will therefore prove a useful addition to nontargeted functional genomics research.

Automation↗

Appendectomy for appendicitis in patients with schizophrenia.

BACKGROUND: Anecdotal evidence suggests that schizophrenia patients who require surgery have a high rate of adverse outcomes. We searched the Department of Veterans Affairs national datasets to determine the clinical course of schizophrenia patients with appendicitis who underwent appendectomy. METHODS: The Patient Treatment File (the nationwide inpatient database for the Department of Veterans Affairs) and the Beneficiary Identification and Records Location System were searched to identify all patients with International Classification of Diseases, Ninth Revision, Clinical Modification diagnostic codes for schizophrenia or schizoaffective disorder diagnosed with appendicitis during fiscal years 1995 to 1999. Computer-based information was supplemented with chart-based data. We sought data on six common preoperative risk factors and 25 specific adverse outcomes, including death. RESULTS: There were 55 patients identified. The mean age was 49, and 96% were men. The median time from symptom onset to diagnosis of appendicitis was 3 days. A history of substance abuse was obtained in 16 (29%). Disruptive behavior was documented in 16 (29%). Restraints were used in 9 (9%). The appendix was perforated in 36 (66%) and gangrenous in 9 (16%). Thirty-one (56%) had > or = 1 complication; there were 2 in-hospital deaths (4%). CONCLUSIONS: This is the first report on this topic in the medical literature. Appendicitis is typically diagnosed late in schizophrenic patients. Adverse patient behaviors are frequent. The complication and death rates are high.

Appendectomy↗

The classification of subjects with joint complaints on incomplete biochemical and haematological datasets.

We performed a retrospective study on 163 subjects suffering from rheumatic fever (16), rheumatoid arthritis (36), lupus erythematosus (17), gout (21), arthrosis (50) and osteomyelitis (23). The number of variables evaluated was 39. These were all of a general biochemical and haematological nature. A feature reduction resulted in sixteen variables that matched well with those known from the literature. Linear discriminant analysis yielded poor results in classifying the six disease categories (with 18 variables 61.8%). A reduction to three disease categories improved the classification results remarkably. This, and the excellent discriminating power between patients and the reference group, shows that the selected variables are illustrative only for general clinical pictures, such as infection, and not for the desired differential diagnosis.

Arthritis, Rheumatoid↗

RankProd: a bioconductor package for detecting differentially expressed genes in meta-analysis.

UNLABELLED: While meta-analysis provides a powerful tool for analyzing microarray experiments by combining data from multiple studies, it presents unique computational challenges. The Bioconductor package RankProd provides a new and intuitive tool for this purpose in detecting differentially expressed genes under two experimental conditions. The package modifies and extends the rank product method proposed by Breitling et al., [(2004) FEBS Lett., 573, 83-92] to integrate multiple microarray studies from different laboratories and/or platforms. It offers several advantages over t-test based methods and accepts pre-processed expression datasets produced from a wide variety of platforms. The significance of the detection is assessed by a non-parametric permutation test, and the associated P-value and false discovery rate (FDR) are included in the output alongside the genes that are detected by user-defined criteria. A visualization plot is provided to view actual expression levels for each gene with estimated significance measurements. AVAILABILITY: RankProd is available at Bioconductor http://www.bioconductor.org. A web-based interface will soon be available at http://cactus.salk.edu/RankProd

Computational Biology↗

Dupuytren's disease risk factors.

Dupuytren's is a common problem, but little is known about its aetiology. We have undertaken a large case-control study to assess and quantify the relative contributions of diabetes and epilepsy as risk factors for Dupuytren's in the community. Cases were patients with a diagnosis of Dupuytren's disease and, for each, two controls were individually matched by age, sex, and general practice. Our dataset included 821 cases and 1,642 controls. Five hundred and eighty-eight (72%) of the cases were men. The mean age at diagnosis was 62 (range 24-97) years. Diabetes was a significant risk factor for Dupuytren's disease (OR=1.75) and there was an increased risk for medicinally treated diabetes (metformin--OR=3.56; sulphonylureas--OR=1.75) and particularly insulin controlled (OR=4.38) rather than diet-controlled diabetes. Epilepsy (OR=1.12) and anti-epileptic medications were not associated with Dupuytren's disease. Ascertainment bias in previous studies may explain the reported association with epilepsy.

Adult↗

Multi-species assessment of electrical resistance as a skin integrity marker for in vitro percutaneous absorption studies.

Assessment of percutaneous absorption in vitro provides key information when predicting dermal absorption in vivo. Confirmation of skin membrane integrity is an essential component of the in vitro method, as described in test guideline OECD 428. Historically, assessment of the membrane's permeability to tritiated water (T2O) and the generation of a permeability coefficient (Kp) were used to confirm that the skin membrane was intact prior to application of the test penetrant. Measuring electrical resistance (ER) across the membrane is a simpler, quicker, safer and more cost effective method. To investigate the robustness of the ER integrity measure, the Kp values for T2O for a range of human and animal skin membranes were compared with corresponding ER data. Overall, for human, rat, pig, mouse, rabbit and guinea pig skin, the ER data gave a good inverse association with the corresponding Kp values; the higher the Kp the lower the ER values. In addition, the distribution across a large dataset for individual skin samples was similar for Kp and ER, allowing a cut-off value for ER to be established for each skin type. Based on CTL's (Syngenta Central Toxicology Laboratory) standard static diffusion cells and databridge, we propose that intact skin should have an ER equal to or above (in kOmega): human (10), mouse (5) guinea pig (5), pig (4) rat (3), and rabbit (0.8). We conclude that measurement of ER across in vitro skin membranes provides a robust measurement of skin barrier integrity and is an appropriate alternative to Kp for T2O in order to identify intact membranes that have acceptable permeability characteristics for in vitro percutaneous absorption studies.

Administration, Topical↗

Serotonergic brainstem abnormalities in Northern Plains Indians with the sudden infant death syndrome.

The rate of the sudden infant death syndrome (SIDS) among American Indian infants in the Northern Plains is almost 6 times higher than in U.S. white infants. In a study of infant mortality among Northern Plains Indians, we tested the hypothesis that receptor binding abnormalities to the neurotransmitter serotonin (5-HT) in SIDS cases, compared with autopsied controls, occur in regions of the medulla oblongata that contain 5-HT neurons and that are critical for the regulation of cardiorespiration and central chemosensitivity during sleep, i.e. the medullary 5-HT system. Tritiated-lysergic acid diethylamide binding to 5-HT(1A-D) and 5-HT2 receptors was measured in 19 brainstem nuclei in 23 SIDS and 6 control infants using tissue receptor autoradiography. Binding in the arcuate nucleus, a part of the medullary 5-HT system along the ventral surface, in the SIDS infants (mean age-adjusted binding 7.1 +/- 0.8 fmol/mg tissue, n = 23) was significantly lower than in controls (mean age-adjusted binding 13.1 +/- 1.6 fmol/mg tissue, n = 5) (p = 0.003). Binding also demonstrated significant diagnosis x age interactions (p < 0.04) in 4 other nuclei that are components of the 5-HT system. These data suggest that medullary 5-HT dysfunction can lead to sleep-related, sudden death in affected SIDS infants, and confirm the same binding abnormalities reported by us in a larger dataset of non-American Indian SIDS and control infants. This study also links 5-HT abnormalities in the arcuate nucleus with exposure to adverse prenatal exposures, i.e. cigarette smoking (p = 0.011) and alcohol (p = 0.075), during the periconceptional period or throughout pregnancy. Prenatal exposure to cigarette smoke and/or alcohol may contribute to abnormal fetal medullary 5-HT development in SIDS infants.

Age Factors↗

Effect of smoking status on outcome after acute ischemic stroke.

BACKGROUND: The status of smoking as a risk factor for the occurrence of stroke is well established. However, there is a paucity of data on the relationship between smoking status and acute stroke outcomes. We evaluated the role of recent smoking as a prognostic factor following acute ischemic stroke. METHODS: We analyzed data from patients enrolled in the Intravenous Magnesium Efficacy in Stroke (IMAGES) trial. Outcome measures studied included change in IMAGES stroke score, poor functional outcomes at day 30 and 90 (defined as Rankin Scale >1 and Barthel Index <95), and survival over the first 3 months after stroke. The independent effect of smoking status (subjects who had smoked in the past year) on outcome was evaluated by logistic regression analysis and Cox's proportional hazards model, adjusting for variables known to predict outcome after ischemic stroke. RESULTS: There were 2,386 subjects in the IMAGES efficacy dataset, including 615 recent or current smokers and 1,771 nonsmokers, among whom smokers were younger (p < 0.0001). After adjusting for covariates, smokers had increased odds of poor 90-day functional outcome independently of other statistically significant predictor variables, as assessed by Rankin Scale (odds ratio 1.38; 95% confidence interval 1.09-1.75) and Barthel Index (odds ratio 1.42; 95% confidence interval 1.13-1.79) at day 90. Smoking status did not affect survival at day 90. CONCLUSIONS: Current or recent smokers experience poorer functional outcomes than nonsmokers 3 months after acute ischemic stroke.

Aged↗

Metrics for external model evaluation with an application to the population pharmacokinetics of gliclazide.

PURPOSE: The aim of this study is to define and illustrate metrics for the external evaluation of a population model. MATERIALS AND METHODS: In this paper, several types of metrics are defined: based on observations (standardized prediction error with or without simulation and normalized prediction distribution error); based on hyperparameters (with or without simulation); based on the likelihood of the model. All the metrics described above are applied to evaluate a model built from two phase II studies of gliclazide. A real phase I dataset and two datasets simulated with the real dataset design are used as external validation datasets to show and compare how metrics are able to detect and explain potential adequacies or inadequacies of the model. RESULTS: Normalized prediction errors calculated without any approximation, and metrics based on hyperparameters or on objective function have good theoretical properties to be used for external model evaluation and showed satisfactory behaviour in the simulation study. CONCLUSIONS: For external model evaluation, prediction distribution errors are recommended when the aim is to use the model to simulate data. Metrics through hyperparameters should be preferred when the aim is to compare two populations and metrics based on the objective function are useful during the model building process.

Algorithms↗

Design and analysis of trials with quality of life as an outcome: a practical guide.

Health Related Quality of Life (HRQoL) measures are becoming more frequently used in clinical trials, as both primary and secondary endpoints. Investigators are now asking statisticians for advice on how to plan (e.g., sample size) and analyze studies using HRQoL measures. HRQoL measures such as the SF-36 are usually measured on an ordered categorical (ordinal) scale. In the designing stages and when analyzing, the scales are often scored and the scores treated as if they were continuous and normally distributed. However the ordinal scaling of HRQoL measures leads to problems in determining sample size, and conventional parametric methods of estimation and hypothesis testing may not be appropriate for such outcomes. We present practical guidelines for the design and analysis of trials with HRQoL measures as outcomes. We used conventional statistical methods (i.e., t-tests and multiple regression), various ordinal regression models (proportional odds, continuation ratio, polytomous and stereotype) and bootstrap methods to analyze an HRQoL dataset. To illustrate the various methods we used HRQoL data on the SF-36 Role Limitations Emotional dimension for two groups of patients with leg ulcers. The bootstrap, t-test, and multiple regression methods gave similar results. The various ordinal regression models also gave similar results. If the HRQoL measure has a large number of ordered categories, most of which are occupied, and the underlying scale really is continuous but measured imperfectly by an instrument with a limited number of discrete values, then an informal rule of thumb is that this discrete scale should be treated as continuous if it has seven or more categories and as ordinal otherwise.

Algorithms↗

Use of narrative analysis for comparisons of the causes of fatal accidents in three countries: New Zealand, Australia, and the United States.

OBJECTIVES: To investigate the utility of narrative analysis of text information for describing the mechanism of injury and to compare the patterns of the mechanism of injury for work related fatalities in three countries. METHODS: Three national collections of data on work related fatalities were used in this study including those for New Zealand, 1985-94 (n=723), for Australia, 1989-92 (n=1,220), and for the United States, 1989-92 (16,383). The New Zealand and Australian collections used the type of occurrence standard code for the mechanism of injury, however the United States collection did not. All three databases included a text description of the circumstances of the fatality so a text based analysis was developed to enable a comparison of the mechanisms of injury in each of the three countries. A test set of 200 cases from each country dataset was used to develop the narrative analysis and to allow comparison of the narrative and standard approaches to mechanism coding. RESULTS: The narrative coding was more useful for some types of injury than others. Differences in coding the narrative codes compared with the standard code were mainly due to lack of sensitivity in detecting cases for all three datasets, although specificity was always high. The pattern of causes was very similar between the two coding methods and between the countries. Hit by moving objects, falls, and rollovers were among the five most common mechanisms of workplace fatalities for all countries. More common mechanisms that distinguished the three countries were electrocutions for Australia, drowning for New Zealand, and gunshot for the United States. CONCLUSION: Narrative analysis shows some promise as an alternative approach for investigating the causes of fatalities.

Accidents, Occupational↗

Linkage disequilibrium mapping via cladistic analysis of phase-unknown genotypes and inferred haplotypes in the Genetic Analysis Workshop 14 simulated data.

We recently described a method for linkage disequilibrium (LD) mapping, using cladistic analysis of phased single-nucleotide polymorphism (SNP) haplotypes in a logistic regression framework. However, haplotypes are often not available and cannot be deduced with certainty from the unphased genotypes. One possible two-stage approach is to infer the phase of multilocus genotype data and analyze the resulting haplotypes as if known. Here, haplotypes are inferred using the expectation-maximization (EM) algorithm and the best-guess phase assignment for each individual analyzed. However, inferring haplotypes from phase-unknown data is prone to error and this should be taken into account in the subsequent analysis. An alternative approach is to analyze the phase-unknown multilocus genotypes themselves. Here we present a generalization of the method for phase-known haplotype data to the case of unphased SNP genotypes. Our approach is designed for high-density SNP data, so we opted to analyze the simulated dataset. The marker spacing in the initial screen was too large for our method to be effective, so we used the answers provided to request further data in regions around the disease loci and in null regions. Power to detect the disease loci, accuracy in localizing the true site of the locus, and false-positive error rates are reported for the inferred-haplotype and unphased genotype methods. For this data, analyzing inferred haplotypes outperforms analysis of genotypes. As expected, our results suggest that when there is little or no LD between a disease locus and the flanking region, there will be no chance of detecting it unless the disease variant itself is genotyped.

Chromosome Mapping↗

A new multiparameter flow cytometer: optical and electrical cell analysis in combination with video microscopy in flow.

BACKGROUND: Flow cytometers, which are commercially available, do not necessarily meet all demands of actual biomedical research. This is the case for the investigation of mechanisms involved in cell volume regulation, which requires electrical volume measurement and ratiometric multichannel fluorescence analysis for the simultaneous assessment of different physiologic parameters (intracellular pH and the intracellular concentration of calcium ions, etc). METHODS AND RESULTS: We describe the construction of a new nonsorting flow cytometer designed for the simultaneous acquisition of seven parameters including fluorescence signals, forward and perpendicular light scatter, cell volume according to the electrical Coulter principle, and flow cytometric imaging. The instrument is equipped with three different light sources. A tunable argon-ion laser generates efficient excitation of the most standard fluorescent probes in the visible spectral range, and an arc lamp provides the means for ultraviolet excitation at low cost. Because of the spatial filtering by the excitation and detection optics, two independent sets of dual fluorescence measurements can be performed, a prerequisite for flexible ratiometric fluorescence analysis. A flow video microscope integrated into the optical system optionally generates either brightfield or phase images of selected flowing particles. Only particles whose individual datasets meet predefined gating conditions are imaged in real time. To avoid smear effects, the motion of the object to be imaged (speed approximately 8 m/s) is frozen on the target of a CCD camera by flash illumination. For this purpose, a high radiance gas discharge lamp with 25-mJ electric pulse energy provides an illumination time of 18 ns (full width half maximum). Test results obtained from latex spheres and cells are shown. CONCLUSIONS: Test results indicate that our instrument can perform Coulter measurements in combination with flexible optical analysis. Moreover, integration of an adapted video microscope into a flow cytometer is an approach to overcome the gap between flow and image cytometry.

Animals↗

Non-major-histocompatibility-complex genetics of ankylosing spondylitis.

There is strong evidence from twin and family studies indicating that a substantial proportion of the heritability of susceptibility to ankylosing spondylitis (AS) and its clinical manifestations is encoded by non-major-histocompatibility-complex genes. Efforts to identify these genes have included genomewide linkage studies and candidate gene association studies. One region, the interleukin (IL)-1 gene complex on chromosome 2, has been repeatedly associated with AS in both Caucasians and Asians. It is likely that more than one gene in this complex is involved in AS, with the strongest evidence to date implicating IL-1A. Identifying the genes underlying other linkage regions has been difficult due to the lack of obvious candidates and the low power of most studies to date to identify genes of the small to moderate magnitude that are likely to be involved. The field is moving towards genomewide association analysis, involving much larger datasets of unrelated cases and controls. Early successes using this approach in other diseases indicates that it is likely to identify genes in common diseases like AS, but there remains the risk that the common-variant, common-disease hypothesis will not hold true in AS. Nonetheless, it is appropriate for the field to be cautiously optimistic that the next few years will bring great advances in our understanding of the genetics of this condition.

Asian People↗