Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Personalised risk communication for informed decision making about entering screening programs.

BACKGROUND: There is a trend towards greater patient involvement in health care decisions. Adequate discussion of the risks and benefits associated with different choices is often required if involvement is to be genuine and effective. Achieving adequate involvement of consumers and informed decision making are now seen as important goals for any screening programme. Individualised risk estimates have been shown to be effective methods of risk communication in general, but the effectiveness of different strategies has not previously been examined. OBJECTIVES: To assess the effects of different types of individualised risk communication for consumers making decisions about participating in screening. SEARCH STRATEGY: We searched the Cochrane Consumers and Communication Review Group specialised register (searched March 2001), MEDLINE (1985 to 2001), EMBASE (1985 to 2001), CancerLit (1985 to 2001), CINAHL (1985 to 2001), ClinPSYC (1989 to 2001), and the Science Citation Index Expanded (searched March 2002). Follow-up searches involved hand searching Preventive Medicine, citation searches on seven authors, and searching reference lists of articles. SELECTION CRITERIA: Randomised controlled trials addressing the decision by consumers of whether or not to undergo screening, incorporating an intervention with a 'personalised risk communication element' and reporting cognitive, affective, or behavioural outcomes. A 'personalised risk communication element' is based on the individual's own risk factors for a condition (such as age or family history). It may be calculated from an individual's risk factors using formulae derived from epidemiological data, and presented as an absolute risk or as a risk score, or it may be categorised into, for example, high, medium or low risk groups. It may be less detailed still, involving a listing, for example, of a consumer's risk factors as a focus for discussion and intervention. DATA COLLECTION AND ANALYSIS: Two reviewers independently assessed trial quality and extracted data. Data about the nature and setting of the intervention, and the relevant outcome data were extracted, along with items relating to methodological quality. MAIN RESULTS: Thirteen studies were included. Personalised risk communication (whether written, spoken or visually presented) was associated with increased uptake of screening tests (odds ratio (OR) 1.5 (95% confidence interval (CI) 1.11 to 2.03). There was no evidence from these studies that this increase in uptake of tests was related to informed decision making by consumers. More detailed personalised risk communication was associated with a smaller increase in uptake of tests. That is, for personalised risk communication which used and presented numerical calculations of risk, the OR for test uptake was 1.22 (95% CI 0.56 to 2.68). For risk estimates or calculations which were categorised into high, medium or low strata of risk, the OR was 1.42 (95% CI 1.07 to 1.88). For risk communication that simply listed risk personal risk factors the OR was 1.7 (95% CI 1.17 to 2.48). Most of the included studies addressed mammography programmes. These studies showed slightly smaller effects than the overall dataset, again with numerical calculated risk estimates being associated with lower ORs for uptake of tests (OR 1.13; 95% CI 0.98 to 1.29) than the other categories of (less detailed) personalised risk communication. The four studies examining risk communication in high risk individuals showed larger odds ratios for uptake of tests than the other studies. The OR for numerical calculated risk estimates was 1.48 (95% CI 1.06 to 2.07), compared to 4.66 (95% CI 2.24 to 9.69) for categorised risk estimates and 2.64 (95% CI 1.42 to 4.9) for listed personal risk factors. There were insufficient data from the included studies to report odds ratios on other key outcomes such as: intention to take tests, anxiety, satisfaction with decisions, decisional conflict, knowledge and risk perception. REVIEWER'S CONCLUSIONS: Personalised risk communication (as currently implemented in the included studies) is associated with increased uptake of screening programmes, but this may not be interpretable as evidence of informed decision making by consumers.

Communication↗

The prevalence of problem drug misuse in a rural county of England.

Previous capture-recapture studies have estimated the prevalence of problem drug misuse in urban areas. This study estimates the prevalence in a rural county, Norfolk, using data from four sources: drug treatment agencies, probation, the arrest referral service, and police (drug-related crime with/without acquisitive crime). Careful consideration was given to methods of matching datasets and sensitivity analyses involved altering matching rules and postcode criteria. Whilst it is recognised that acquisitive crime is often related to drug use, this is the first capture-recapture study to incorporate acquisitive crime data. In further sensitivity analyses the proportion of acquisitive crime assumed to be drug-related was varied from 25-60%. The main analysis provided an estimated prevalence of problem drug use in Norfolk of 2.05% (95% confidence interval: 1.66%-2.56%) for ages 15-54 years, considerably higher than the 1.1% currently suggested for the UK. Sensitivity analyses based on varied matching and postcode criteria produced estimates ranging from 2.41%-3.37%, suggesting our estimate may be conservative. Sensitivity analyses assuming that 25-60% of acquisitive crimes were drug-related, produced estimates ranging from 2.02% to 5.73%, further supporting our main analysis. In conclusion, this study provides evidence that problem drug misuse is more prevalent in this rural population than previously thought.

Adolescent↗

Assessment of analysis-of-variance-based methods to quantify the random variations of observers in medical imaging measurements: guidelines to the investigator.

The random variations of observers in medical imaging measurements negatively affect the outcome of cancer treatment, and should be taken into account during treatment by the application of safety margins that are derived from estimates of the random variations. Analysis-of-variance- (ANOVA-) based methods are the most preferable techniques to assess the true individual random variations of observers, but the number of observers and the number of cases must be taken into account to achieve meaningful results. Our aim in this study is twofold. First, to evaluate three representative ANOVA-based methods for typical numbers of observers and typical numbers of cases. Second, to establish guidelines to the investigator to determine which method, how many observers, and which number of cases are required to obtain the a priori chosen performance. The ANOVA-based methods evaluated in this study are an established technique (pairwise differences method: PWD), a new approach providing additional statistics (residuals method: RES), and a generic technique that uses restricted maximum likelihood (REML) estimation. Monte Carlo simulations were performed to assess the performance of the ANOVA-based methods, which is expressed by their accuracy (closeness of the estimates to the truth), their precision (standard error of the estimates), and the reliability of their statistical test for the significance of a difference in the random variation of an observer between two groups of cases. The highest accuracy is achieved using REML estimation, but for datasets of at least 50 cases or arrangements with 6 or more observers, the differences between the methods are negligible, with deviations from the truth well below +/-3%. For datasets up to 100 cases, it is most beneficial to increase the number of cases to improve the precision of the estimated random variations, whereas for datasets over 100 cases, an improvement in precision is most efficiently achieved by increasing the number of observers. For datasets of at least 50 cases, the standard error ranges between 30% or less with 3 observers down to 10% or less with 8 observers, and the differences in precision between the methods are negligible. The F test (PWD) is very anticonservative and should not be used, while the t test (RES) is reliable for datasets of at least 2 x 50 cases evaluated by 4 or more observers. The likelihood-ratio-test (REML estimation) consistently indicates the significance of a difference in the random variation of an observer between two groups of cases, regardless of the number of cases, and regardless of the number of observers. If a statistical package to perform REML estimation is available, and the investigator feels confident using it, this is the preferred method for studies that involve less than 50 cases evaluated by less than 6 observers. Otherwise, the RES method is an excellent alternative, because of its straightforward implementation, its completeness with respect to the provided statistics, and its overall sufficient accuracy, precision, and reliability of the provided statistical test. If neither the RES method nor REML estimation can provide sufficient performance, either more observers or more cases must be included.

Algorithms↗

A cryptologic based trust center for medical images.

OBJECTIVE: To investigate practical solutions that can integrate cryptographic techniques and picture archiving and communication systems (PACS) to improve the security of medical images. DESIGN: The PACS at the University of California San Francisco Medical Center consolidate images and associated data from various scanners into a centralized data archive and transmit them to remote display stations for review and consultation purposes. The purpose of this study is to investigate the model of a digital trust center that integrates cryptographic algorithms and protocols seamlessly into such a digital radiology environment to improve the security of medical images. MEASUREMENTS: The timing performance of encryption, decryption, and transmission of the cryptographic protocols over 81 volumetric PACS datasets has been measured. Lossless data compression is also applied before the encryption. The transmission performance is measured against three types of networks of different bandwidths: narrow-band Integrated Services Digital Network, Ethernet, and OC-3c Asynchronous Transfer Mode. RESULTS: The proposed digital trust center provides a cryptosystem solution to protect the confidentiality and to determine the authenticity of digital images in hospitals. The results of this study indicate that diagnostic images such as x-rays and magnetic resonance images could be routinely encrypted in PACS. However, applying encryption in teleradiology and PACS is a tradeoff between communications performance and security measures. CONCLUSION: Many people are uncertain about how to integrate cryptographic algorithms coherently into existing operations of the clinical enterprise. This paper describes a centralized cryptosystem architecture to ensure image data authenticity in a digital radiology department. The system performance has been evaluated in a hospital-integrated PACS environment.

Algorithms↗

IUD--related uterine perforation: an epidemiologic analysis of a rare event using an international dataset.

An international dataset of 21,610 IUD insertions revealed 41 uterine perforations occurring at the time of, or subsequent to, insertion. Dat were collected on standard forms. Perforations were classified as confirmed, probable, or possible, based on the clinician's judgement and subsequent management. The uterine perforation rate was estimated as between 1.9 and 3.6/1000 insertions. 13 of the total 41 perforations were reported from 2 of 72 cooperating clinics. In 1 clinic, this center-clustering phenomenon suggested an effect of a high risk device. In the other clinic, the effect of insertor (operator) inexperience was indicated. A case-control analysis delineated previous cesarean section as a host risk factor (P0.05, McNemar's chi square). Further investigation of this association by similar case-control studies is recommended.

Age Factors↗

Do supplementary items on the eating disorder examination improve the assessment of adolescents with anorexia nervosa?

OBJECTIVE: Given that adolescents with anorexia nervosa (AN) typically have lower scores on the Eating Disorder Examination (EDE) than expected, the current study examined whether the inclusion of eight supplementary items developed by the authors of the EDE better captured the symptoms of adolescents with AN. METHOD: A dataset consisting of EDEs from 86 adolescents was examined by 3 primary methods: (1) baseline subscale scores were compared before and after the addition of the supplementary items, (2) the internal consistency of the EDE with the addition of these items was examined, and (3) each of these items was compared before and after treatment. RESULTS: After the addition of the supplementary items, the Eating Concern and Weight Concern subscales were significantly increased, whereas the Restraint subscale was significantly decreased, and the Shape Concern subscale was unchanged. Internal consistency was improved on the Eating Concern, Weight Concern, and Shape Concern subscales, and was decreased on the Restraint subscale. Three of eight items showed a significant decrease with treatment. CONCLUSION: Although the addition of some of these eight supplementary items better captured the psychopathology of adolescents with AN, scores were still substantially below expected, indicating that the exploration of other methods of assessment is needed.

Adolescent↗

Quantifying psychiatric comorbidity--lessions from chronic disease epidemiology.

BACKGROUND: Comorbidity research in psychiatric epidemiology mostly uses measures of association like odds or risk ratios to express how strongly disorders are linked. In contrast, chronic disease epidemiologists increasingly use measures of clustering, like multimorbidity (cluster) coefficients, to study comorbidity. This article compares measures of association and clustering. METHODS: Narrative review, algebraical examples, a secondary analysis of an existing dataset and a pooled analysis of published data. RESULTS: Odds and risk ratios, but the former more than the latter, confound clustering with coincidental comorbidity. Multimorbidity coefficients provide a pure estimate of clustering which is the proportion of the association between disorders that is of etiological interest. Odds and risk ratios can express comorbidity between no more than two disorders, whilst clustering coefficients, although computationally laboursome, can capture multimorbidity of any number of disorders. Cluster coefficients depend less on the prevalence of illness in study groups than measures of association. CONCLUSION: Odds and risk ratios are well suited for comorbidity research which focuses on which sets of disorders or syndromes tend to occur in combination and the implications of this for, for instance, nosological classification, a traditional interest of psychiatric epidemiology. However, the cluster coefficient is to be preferred if the interest is more aetiological, addressing for example why certain individuals are prone to multiple health problems.

Alcoholism↗

Reliability and validity of an algorithm for fuzzy tissue segmentation of MRI.

PURPOSE: A new multistep, volumetric-based tissue segmentation algorithm that results in fuzzy (or probabilistic) voxel description is described. This algorithm is designed to accurately segment gray matter, white matter, and CSF and can be applied to both single channel high resolution and multispectral (multiecho) MR images. METHOD: The reliability and validity of this method are evaluated by assessing (a) the stability of the algorithm across time, rater, and pulse sequence; (b) the accuracy of the method when applied to both real and synthetic image datasets; and (c) differences in specific tissue volumes between individuals with a specific genetic condition (fragile X syndrome) and normal control subjects. RESULTS: The algorithm was found to have high reliability, accuracy, and validity. The finding of increased caudate gray matter volume associated with the fragile X syndrome is replicated in this sample. CONCLUSION: Since this segmentation approach incorporates "fuzzy" or probabilistic methods, it has the potential to more accurately address partial volume effects, anatomical variation within "pure" tissue compartments, and more subtle changes in tissue volumes as a result of disease and treatment. The method is a component of software that is available in the public domain and has been implemented on an inexpensive personal computer thus offering an attractive and promising method for determining the status and progression of both normal development and pathology of the CNS.

Adolescent↗

Exercise training meta-analysis of trials in patients with chronic heart failure (ExTraMATCH).

OBJECTIVE: To determine the effect of exercise training on survival in patients with heart failure due to left ventricular systolic dysfunction. DESIGN: Collaborative meta-analysis. Inclusion criteria Randomised parallel group controlled trials of exercise training for at least eight weeks with individual patient data on survival for at least three months. Studies reviewed Nine datasets, totalling 801 patients: 395 received exercise training and 406 were controls. MAIN OUTCOME MEASURE: Death from all causes. RESULTS: During a mean (SD) follow up of 705 (729) days there were 88 (22%) deaths in the exercise arm and 105 (26%) in the control arm. Exercise training significantly reduced mortality (hazard ratio 0.65, 95% confidence interval, 0.46 to 0.92; log rank chi(2) = 5.9; P = 0.015). The secondary end point of death or admission to hospital was also reduced (0.72, 0.56 to 0.93; log rank chi(2) = 6.4; P = 0.011). No statistically significant subgroup specific treatment effect was observed. CONCLUSION: Meta-analysis of randomised trials to date gives no evidence that properly supervised medical training programmes for patients with heart failure might be dangerous, and indeed there is clear evidence of an overall reduction in mortality. Further research should focus on optimising exercise programmes and identifying appropriate patient groups to target.

Exercise Therapy↗

Simulation-based homozygosity mapping with the GAW14 COGA dataset on alcoholism.

BACKGROUND: We have developed a simulation-based approach to the analysis of shared homozygous chromosomal segments and have applied it to data on allele sharing among alcoholics in a single Collaborative Study on the Genetics of Alcoholism pedigree. Our assessment of sharing involved the use of a single-nucleotide polymorphism (SNP) marker map provided by Affymetrix. RESULTS: All 11 affected individuals in the selected pedigree shared 2 copies of an allele at 4 adjacent SNPs in a region on chromosome 5. Via simulation, we determined that the probability that such sharing is caused by mere chance is less than 0.0000001. After correcting for undocumented inbreeding, this probability rose to 0.0016. The probability that the shared segment emanates from a single ancestor and is unrelated to the affection status is less than 0.0000001 in the corrected pedigree. Haplotype association analysis and a search for a protective locus using unaffected individuals yielded no significant results. CONCLUSION: Homozygosity mapping results on chromosome 5 provide suggestive evidence of the region's role as one that may harbor a genetic determinant of alcoholism. Furthermore, the probabilities of chance homozygous allele sharing for the original and for the inbreeding-corrected pedigree provide insight into the impact that inbreeding can have on such calculations.

Alcoholism↗

Bio-Ontology and text: bridging the modeling gap.

MOTIVATION: Natural language processing (NLP) techniques are increasingly being used in biology to automate the capture of new biological discoveries in text, which are being reported at a rapid rate. Yet, information represented in NLP data structures is classically very different from information organized with ontologies as found in model organisms or genetic databases. To facilitate the computational reuse and integration of information buried in unstructured text with that of genetic databases, we propose and evaluate a translational schema that represents a comprehensive set of phenotypic and genetic entities, as well as their closely related biomedical entities and relations as expressed in natural language. In addition, the schema connects different scales of biological information, and provides mappings from the textual information to existing ontologies, which are essential in biology for integration, organization, dissemination and knowledge management of heterogeneous phenotypic information. A common comprehensive representation for otherwise heterogeneous phenotypic and genetic datasets, such as the one proposed, is critical for advancing systems biology because it enables acquisition and reuse of unprecedented volumes of diverse types of knowledge and information from text. RESULTS: A novel representational schema, PGschema, was developed that enables translation of phenotypic, genetic and their closely related information found in textual narratives to a well-defined data structure comprising phenotypic and genetic concepts from established ontologies along with modifiers and relationships. Evaluation for coverage of a selected set of entities showed that 90% of the information could be represented (95% confidence interval: 86-93%; n = 268). Moreover, PGschema can be expressed automatically in an XML format using natural language techniques to process the text. To our knowledge, we are providing the first evaluation of a translational schema for NLP that contains declarative knowledge about genes and their associated biomedical data (e.g. phenotypes). AVAILABILITY: http://zellig.cpmc.columbia.edu/PGschema

Abstracting and Indexing↗

Statistical analysis of blood- to breath-alcohol ratio data in the logarithm-transformed and non-transformed modes.

The statistical analysis of non-transformed and logarithm-transformed blood- to breath-alcohol ratios ("blood/breath ratios") is detailed. The data analyzed were derived from 137 simultaneous blood-alcohol and breath-alcohol concentration measurements made between 15 and 179 min after the end of drinking, with 136 of the measurements obtained during the 15- to 124-min time frame. Although the distribution of the non-transformed ratios is positively skewed, and that of the logarithm-transformed data more closely approximates the normal distribution upon visual inspection, both analyses generated results that do not differ significantly from each other when considered in the context of "mean ratios + or - 2SD". This is in accord with the results of the Kolmogorov-Smirnov goodness-of-fit test, which does not reject either dataset and demonstrates that both are approximately normal. Since the logarithm-transformed data generate more conservative statistical blood/breath ratio ranges than the non-transformed data, they were selected as the basis for the principal conclusion of this work. That conclusion is a refutation of the argument that, breath-alcohol analyzers relying on a 2100:1 blood/breath ratio tend to underestimate the blood-alcohol concentrations of driving-while-intoxicated arrestees because the commonly accepted mean postabsorptive ratio is 2300:1. In fact, whenever the absorption status of a driving-while-intoxicated arrestee at the time of a breath test cannot be definitively established, the results of this work support the application of a relative error range of -40% to +28% for 95% of the population, based on a statistical blood/breath ratio range of 1259:1 to 2679:1, and -46% to +42% for 99% of the population, based on a statistical range of 1128:1 to 2989:1.

Absorption↗

The prognostic value of serum myoglobin in patients with non-ST-segment elevation acute coronary syndromes. Results from the TIMI 11B and TACTICS-TIMI 18 studies.

OBJECTIVES: The goal of this study was to define the prognostic value of serum myoglobin in patients with non-ST-elevation acute coronary syndromes (ACS). BACKGROUND: While myoglobin is useful for the early diagnosis of myocardial infarction (MI), its role in the early risk-stratification of patients with ACS has not been established. METHODS: Myoglobin, creatine kinase-MB subfraction (CK-MB) and troponin I (cTnI) were measured at randomization in 616 patients from the Thrombolysis In Myocardial Ischemia/Infarction (TIMI) 11B study and 1,841 patients from the Treat Angina with Aggrastat and Determine Cost of Therapy with an Invasive or Conservative Therapy-Thrombolysis In Myocardial Ischemia/Infarction (TACTICS-TIMI) 18 study. The risks for death and nonfatal MI through six months of follow-up were compared between patients with and without myoglobin elevation (>110 microg/l) in each study and in a dataset combining all eligible patients from both studies (n = 2,457). RESULTS: In a multivariate model adjusting for baseline characteristics, ST changes and CK-MB and cTnI levels, an elevated baseline myoglobin was associated with increased six-month mortality in TIMI 11B (adjusted odds ratio [OR] 2.9 [95% confidence interval [CI] 1.2 to 7.1]), TACTICS-TIMI 18 (adjusted OR 3.0 [95% CI 1.5 to 5.9]) and the combined dataset (adjusted OR 3.0 [95% CI 1.8 to 5.0]). In contrast, there was no significant association between myoglobin elevation and nonfatal MI (combined dataset adjusted OR 1.55, 95% CI 0.9 to 2.6). In TACTICS-TIMI 18, patients with versus those without myoglobin elevation were more likely to have an occluded culprit artery (28% vs. 10%; p < 0.0001) and visible thrombus (49% vs. 34%; p = 0.006) and less likely to have TIMI 3 flow (53% vs. 68%; p = 0.009). CONCLUSIONS: A serum concentration of myoglobin above the MI detection threshold (>110 microg/l) is associated with an increased risk of six-month mortality, independent of baseline clinical characteristics, electrocardiographic changes and elevation in CK-MB and cTnI. These findings suggest that myoglobin may be a useful addition to cardiac biomarker panels for early risk-stratification in ACS.

Acute Disease↗

Diagnostic and neural analysis of skin cancer (DANAOS). A multicentre study for collection and computer-aided analysis of data from pigmented skin lesions using digital dermoscopy.

BACKGROUND: Early detection of melanomas by means of diverse screening campaigns is an important step towards a reduction in mortality. Computer-aided analysis of digital images obtained by dermoscopy has been reported to be an accurate, practical and time-saving tool for the evaluation of pigmented skin lesions (PSLs). A prototype for the computer-aided diagnosis of PSLs using artificial neural networks (NNs) has recently been developed: diagnostic and neural analysis of skin cancer (DANAOS). OBJECTIVES: To demonstrate the accuracy of PSL diagnosis by the DANAOS expert system, a multicentre study on a diverse multinational population was conducted. METHODS: A calibrated camera system was developed and used to collect images of PSLs in a multicentre study in 13 dermatology centres in nine European countries. The dataset was used to train an NN expert system for the computer-aided diagnosis of melanoma. We analysed different aspects of the data collection and its influence on the performance of the expert system. The NN expert system was trained with a dataset of 2218 dermoscopic images of PSLs. RESULTS: The resulting expert system showed a performance similar to that of dermatologists as published in the literature. The performance depended on the size and quality of the database and its selection. CONCLUSIONS: The need for a large database, the usefulness of multicentre data collection, as well as the benefit of a representative collection of cases from clinical practice, were demonstrated in this trial. Images that were difficult to classify using the NN expert system were not identical to those found difficult to classify by clinicians. We suggest therefore that the combination of clinician and computer may potentially increase the accuracy of PSL diagnosis. This may result in improved detection of melanoma and a reduction in unnecessary excisions.

Adult↗

MArray: analysing single, replicated or reversed microarray experiments.

UNLABELLED: MArray is a Matlab toolbox with a graphical user interface that allows the user to analyse single or paired microarray datasets by direct input of the raw data output file from image analysis packages, such as QuantArray or GenePiX. The application provides simple procedures to manually evaluate the quality of each measurement, multiple approaches to both ratio normalization (simple normalization, intensity dependent normalization) and evaluation of the reproducibility of paired experiments (using the techniques 'simple statistical method' and 'quality control ellipse' and 'significance analysis of microarrays'). Specifically, interactive spot evaluation functions are available in MArray and an online gene information database (NCBI UniGene) is linked. The application may provide a valuable aid in selecting and optimizing experimental procedures, as well as serving as an analytical tool for two-state biological comparisons, such as a study of single-dose activation. It is entirely platform independent, and only requires Matlab installed. AVAILABILITY: http://matrise.uio.no/marray/marray.html

Computer Graphics↗

Evaluation of methods for oligonucleotide array data via quantitative real-time PCR.

BACKGROUND: There are currently many different methods for processing and summarizing probe-level data from Affymetrix oligonucleotide arrays. It is of great interest to validate these methods and identify those that are most effective. There is no single best way to do this validation, and a variety of approaches is needed. Moreover, gene expression data are collected to answer a variety of scientific questions, and the same method may not be best for all questions. Only a handful of validation studies have been done so far, most of which rely on spike-in datasets and focus on the question of detecting differential expression. Here we seek methods that excel at estimating relative expression. We evaluate methods by identifying those that give the strongest linear association between expression measurements by array and the "gold-standard" assay. Quantitative reverse-transcription polymerase chain reaction (qRT-PCR) is generally considered the "gold-standard" assay for measuring gene expression by biologists and is often used to confirm findings from microarray data. Here we use qRT-PCR measurements to validate methods for the components of processing oligo array data: background adjustment, normalization, mismatch adjustment, and probeset summary. An advantage of our approach over spike-in studies is that methods are validated on a real dataset that was collected to address a scientific question. RESULTS: We initially identify three of six popular methods that consistently produced the best agreement between oligo array and RT-PCR data for medium- and high-intensity genes. The three methods are generally known as MAS5, gcRMA, and the dChip mismatch mode. For medium- and high-intensity genes, we identified use of data from mismatch probes (as in MAS5 and dChip mismatch) and a sequence-based method of background adjustment (as in gcRMA) as the most important factors in methods' performances. However, we found poor reliability for methods using mismatch probes for low-intensity genes, which is in agreement with previous studies. CONCLUSION: We advocate use of sequence-based background adjustment in lieu of mismatch adjustment to achieve the best results across the intensity spectrum. No method of normalization or probeset summary showed any consistent advantages.

Algorithms↗

Hybridization interactions between probesets in short oligo microarrays lead to spurious correlations.

BACKGROUND: Microarrays measure the binding of nucleotide sequences to a set of sequence specific probes. This information is combined with annotation specifying the relationship between probes and targets and used to make inferences about transcript- and, ultimately, gene expression. In some situations, a probe is capable of hybridizing to more than one transcript, in others, multiple probes can target a single sequence. These 'multiply targeted' probes can result in non-independence between measured expression levels. RESULTS: An analysis of these relationships for Affymetrix arrays considered both the extent and influence of exact matches between probe and transcript sequences. For the popular HGU133A array, approximately half of the probesets were found to interact in this way. Both real and simulated expression datasets were used to examine how these effects influenced the expression signal. It was found not only to lead to increased signal strength for the affected probesets, but the major effect is to significantly increase their correlation, even in situations when only a single probe from a probeset was involved. By building a network of probe-probeset-transcript relationships, it is possible to identify families of interacting probesets. More than 10% of the families contain members annotated to different genes or even different Unigene clusters. Within a family, a mixture of genuine biological and artefactual correlations can occur. CONCLUSION: Multiple targeting is not only prevalent, but also significant. The ability of probesets to hybridize to more than one gene product can lead to false positives when analysing gene expression. Comprehensive annotation describing multiple targeting is required when interpreting array data.

Artifacts↗

Challenges in measuring nursing home and home health quality: lessons from the First National Healthcare Quality Report.

BACKGROUND: The availability of patient assessment data collected by all Medicare- and Medicaid-certified nursing homes (NHs) (the Minimum Data Set [MDS]) and home health agencies (HHAs) (the Outcome and Assessment Information Set [OASIS]) provides an opportunity to measure quality of care in these settings. OBJECTIVE: The objective of this study was to examine methodologic issues encountered as these datasets are used to report the nation's health care in the National Healthcare Quality Report (NHQR) at national and state levels. FINDINGS: Although the reliability of most data elements from MDS and OASIS are considered acceptable in research studies, mixed evidence exists for the reliability and validity of the quality measures themselves. Detection bias can affect the quality measures, particularly for pain and pressure ulcers. Although risk adjustment is used for all measures, effectiveness varies among measures and methods. Additional quality measures such as patient satisfaction, quality of life, and structural measures would be desirable but will require additional data collection efforts. Although the NH measures represent most NH residents, the HHA measures only apply to Medicare and Medicaid patients served by Medicare-certified agencies. Finally, the absence of clinical benchmarks limits the interpretation of the NHQR HHA and NH measures. CONCLUSIONS: Further developmental work is needed to address many of these issues to improve the usefulness of these quality measures in future NHQR reports.

Activities of Daily Living↗