Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Monte Carlo calculation of the TG-43 dosimetric parameters of a new BEBIG Ir-192 HDR source.

BACKGROUND AND PURPOSE: High dose rate (HDR) brachytherapy is a highly extended practice in clinical brachytherapy today. Quality dose rate distribution datasets of the HDR sources used in a clinical treatment are required. Because of the different source designs, a specific dosimetry dataset is required for each source model. In the recently published BRAPHYQS-ESTRO Report, an overview of available dosimetric data for all HDR Ir-192 sources is given, pointing out the lack of data for one of the sources that is used by the BEBIG MultiSource afterloading system (BEBIG GmbH, Germany). The purpose of this study is to obtain detailed dose rate distributions in liquid water media around this source. MATERIAL AND METHODS: The Monte Carlo code GEANT4 was used to estimate dose rate in water and air-kerma strength around the Ir-192 source. All the details of the stainless steel encapsulated BEBIG HDR 1.1mm in external diameter has been included in the simulation. RESULTS: A complete dosimetric dataset for the BEBIG Ir-192 HDR source is presented. TG43 dosimetric functions and parameters have been obtained as well as a 2D rectangular dose rate table, consistent with the TG43 dose calculation formalism. The dosimetric parameters and functions obtained for the BEBIG HDR source have been compared with that obtained in the literature for others HDR sources, showing that the use of specific datasets for this new source is justified. CONCLUSIONS: This dataset can be used as input in the TPS and to validate its calculations. As policy of BRAPHYQS-ESTRO task group, this dataset will be incorporated to the website: available to users in excel format.

Air↗

A realistic closed-form radiobiological model of clinical tumor-control data incorporating intertumor heterogeneity.

PURPOSE: To investigate the role of intertumor heterogeneity in clinical tumor control datasets and the relationship to in vitro measurements of tumor biopsy samples. Specifically, to develop a modified linear-quadratic (LQ) model incorporating such heterogeneity that it is practical to fit to clinical tumor-control datasets. METHODS AND MATERIALS: We developed a modified version of the linear-quadratic (LQ) model for tumor control, incorporating a (lagged) time factor to allow for tumor cell repopulation. We explicitly took into account the interpatient heterogeneity in clonogen number, radiosensitivity, and repopulation rate. Using this model, we could generate realistic TCP curves using parameter estimates consistent with those reported from in vitro studies, subject to the inclusion of a radiosensitivity (or dose)-modifying factor. We then demonstrated that the model was dominated by the heterogeneity in alpha (tumor radiosensitivity) and derived an approximate simplified model incorporating this heterogeneity. This simplified model is expressible in a compact closed form, which it is practical to fit to clinical datasets. Using two previously analysed datasets, we fit the model using direct maximum-likelihood techniques and obtained parameter estimates that were, again, consistent with the experimental data on the radiosensitivity of primary human tumor cells. This heterogeneity model includes the same number of adjustable parameters as the standard LQ model. RESULTS: The modified model provides parameter estimates that can easily be reconciled with the in vitro measurements. The simplified (approximate) form of the heterogeneity model is a compact, closed-form probit function that can readily be fitted to clinical series by conventional maximum-likelihood methodology. This heterogeneity model provides a slightly better fit to the datasets than the conventional LQ model, with the same numbers of fitted parameters. The parameter estimates of the clinically important time factors and lag periods are very similar to those obtained from the conventional LQ model, but with slightly narrower confidence intervals, reflecting the better fit to the clinical data. DISCUSSION: We have demonstrated, as have others, the importance of intertumor heterogeneity in the response of patient populations to radiotherapy. With the possible inclusion of a radiosensitivity-modifying factor (in vitro/in vivo) of around 1.7, the in vivo data can be made consistent with the in vitro SF2 and Tpot data. Fitting two previously analyzed multicenter datasets indicated that previous analyses based on conventional LQ models gave results for clinically important time factors and lags periods that were not significantly biased by the failure to include intertumor heterogeneity, with slightly narrower confidence intervals, reflecting the better fit to the clinical data. The simple closed-form model we have developed allows direct estimation of the heterogeneity in radiosensitivity within clinical series, and should prove useful in the analysis of other clinical series.

Dose-Response Relationship, Radiation↗

Attrition in longitudinal studies. How to deal with missing data.

The purpose of this paper was to illustrate the influence of missing data on the results of longitudinal statistical analyses [i.e., MANOVA for repeated measurements and Generalised Estimating Equations (GEE)] and to illustrate the influence of using different imputation methods to replace missing data. Besides a complete dataset, four incomplete datasets were considered: two datasets with 10% missing data and two datasets with 25% missing data. In both situations missingness was considered independent and dependent on observed data. Imputation methods were divided into cross-sectional methods (i.e., mean of series, hot deck, and cross-sectional regression) and longitudinal methods (i.e., last value carried forward, longitudinal interpolation, and longitudinal regression). Besides these, also the multiple imputation method was applied and discussed. The analyses were performed on a particular (observational) longitudinal dataset, with particular missing data patterns and imputation methods. The results of this illustration shows that when MANOVA for repeated measurements is used, imputation methods are highly recommendable (because MANOVA as implemented in the software used, uses listwise deletion of cases with a missing value). Applying GEE analysis, imputation methods were not necessary. When imputation methods were used, longitudinal imputation methods were often preferable above cross-sectional imputation methods, in a way that the point estimates and standard errors were closer to the estimates derived from the complete dataset. Furthermore, this study showed that the theoretically more valid multiple imputation method did not lead to different point estimates than the more simple (longitudinal) imputation methods. However, the estimated standard errors appeared to be theoretically more adequate, because they reflect the uncertainty in estimation caused by missing values.

Analysis of Variance↗

Estimation of parameters of dose-volume models and their confidence limits.

Predictions of the normal-tissue complication probability (NTCP) for the ranking of treatment plans are based on fits of dose-volume models to clinical and/or experimental data. In the literature several different fit methods are used. In this work frequently used methods and techniques to fit NTCP models to dose response data for establishing dose-volume effects, are discussed. The techniques are tested for their usability with dose-volume data and NTCP models. Different methods to estimate the confidence intervals of the model parameters are part of this study. From a critical-volume (CV) model with biologically realistic parameters a primary dataset was generated, serving as the reference for this study and describable by the NTCP model. The CV model was fitted to this dataset. From the resulting parameters and the CV model, 1000 secondary datasets were generated by Monte Carlo simulation. All secondary datasets were fitted to obtain 1000 parameter sets of the CV model. Thus the 'real' spread in fit results due to statistical spreading in the data is obtained and has been compared with estimates of the confidence intervals obtained by different methods applied to the primary dataset. The confidence limits of the parameters of one dataset were estimated using the methods, employing the covariance matrix, the jackknife method and directly from the likelihood landscape. These results were compared with the spread of the parameters, obtained from the secondary parameter sets. For the estimation of confidence intervals on NTCP predictions, three methods were tested. Firstly, propagation of errors using the covariance matrix was used. Secondly, the meaning of the width of a bundle of curves that resulted from parameters that were within the one standard deviation region in the likelihood space was investigated. Thirdly, many parameter sets and their likelihood were used to create a likelihood-weighted probability distribution of the NTCP. It is concluded that for the type of dose response data used here, only a full likelihood analysis will produce reliable results. The often-used approximations, such as the usage of the covariance matrix, produce inconsistent confidence limits on both the parameter sets and the resulting NTCP values.

Dose-Response Relationship, Radiation↗

Binary state pattern clustering: a digital paradigm for class and biomarker discovery in gene microarray studies of cancer.

Class and biomarker discovery continue to be among the preeminent goals in gene microarray studies of cancer. We have developed a new data mining technique, which we call Binary State Pattern Clustering (BSPC) that is specifically adapted for these purposes, with cancer and other categorical datasets. BSPC is capable of uncovering statistically significant sample subclasses and associated marker genes in a completely unsupervised manner. This is accomplished through the application of a digital paradigm, where the expression level of each potential marker gene is treated as being representative of its discrete functional state. Multiple genes that divide samples into states along the same boundaries form a kind of gene-cluster that has an associated sample-cluster. BSPC is an extremely fast deterministic algorithm that scales well to large datasets. Here we describe results of its application to three publicly available oligonucleotide microarray datasets. Using an alpha-level of 0.05, clusters reproducing many of the known sample classifications were identified along with associated biomarkers. In addition, a number of simulations were conducted using shuffled versions of each of the original datasets, noise-added datasets, as well as completely artificial datasets. The robustness of BSPC was compared to that of three other publicly available clustering methods: ISIS, CTWC and SAMBA. The simulations demonstrate BSPC's substantially greater noise tolerance and confirm the accuracy of our calculations of statistical significance.

Algorithms↗

The Parkinson's Disease Questionnaire (PDQ-39): evidence for a method of imputing missing data.

BACKGROUND: The Parkinson's Disease Questionnaire (PDQ-39) is the most widely used Parkinson's specific measure of health status. It is increasingly used in treatment trials, sometimes as a primary end-point, where any missing data can potentially cause difficulties in analyses. OBJECTIVES: The purpose of this article is to evaluate the Expectation Maximisation (EM) algorithm for the imputation of missing dimension scores on the 39-item PDQ-39. METHODS: A postal survey of patients diagnosed with Parkinson's disease (PD). A total of 1,372 patients were surveyed and 839 (61.15%) questionnaires returned completed or partially completed. Of these, complete PDQ data were available in 715 (85.22%) cases. Data were deleted from this complete dataset and a sub-set of 200 respondents from this dataset and then imputed using the EM algorithm; results were then compared to the dataset before data deletion. RESULTS: Results gained from imputation of data closely mirrored that of the complete dataset in each case. Descriptive statistics, mean scores and spread of scores were almost identical between original and imputed datasets. Furthermore, original and imputed datasets were highly correlated [intra-class correlation coefficient (ICC) = 0.93 or greater], and mean differences were small (+/-1.00). CONCLUSIONS: The results suggest that the use of EM for the PDQ-39 provides data that closely mirrors the original when this has been deliberately removed. Consequently, EM is likely to be appropriate for trials using the PDQ that contains missing data points.

Adult↗

T-SMmOTE: tweaked synthetic majority minority oversampling technique for data scarcity issue in multi omics studies.

MOTIVATION: Multiomics data offer a rich data mine for modeling complex as well as day-to-day diseases, but their practical deployment is constrained by the limited sample availability. To this end, generating synthetic samples is a viable remedy. Extant schemes operating along this line, however, are mostly limited to augmenting the minority class in imbalanced datasets and often produce synthetic samples that lack sufficient diversity and fail to faithfully capture the underlying data distribution. As a result, the full potential of synthetic augmentation in multi-omics learning remains underexplored. The aim is to address the data scarcity problem in multi-omics domain. We propose a synthetic oversampling framework, which is dedicated to addressing overall data scarcity in multi-omics datasets and the lack of diversity in synthetic samples. Contrary to conventional methods that restrict augmentation to minority classes and rely on interpolation of two neighbors, our method generates diverse yet distribution-aligned synthetic samples by interpolating three neighbors and extends this augmentation paradigm to the majority class. The framework first balances the dataset by generating synthetic minority samples, and subsequently augments the balanced dataset by oversampling both majority and minority classes. RESULTS: Empirical evaluation on multi-omics data obtained from three heterogeneous health scenarios-inflammatory bowel disease, multi-organ dysfunction syndrome, and colorectal cancer-substantiates the utility of the proposed scheme in improving the predictive performance. The models trained on T-SMmOTE-augmented data achieve higher Matthews correlation coefficient values, along with improvedscores for both majority and minority classes. Notably, oversampling of the majority class improves the cognition of the minority class as well. We also explore the consistency of the class distributions between the original and augmented class-specific datasets. These findings confirm the capability of our scheme to learn from small, high-dimensional multi-omics datasets and highlight its potential for non-invasive disease detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/payelu/TSMm.

Journal Article↗

RLBWT-based LCP computation in compressed space for terabase-scale pangenome analysis.

MOTIVATION: Lossless full text indexes are utilized in a myriad of applications in bioinformatics. The continuously decreasing cost of generating biological data has resulted in the need to build full text indexes on biological datasets of increasing size. Many compressed full text indexes have been developed to address this problem. In particular, run-length Burrows-Wheeler transform (RLBWT) based compressed full text indexes have seen wide development and adoption. However, the construction of these RLBWT-based compressed full text indexes is still computationally expensive, sometimes prohibitively so, even for current dataset sizes. RESULTS: Therefore, we present algorithms for the construction of RLBWT-based compressed full text indexes and their supporting data structures in compressed space. The algorithms have a space complexity of O(r) words and run in O(n) time for repetitive datasets, where r is the number of runs in the BWT, n is the length of the text, and repetitive datasets implies nr∈Ω(log n). We provide the first algorithm to compute LCP-related information for repetitive datasets in optimal time and O(r) space, greatly reducing memory requirements. The key idea behind this algorithm is the utilization of r samples of the inverse suffix array at regular intervals. For example, on the Human Pangenome Reference Consortium Release 2 dataset, this reduces peak memory from 2135 GiB to 170 GiB (12.6x reduction) compared to the previous best method (pfp-thresholds). AVAILABILITY AND IMPLEMENTATION: The implementation is available at https://github.com/ucfcbb/TeraTools.

Algorithms↗

Making multi-axis Gaussian graphical models scalable to millions of cells.

MOTIVATION: Networks underlie the generation and interpretation of many biological datasets: gene networks shed light on the regulatory structure of the genome, and cell networks can capture structure of the tumor micro-environment. However, most methods that learn such networks make the faulty "independence assumption"; to learn the gene network, they assume that no cell network exists. "Multi-axis" methods, which do not make this assumption, fail to scale beyond a few thousand cells or genes. This limits their applicability to only the smallest datasets. RESULTS: We develop a multi-axis method, which learns conditional dependency networks, capable of processing million-cell datasets within minutes. This was previously impossible, and unlocks the use of such methods on modern scRNA-seq datasets, as well as more complex datasets. We apply the method to a new scRNA-seq dataset for neuronal cell development, and compare the result to an existing state of the art method, hdWGCNA. We demonstrate that the new method yields gene networks that have a more focused biological interpretation and that the simultaneously learned cell network has advantages over a conventional kNN-based clustering. Further, our method yields novel biological insights by identifying long non-coding RNAs that potentially have a role in neuronal development. AVAILABILITY AND IMPLEMENTATION: Our methodology is available as a Python package GmGM on PyPI (https://pypi.org/project/GmGM/0.5.3/). The code for all experiments performed in this article is available on GitHub (https://github.com/BaileyAndrew/GmGM-Bioinformatics) and Zenodo (10.5281/zenodo.20384566).

Gene Regulatory Networks↗

Selective integration of multiple biological data for supervised network inference.

MOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request.

Algorithms↗

Fast tandem mass spectra-based protein identification regardless of the number of spectra or potential modifications examined.

MOTIVATION: Comparing tandem mass spectra (MSMS) against a known dataset of protein sequences is a common method for identifying unknown proteins; however, the processing of MSMS by current software often limits certain applications, including comprehensive coverage of post-translational modifications, non-specific searches and real-time searches to allow result-dependent instrument control. This problem deserves attention as new mass spectrometers provide the ability for higher throughput and as known protein datasets rapidly grow in size. New software algorithms need to be devised in order to address the performance issues of conventional MSMS protein dataset-based protein identification. METHODS: This paper describes a novel algorithm based on converting a collection of monoisotopic, centroided spectra to a new data structure, named 'peptide finite state machine' (PFSM), which may be used to rapidly search a known dataset of protein sequences, regardless of the number of spectra searched or the number of potential modifications examined. The algorithm is verified using a set of commercially available tryptic digest protein standards analyzed using an ABI 4700 MALDI TOFTOF mass spectrometer, and a free, open source PFSM implementation. It is illustrated that a PFSM can accurately search large collections of spectra against large datasets of protein sequences (e.g. NCBI nr) using a regular desktop PC; however, this paper only details the method for identifying peptide and subsequently protein candidates from a dataset of known protein sequences. The concept of using a PFSM as a peptide pre-screening technique for MSMS-based search engines is validated by using PFSM with Mascot and XTandem. AVAILABILITY: Complete source code, documentation and examples for the reference PFSM implementation are freely available at the Proteome Commons, http://www.proteomecommons.org and source code may be used both commercially and non-commercially as long as the original authors are credited for their work.

Algorithms↗

Epileptic seizures in an Andean region of Ecuador. Incidence and prevalence and regional variation.

A large-scale neuro-epidemiological study was carried out in a population of 72,121 inhabitants of a region of Northern Ecuadorian Andean Sierra, to identify prevalence and incidence rates of epileptic seizures and to identify demographic and geographic variations in these rates. Calculations were made using three datasets. First, rates were calculated from all cases identified in the field (raw dataset); secondly, lower rates were calculated based on a further diagnostic and reclassification procedure (minimum estimated dataset); thirdly, higher rates were derived by calculating false negative rates from the screening procedures, and adding these to the cases actually identified (maximum estimated dataset). Lifetime point-prevalence rates between 12.2/1000 and 19.5/1000 were recorded (minimum and maximum estimated rates), and the prevalence of active epileptic seizures was between 6.7/1000 and 8.0/1000 (minimum estimated and raw datasets). Incidence rate ranging between 122/100,000/year and 190/100,000/year were found (minimum, estimated and raw datasets). A marked difference in prevalence rates was found in two subregions of the survey area, and also in urban and rural areas. The reasons for these differences were not identified.

Adolescent↗

Reanalysis of two studies with contrasting results on the association between statin use and fracture risk: the General Practice Research Database.

BACKGROUND: Two recent case-control studies by Meier et al. and van Staa et al. used the UK General Practice Research Database (GPRD) to examine the association between the use of statins and the risk of fractures, with different results. The objective of the present study was to examine methodological explanations for the discrepant results. METHODS: We created two datasets, which mimicked the previous study designs: a 'selected population' (SP) case-control dataset, with fracture cases matched to controls nested within a selected cohort (Meier et al.), and an 'entire population' (EP) case-control dataset, with both cases and controls sampled from the total GPRD population (van Staa et al.). Cases and controls were matched by gender, age (year of birth or 5 year age bands), and general practice. RESULTS: The study included 131 855 fracture cases. The crude odds ratio (OR) for hip fracture in statin users was 0.37 (95% CI 0.27-0.52) in the SP and 0.54 (95% CI 0.39-0.74) in the EP dataset. This difference was reduced when matching by year of birth, rather than by 5 year age bands: crude ORs were 0.58 (95% CI 0.43-0.79) and 0.61 (95% CI 0.44-0.88), respectively. In the SP dataset, 37% of the cases could be matched by year of birth, while this was achieved for 99% in the 'EP' dataset. The exposure time-window, the selection of confounders, and exclusion of high-risk patients also influenced results. CONCLUSION: Residual confounding by a matching variable and different definitions of the exposure time window explained differences in results. In case-control studies of drug use and fracture risk, broad matching criteria for age should be avoided and the selection of the time-window for exposure should be carefully considered.

Aged↗

Using artificial neural networks to model the urinary excretion of total and purine derivative nitrogen fractions in cows.

A dataset of 177 individual nitrogen balances from dry and lactating cows was split in two independent groups: training dataset (n = 130) and challenge dataset (n = 47). The training dataset was used to develop multiple linear regressions (MLR) and artificial neural networks (ANN) aimed at predicting the urinary excretion of total (NURI) and that of purine derivative nitrogen (PDN). Input variables for the prediction of NURI were crude protein (CP) intake, effective degradability of non-protein dry matter (DM), neutral detergent fiber (NDF) content of the diet, live weight and milk yield. Live weight, total carbohydrate intake, the ratio of non-protein DM degraded to CP degraded and milk yield corrected for DM intake were entered to predict PDN. The regression between predicted and observed values for the training dataset showed a better statistical accuracy of ANN than did MLR models, especially for PDN. The evaluation of the two models on the challenge dataset showed similar determination coefficients, either when predicting total nitrogen excretion (0.623 and 0.614 for ANN and MLR, respectively) or PDN (0.688 and 0.666, for ANN and MLR, respectively). Moreover, both approaches were affected by a tendency to under-predict both targets at high levels of NURI and PDN. However, with the ANN approach, it is possible to study the response of the model to modifications of individual inputs by the so-called response analysis. This unique feature could be used to study the effect of different physiological situations as well as providing hypotheses for additional research.

Animals↗

Pooled analysis of prognostic impact of urokinase-type plasminogen activator and its inhibitor PAI-1 in 8377 breast cancer patients.

BACKGROUND: Urokinase-type plasminogen activator (uPA) and its inhibitor (PAI-1) play essential roles in tumor invasion and metastasis. High levels of both uPA and PAI-1 are associated with poor prognosis in breast cancer patients. To confirm the prognostic value of uPA and PAI-1 in primary breast cancer, we reanalyzed individual patient data provided by members of the European Organization for Research and Treatment of Cancer-Receptor and Biomarker Group (EORTC-RBG). METHODS: The study included 18 datasets involving 8377 breast cancer patients. During follow-up (median 79 months), 35% of the patients relapsed and 27% died. Levels of uPA and PAI-1 in tumor tissue extracts were determined by different immunoassays; values were ranked within each dataset and divided by the number of patients in that dataset to produce fractional ranks that could be compared directly across datasets. Associations of ranks of uPA and PAI-1 levels with relapse-free survival (RFS) and overall survival (OS) were analyzed by Cox multivariable regression analysis stratified by dataset, including the following traditional prognostic variables: age, menopausal status, lymph node status, tumor size, histologic grade, and steroid hormone-receptor status. All P values were two-sided. RESULTS: Apart from lymph node status, high levels of uPA and PAI-1 were the strongest predictors of both poor RFS and poor OS in the analyses of all patients. Moreover, in both lymph node-positive and lymph node-negative patients, higher uPA and PAI-1 values were independently associated with poor RFS and poor OS. For (untreated) lymph node-negative patients in particular, uPA and PAI-1 included together showed strong prognostic ability (all P<.001). CONCLUSIONS: This pooled analysis of the EORTC-RBG datasets confirmed the strong and independent prognostic value of uPA and PAI-1 in primary breast cancer. For patients with lymph node-negative breast cancer, uPA and PAI-1 measurements in primary tumors may be especially useful for designing individualized treatment strategies.

Adult↗

Comparison of UK and US methods for weighting and scoring the SF-36 summary measures.

BACKGROUND: The SF-36 is a widely used measure of health status that can be scored to provide either a profile of eight scores or two summary measures of health, the Physical Component Summary and Mental Component Summary (PCS and MCS). Scoring of the summary scales is undertaken by weighting and summing the original eight dimensions. These weights are gained from factor analysis of data from a general population and have been assumed to be country specific. However, it has been suggested that the weights gained from the US developers could be applied to all datasets, throughout the world, for purposes of comparability and simplicity. The purpose of this study is to evaluate US and UK scoring schemes in a UK population dataset, and in a cohort study of elderly congestive heart failure patients receiving standard therapy and a trial of open vs laparoscopic surgery for hernia repair. METHODS: This paper compares algorithms developed in the USA and the UK for the calculation of the Physical and Mental Health Summary scores (PCS and MCS) for the SF-36 health status measure. In this study the PCS and MCS were calculated using a weighting scheme recommended by the developers and derived from a US population sample dataset, as well as being calculated from weights derived from an UK population sample dataset. RESULTS: The two methods produced similar results, both cross-sectionally and in the assessment of change. CONCLUSIONS: It is suggested that it may be necessary to weight the PCS and MCS using only the original US algorithms, which will lead to more uniform analysis of datasets and may also lead to greater uptake of the summary measures. Furthermore, the results would suggest that in international trials the SF-36 can be adopted and summary scores calculated for countries where no large-scale normative dataset is available. However, further research is needed to determine that the similarity of results gained using UK and US algorithms is not an idiosyncratic feature of the UK data. Studies to verify the findings reported here are required from other countries.

Aged↗

A comparison of ratio distributions based on the NOAEL and the benchmark approach for subchronic-to-chronic extrapolation.

One approach to derive a data-based assessment factor (AF) for subchronic-to-chronic extrapolation is to determine ratios between the NOAEL(subchronic) and NOAEL(chronic) for the same compounds. Instead of using ratios of NOAELs, the distribution can also be estimated by ratios of subchronic and chronic Benchmark Doses (or Critical Effect Doses, CEDs, for continuous data). In this study 314 dose-response datasets on body weights and liver weights of mice and rats were selected providing dose-response information after both subchronic and chronic exposure. NOAEL ratios could be derived in only 68 of these datasets, while CED ratios could be derived in 189 datasets. When only the (53) datasets suitable for both approaches were evaluated the variation of the CED ratio distribution (GSD [geometric standard deviation]: 2.9) was smaller than the one of the NOAEL ratio distribution (GSD: 3.3). After correcting for the estimation error of the individual CED ratios the GSD of the CED distribution decreased to 2.3. The geometric means (GMs) of the NOAEL and CED distributions were similar (1.2 and 1.6, respectively). Comparing the NOAEL distribution based on all 68 datasets suitable for deriving NOAEL ratios with the CED distribution based on the 189 ratios suitable for deriving CED ratios resulted in similar GMs (1.5 and 1.7, respectively), but the GSDs differed considerably (5.3 and 2.3 respectively). It is concluded that usage of the CED approach results in less wide distributions. Furthermore, a larger fraction of available datasets is useful to inform the ratio distribution. This results in more accurate, and less conservative distributions of AFs in general compared to the distributions based on NOAEL ratios that have been proposed so far.

Algorithms↗

Evaluation of the timeliness and completeness of a Web-based notifiable disease reporting system by a local health department.

OBJECTIVE: To evaluate the completeness and timeliness of the Colorado statewide Web-based system for reporting notifiable diseases, called the Colorado Electronic Disease Reporting System. This project demonstrates how a local health department can conduct a surveillance evaluation to identify areas of improvement. METHODS: Reports received by Colorado for 2004 were categorized as Tri-County Health Department (TCHD) reports and reports received for the rest of Colorado. Report completeness and timeliness were compared for all diseases routinely followed up by TCHD for both datasets. A data field was considered complete if there was data entry for that field. Timeliness in this study was defined as the interval between "specimen collection date" and "report date" for each record. RESULTS: Six of 12 selected data fields were 95% or more complete for both datasets. Twenty-four-hour notifiable diseases were reported a median of 2.0 days for reports in the TCHD dataset and a median of 3.0 days for reports in the dataset for the rest of Colorado. Seven-day notifiable diseases were reported a median of 4.0 days for both datasets. CONCLUSIONS: Both Colorado datasets were found to be relatively complete and timely. Improved data collection by interviewers will help better determine demographic information of reported cases and timeliness of reports.

Colorado↗