Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Permutation tests to assess sex differences in omics data.

It is common to sex-stratify analyses of omics data and to report effects as 'sex-specific' when they are significant in only one sex. However, when analysing hundreds or thousands of molecules, this approach will yield many spurious 'sex-specific' effects if not supported by significant interactions. I illustrate this problem using an RNA sequencing dataset showing almost no significant sex by treatment interactions, but where sex-stratified analyses yield hundreds of 'sex-specific' effects of treatment. These 'sex-specific' effects could be spurious or could be real but not show interactions due to low statistical power. To distinguish these possibilities, I describe permutation tests, which provide an intuitive way to determine if a pattern of observations differs from what would be expected due to chance. For this dataset, assigning sex at random often generates more 'sex-specific' effects than the real data, demonstrating that there is little evidence of sex differences. Next, I simulate an RNA sequencing dataset that includes genes modelled to have sex-specific effects of a condition. As expected, analysis of this simulated dataset yields both significant interactions and sex-specific effects in sex-stratified analyses. While stratified analyses detect a higher number of sex-specific effects than the analysis of interactions, they erroneously identify genes not modelled to show sex-specific effects more often than interactions. A permutation test confirms that the number of sex-specific effects observed in the simulated dataset is greater than expected due to chance. Permutation tests can be applied to omics studies of sex differences, simultaneously providing (i) a clear and simple demonstration of the problems of sex-stratified analyses, and (ii) additional evidence of sex-specific effects where these are present. R code is provided for permutations, simulations, and plots to visualize potential sex-specific effects, which can be adapted to other types of data.

Female↗

PSORTdb: a protein subcellular localization database for bacteria.

Information about bacterial subcellular localization (SCL) is important for protein function prediction and identification of suitable drug/vaccine/diagnostic targets. PSORTdb (http://db.psort.org/) is a web-accessible database of SCL for bacteria that contains both information determined through laboratory experimentation and computational predictions. The dataset of experimentally verified information (approximately 2000 proteins) was manually curated by us and represents the largest dataset of its kind. Earlier versions have been used for training SCL predictors, and its incorporation now into this new PSORTdb resource, with its associated additional annotation information and dataset version control, should aid researchers in future development of improved SCL predictors. The second component of this database contains computational analyses of proteins deduced from the most recent NCBI dataset of completely sequenced genomes. Analyses are currently calculated using PSORTb, the most precise automated SCL predictor for bacterial proteins. Both datasets can be accessed through the web using a very flexible text search engine, a data browser, or using BLAST, and the entire database or search results may be downloaded in various formats. Features such as GO ontologies and multiple accession numbers are incorporated to facilitate integration with other bioinformatics resources. PSORTdb is freely available under GNU General Public License.

Bacterial Proteins↗

Is it better to combine predictions?

We have compared the accuracy of the individual protein secondary structure prediction methods: PHD, DSC, NNSSP and Predator against the accuracy obtained by combing the predictions of the methods. A range of ways of combing predictions were tested: voting, biased voting, linear discrimination, neural networks and decision trees. The combined methods that involve 'learning' (the non-voting methods) were trained using a set of 496 non-homologous domains; this dataset was biased as some of the secondary structure prediction methods had used them for training. We used two independent test sets to compare predictions: the first consisted of 17 non-homologous domains from CASP3 (Third Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction); the second set consisted of 405 domains that were selected in the same way as the training set, and were non-homologous to each other and the training set. On both test datasets the most accurate individual method was NNSSP, then PHD, DSC and the least accurate was Predator; however, it was not possible to conclusively show a significant difference between the individual methods. Comparing the accuracy of the single methods with that obtained by combing predictions it was found that it was better to use a combination of predictions. On both test datasets it was possible to obtain a approximately 3% improvement in accuracy by combing predictions. In most cases the combined methods were statistically significantly better (at P = 0.05 on the CASP3 test set, and P = 0.01 on the EBI test set). On the CASP3 test dataset there was no significant difference in accuracy between any of the combined method of prediction: on the EBI test dataset, linear discrimination and neural networks significantly outperformed voting techniques. We conclude that it is better to combine predictions.

Algorithms↗

Detection of obstructive sleep apnea in pediatric subjects using surface lead electrocardiogram features.

STUDY OBJECTIVES: To investigate the feasibility of detecting obstructive sleep apnea (OSA) in children using an automated classification system based on analysis of overnight electrocardiogram (ECG) recordings. DESIGN: Retrospective observational study. SETTING: A pediatric sleep clinic. PARTICIPANTS: Fifty children underwent full overnight polysomnography. INTERVENTION: N/A. MEASUREMENTS AND RESULTS: Expert polysomnography scoring was performed. The datasets were divided into a training set of 25 subjects (11 normal, 14 with OSA) and a withheld test set of 25 subjects (11 normal, 14 with OSA). Features, calculated from the ECG of the 25 training datasets, were empirically chosen to train a modified quadratic discriminant analysis classification system. The selected configuration used a segment length of 60 seconds and processed mean, SD, power spectral density, and serial correlation measures to classify segments as apneic or normal. By combining per-segment classifications and using receiver-operator characteristic analysis, a per-subject classifier was obtained that had a sensitivity of 85.7%, specificity of 90.9%, and accuracy of 88% on the training datasets. The same decision threshold was applied to the withheld datasets and yielded a sensitivity of 85.7%, specificity of 81.8%, and accuracy of 84%. The positive and negative predictive values were 85.7% and 81.8%, respectively, on the test dataset. CONCLUSIONS: The ability to correctly identify 12 out of 14 cases of OSA (with the 2 false negatives arising from subjects with an apnea-hypopnea index less than 10) indicates that the automated apnea classification system outlined may have clinical utility in pediatric patients.

Adult↗

Human muscle gene expression responses to endurance training provide a novel perspective on Duchenne muscular dystrophy.

Global gene expression profiling is used to generate novel insight into a variety of disease states. Such studies yield a bewildering number of data points, making it a challenge to validate which genes specifically contribute to a disease phenotype. Aerobic exercise training represents a plausible model for identification of molecular mechanisms that cause metabolic-related changes in human skeletal muscle. We carried out the first transcriptome-wide characterization of human skeletal muscle responses to 6 wk of supervised aerobic exercise training in 8 sedentary volunteers. Biopsy samples before and after training allowed us to identify approximately 470 differentially regulated genes using the Affymetrix U95 platform (80 individual hybridization steps). Gene ontology analysis indicated that extracellular matrix and calcium binding gene families were most up-regulated after training. An electronic reanalysis of a Duchenne muscular dystrophy (DMD) transcript expression dataset allowed us to identify approximately 90 genes modulated in a nearly identical fashion to that observed in the endurance exercise dataset. Trophoblast noncoding RNA, an interfering RNA species, was the singular exception-being up-regulated by exercise and down-regulated in DMD. The common overlap between gene expression datasets may be explained by enhanced alpha7beta1 integrin signaling, and specific genes in this signaling pathway were up-regulated in both datasets. In contrast to these common features, OXPHOS gene expression is subdued in DMD yet elevated by exercise, indicating that more than one major mechanism must exist in human skeletal muscle to sense activity and therefore regulate gene expression. Exercise training modulated diabetes-related genes, suggesting our dataset may contain additional and novel gene expression changes relevant for the anti-diabetic properties of exercise. In conclusion, gene expression profiling after endurance exercise training identified a range of processes responsible for the physiological remodeling of human skeletal muscle tissue, many of which were similarly regulated in DMD. Furthermore, our analysis demonstrates that numerous genes previously suggested as being important for the DMD disease phenotype may principally reflect compensatory integrin signaling.

Biopsy↗

Adjusting for patient selection suggests the addition of docetaxel to 5-fluorouracil-cisplatin induction therapy may offer survival benefit in squamous cell cancer of the head and neck.

When induction chemotherapy is used in locally advanced squamous cell cancer of the head and neck (SCCHN), patients often receive cisplatin-5-fluorouracil (PF) followed by radical loco-regional therapy. Phase II studies of docetaxel-cisplatin-5-fluorouracil (TPF) induction therapy, with or without leucovorin (L), have achieved high survival rates versus those reported in phase III PF trials. However, the distribution of prognostic factors may vary between phase II and phase III study populations, making the extrapolation of phase II TPF/L results to phase III PF populations difficult. This study used a patient selection standardization method and Cox model to adjust for potential selection bias. Thus, the survival benefit from adding docetaxel into PF induction regimens in SCCHN could be more accurately assessed. The TPF/L dataset comprised 195 patients from six phase II trials. The PF dataset of 585 patients was derived from five large randomized trials included in the Meta-Analysis of Chemotherapy in Head and Neck Cancer (MACH-NC) database. TPF/L and PF datasets differed significantly concerning the distribution of several prognostic factors. Adjusting for these differences, the relative risk of death in the PF versus TPF/L datasets was 1.85 (95% confidence interval 1.37-2.49), corresponding to a 20% 2-year survival benefit (p < 0.0001). Sensitivity analyses confirmed that this improved 2-year survival rate of TPF/L over PF was robust, irrespective of the distribution of studied prognostic factors between treatment datasets. We conclude that this improved survival might be due either to docetaxel's pharmacologic effect or to uncontrolled prognostic factors.

Antineoplastic Combined Chemotherapy Protocols↗

Multiplanar reconstruction in magnetic resonance evaluation of the knee. Comparison with film magnetic resonance interpretation.

RATIONALE AND OBJECTIVES: At many institutions, three-dimensional magnetic resonance imaging (MRI) is routinely used for examination of the knee. Multiplanar reconstruction (MPR) is a method of displaying three-dimensional datasets. The authors assessed the usefulness of MPR for evaluating knee MRI datasets by comparing readers' performance using MPR and conventional film MRIs. METHODS: Eight patients with internal derangement of the knee were studied. All had MRI datasets acquired in the sagittal plane using a three-dimensional gradient-echo fast imaging with steady-state precession (FISP) sequence. Arthroscopic surgery after MRI confirmed the presence of 6 anterior cruciate ligament (ACL) tears, 11 meniscal tears, and 5 normal menisci in this group. Four blinded readers, experienced in MRI of the knee, evaluated the images. The MRI datasets were then loaded onto a three-dimensional workstation and interpreted by the same readers using MPR. The MRI findings were correlated with arthroscopy. RESULTS: For diagnosis of tears of the ACL, sensitivity was 96% and specificity was 100% for both films and MPR. For detecting meniscal tears, sensitivity was 55% and specificity was 90%, using the filmed images, versus 64% and 85%, respectively, with MPR. These differences were not statistically significant by the sign test. Total time (technologist processing time plus radiologist reading time) for MPR was greater than for film interpretation by a factor of 1.12 (P < .05), and radiologist reading time for MPR was greater by a factor of 1.88 (P < .05). CONCLUSIONS: For sagittal three-dimensional FISP MRI datasets, real-time MPR is comparable with film interpretation for evaluation of ACL and meniscal injuries, but it increases the time required for diagnosis.

Adolescent↗

Development and validation of a derived measure of research utilization by nurses.

BACKGROUND: Theoretical models are needed to guide strategies for the implementation of research into clinical practice. To develop and test such models, including analyses of complex theoretical constructs and causal relationships, rich datasets are needed. Working with existing datasets may mean that important variables are lacking. OBJECTIVE: The aim of this study was to derive a nursing research utilization variable and validate it using the Promoting Action on Research Implementation in Health Services (PARIHS) conceptual framework on research implementation. METHODS: This study was based on data from two surveys of registered nurses. The first survey (1996; N = 600) contained robust research utilization variables but few organizational variables. The second (1998; N = 6,526) was rich in organizational variables but contained no research utilization variables. A linear regression model with predictors common to both datasets was used to derive a research utilization variable in the 1998 dataset. To validate these scores, four separate procedures based on the hypothesis of a positive relationship between context and research utilization were completed. Mutually exclusive groups reflecting various levels of context were created to accomplish these procedures. RESULTS: The derived research utilization variable was successfully mapped onto the cases in the 1998 dataset. The derived scores ranged from 0.21 to 21.40, with a mean of 10.85 (SD = 3.23). The mean score per subgroup ranged from 8.28 for the lowest context group to 12.75 for the highest context group. One of the validation procedures showed that significant differences in mean research utilization existed only among four conceptually unique context groups (p < .001). These groups showed a positive incremental relationship in research utilization (p < .001; the better the context, the higher the research utilization score). The validity of the derived variable was supported by using the three remaining validation procedures. DISCUSSION: The successful creation and validation of a derived research utilization variable will enable advanced modeling of the relationships between research utilization and individual and organizational characteristics. The findings also support the construct validity of the context element of the PARIHS theoretical framework.

Adult↗

Multinational study of pneumococcal serotypes causing acute otitis media in children.

BACKGROUND: Streptococcus pneumoniae is a major cause of acute otitis media (AOM) in young children. More than 90 immunologically distinct pneumococcal serotypes have been identified, but limited information is available regarding their relative importance in AOM. METHODS: We analyzed nine existing datasets comprising pneumococcal isolates from middle ear fluid samples collected from 1994 through 2000 from 3,232 children with AOM from Finland, France, Greece, Israel, several East European countries, the US and Argentina. We examined the distribution of pneumococcal serotypes in relation to several demographic and epidemiologic variables, including gender, age, antibiotic resistance and source of culture material. RESULTS: The major serotypes identified included 19F and 23F, each comprising 13 to 25% of pneumococcal middle ear fluid isolates in most datasets; 14 and 6B, comprising 6 to 18%; whereas 6A, 19A and 9V each comprised 5 to 10%. Despite differences in location, study design and antibiotic susceptibility, each major serotype was prominent in most age groups of each dataset. Serotypes represented in the 7-valent pneumococcal conjugate vaccine (PCV-7, 4, 6B, 9V, 14, 18C, 19F, 23F) accounted for 60 to 70% of all pneumococcal isolates in the 6- to 59-month age range, but only 40 to 50% of isolates in children <6 or >/=60 months old. Serotype 3 and, in certain datasets, serotypes 1 and 5, were more important in the <6- and >/=60-month age groups. In each age group vaccine-related serotypes (mainly 6A and 19A) comprised an additional 10 to 15% of all pneumococcal isolates. Four serotypes (23F, 19F, 14 and 6B) accounted for 83% of all penicillin-resistant observations. CONCLUSIONS: This analysis of several geographically diverse datasets indicates that a limited number of serotypes, largely represented in PCV-7, accounted for the majority of episodes of pneumococcal AOM in children between 6 and 59 months of age. Certain serotypes appeared to be relatively more significant in children <6 months or >59 months of age.

Age Factors↗

Phylogeny of the photosynthetic euglenophytes inferred from the nuclear SSU and partial LSU rDNA.

Previous studies using the nuclear SSU rDNA have indicated that the photosynthetic euglenoids are a monophyletic group; however, some of the genera within the photosynthetic lineage are not monophyletic. To test these results further, evolutionary relationships among the photosynthetic genera were investigated by obtaining partial LSU nuclear rDNA sequences. Taxa from each of the external clades of the SSU rDNA-based phylogeny were chosen to create a combined dataset and to compare the individual LSU and SSU rDNA datasets. Conserved areas of the aligned sequences for both the LSU and SSU rDNA were used to generate parsimony, log-det, maximum-likelihood and Bayesian trees. The SSU and LSU rDNA consistently generated the same seven terminal clades; however, the relationship among those clades varied depending on the type of analysis and the dataset used. The combined dataset generated a more robust phylogeny, but the relationships among clades still varied. The addition of the LSU rDNA dataset to the euglenophyte phylogeny supports the view that the genera Euglena, Lepocinclis and Phacus are not monophyletic and substantiates the existence of several well-supported clades. A secondary structural model for the D2 region of the LSU rDNA was proposed on the basis of compensatory base changes found in the alignment.

Animals↗

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids↗

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article↗

Control genes and variability: absence of ubiquitous reference transcripts in diverse mammalian expression studies.

Control genes, commonly defined as genes that are ubiquitously expressed at stable levels in different biological contexts, have been used to standardize quantitative expression studies for more than 25 yr. We analyzed a group of large mammalian microarray datasets including the NCI60 cancer cell line panel, a leukemia tumor panel, and a phorbol ester induction time course as well as human and mouse tissue panels. Twelve housekeeping genes commonly used as controls in classical expression studies (including GAPD, ACTB, B2M, TUBA, G6PD, LDHA, and HPRT) show considerable variability of expression both within and across microarray datasets. Although we can identify genes with lower variability within individual datasets by heuristic filtering, such genes invariably show different expression levels when compared across other microarray datasets. We confirm these results with an analysis of variance in a controlled mouse dataset, showing the extent of variability in gene expression across tissues. The results show the problems inherent in the classical use of control genes in estimating gene expression levels in different mammalian cell contexts, and highlight the importance of controlled study design in the construction of microarray experiments.

Animals↗

Rec-I-DCM3: a fast algorithmic technique for reconstructing large phylogenetic trees.

Phylogenetic trees are commonly reconstructed based on hard optimization problems such as maximum parsimony (MP) and maximum likelihood (ML). Conventional MP heuristics for producing phylogenetic trees produce good solutions within reasonable time on small datasets (up to a few thousand sequences), while ML heuristics are limited to smaller datasets (up to a few hundred sequences). However, since MP (and presumably ML) is NP-hard, such approaches do not scale when applied to large datasets. In this paper, we present a new technique called Recursive-Iterative-DCM3 (Rec-I-DCM3), which belongs to our family of Disk-Covering Methods (DCMs). We tested this new technique on ten large biological datasets ranging from 1,322 to 13,921 sequences and obtained dramatic speedups as well as significant improvements in accuracy (better than 99.99%) in comparison to existing approaches. Thus, high-quality reconstructions can be obtained for datasets at least ten times larger than was previously possible.

Algorithms↗

A support vectors classifier approach to predicting the risk of progression of adolescent idiopathic scoliosis.

A support vector classifier (SVC) approach was employed in predicting the risk of progression of adolescent idiopathic scoliosis (AIS), a condition that causes visible trunk asymmetries. As the aetiology of AIS is unknown, its risk of progression can only be predicted from measured indicators. Previous studies suggest that individual indicators of AIS do not reliably predict its risk of progression. Complex indicators with better predictive values have been developed but are unsuitable for clinical use as obtaining their values is often onerous, involving much skill and repeated measurements taken over time. Based on the hypothesis that combining common indicators of AIS using an SVC approach would produce better prediction results more quickly, we conducted a study using three datasets comprising a total of 44 moderate AIS patients (30 observed, 14 treated with brace). Of the 44 patients, 13 progressed less than 5 degrees and 31 progressed more than 5 degrees. One dataset comprised all the patients. A second dataset comprised all the observed patients and a third comprised all the brace-treated patients. Twenty-one radiographic and clinical indicators were obtained for each patient. The result of testing on the three datasets showed that the system achieved 100% accuracy in training and 65%-80% accuracy in testing. It outperformed a "statistically equivalent" logistic regression model and a stepwise linear regression model on the said datasets. It took less than 20 min per patient to measure the indicators, input their values into the system, and produce the needed results, making the system viable for use in a clinical environment.

Adolescent↗

Segmentation of trabeculated structures using an anisotropic Markov random field: application to the study of the optic nerve head in glaucoma.

The study of the architecture of the optic nerve head (ONH) may provide valuable information about the development and progression of glaucoma. To this end, we have generated three-dimensional datasets from monkey eyes under controlled intraocular pressure (IOP). Segmentation of the connective tissues in this area is crucial to obtain an accurate measurement of geometrical parameters and to build mechanical models. However, this segmentation is made difficult by the complicated geometry and the artifacts introduced in the dataset building process. We present a novel segmentation algorithm, based on expectation-maximization, which incorporates an anisotropic Markov random field (MRF) to introduce prior knowledge about the geometry of the structure. The structure tensor is used to characterize the predominant structure direction and the spatial coherence at each point. The algorithm, which has been validated on an artificial validation dataset that mimics our ONH datasets, shows significant improvement over an isotropic MRF. Results on the real datasets demonstrate the ability of the new algorithm to obtain accurate, spatially consistent segmentations of this structure.

Algorithms↗

Population pharmacokinetics of APOMINE: a meta-analysis in cancer patients and healthy males.

AIMS: 1) To characterize the population pharmacokinetics of apomine in healthy males and in male and female patients with solid tumours and 2) to understand more fully the influence of induction and between- and within-subject variability on exposure to drug using Monte Carlo simulation. METHODS: Apomine was administered once- or twice-daily with or without food in single and multiple oral doses of 30-2100 mg to healthy males (n = 19) and patients with solid tumours (n = 19). The data were divided into model development and validation sets. Models were developed using standard population methods. These were the identification of an appropriate base model, calculation of the empirical Bayes estimates of the primary pharmacokinetic parameters, covariate screening, forward stepwise addition of covariates using the likelihood ratio test as a model selection criteria, and backwards elimination to obtain the final model. To study the influence of data from individual subjects, the model development dataset was subjected to the delete-1 jack-knife and the final model was fitted to each jack-knifed dataset. Principal components analysis of the jack-knifed matrix of model parameters identified two influential subjects who were removed from the dataset, and the final model contained data from the remaining subjects. Model validation was examined using goodness of fit statistics and relative error measures using independent datasets from cancer patients. The model provided a reasonable approximation to the pharmacokinetic measurements in the validation datasets. Computer simulations were undertaken to understand further the pharmacokinetics of apomine in otherwise healthy females, a population not yet studied. RESULTS: Apomine pharmacokinetics were complex and consistent with a two-compartment model with a lag-time. Apparent oral clearance at baseline and apparent volume of distribution at steady-state were larger in healthy males than in cancer patients (41 ml h(-1) and 14.1 l vs 10 ml h(-1) and 8.9 l, respectively, for a 75 kg person). Clearance was time-variant showing a maximal increase with full induction of 320 ml h(-1), independent of patient type. The time to reach 50% maximal induction was about 2 days. The fraction of drug absorbed was relatively constant at doses less than 100-200 mg once daily but decreased at higher doses. Food also decreased relative bioavailability by 36%. Patient characteristics had no effect on apomine pharmacokinetics except for weight, which was proportional to the volume of the central compartment. Between-subject variability (68% for clearance, 30% for central volume, and 141% for peripheral volume) was moderate to large and independent of patient type. Inter-occasion variability was small (18% for both clearance and central volume). Residual variability was modelled with an additive and proportional error model. Cancer patients had slightly higher plasma concentrations than healthy males but this difference was probably not clinically significant. Steady-state was reached in about 3-4 days after once-daily drug administration. The half-life of apomine after three weeks of once-daily dosing was 41 h in cancer patients and 32 h in healthy males. CONCLUSIONS: A population model for apomine has been developed has been developed that characterizes its pharmacokinetics in cancer patients and healthy subjects under a variety of conditions.

Adolescent↗

Cause of death in patients with end-stage renal disease: assessing concordance of death certificates with registry reports.

OBJECTIVES: To assess concordance in reporting, in two Australian national datasets, of cause of death of patients with end-stage renal disease (ESRD). METHODS: For deaths in 1997-99, we compared 'cause of death' and 'primary renal disease', as coded in the Australian and New Zealand Dialysis and Transplant Registry (ANZDATA), with 'underlying' and 'associated' causes of death (based on death certificates), as coded by the Australian Bureau of Statistics (ABS). Dates of birth and death and sex identified the same individuals in the two datasets. Deaths from the three States for which date of birth was not available from death certificates were excluded. Cause of death was compared at the ICD-10 chapter level. RESULTS: Of 1,728 ANZDATA patients from NSW, SA, WA, NT and ACT who died during 1997-99, 1,117 (65%) could be matched to a record in the ABS dataset for the corresponding jurisdictions. The death certificates of 219 (20%) of these 1,117 patients made no mention of chronic renal failure. Overall, agreement on cause of death was poor (kappa = 0.22). Using ANZDATA information on cause of death and ABS underlying cause of death, only 38% of patients had the same cause (at the ICD-10 chapter level) recorded in both datasets. Additional information on primary renal disease (ANZDATA) and up to 12 associated causes of death (ABS) was required to obtain substantial agreement. CONCLUSION AND IMPLICATIONS: Death certificates and ANZDATA records provide differing causes of death for ESRD patients. Information from these sources was not directly comparable. Neither dataset provided a complete picture of renal disease as a cause of death in Australia.

Australia↗