Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Cross-platform comparison and visualisation of gene expression data using co-inertia analysis.

BACKGROUND: Rapid development of DNA microarray technology has resulted in different laboratories adopting numerous different protocols and technological platforms, which has severely impacted on the comparability of array data. Current cross-platform comparison of microarray gene expression data are usually based on cross-referencing the annotation of each gene transcript represented on the arrays, extracting a list of genes common to all arrays and comparing expression data of this gene subset. Unfortunately, filtering of genes to a subset represented across all arrays often excludes many thousands of genes, because different subsets of genes from the genome are represented on different arrays. We wish to describe the application of a powerful yet simple method for cross-platform comparison of gene expression data. Co-inertia analysis (CIA) is a multivariate method that identifies trends or co-relationships in multiple datasets which contain the same samples. CIA simultaneously finds ordinations (dimension reduction diagrams) from the datasets that are most similar. It does this by finding successive axes from the two datasets with maximum covariance. CIA can be applied to datasets where the number of variables (genes) far exceeds the number of samples (arrays) such is the case with microarray analyses. RESULTS: We illustrate the power of CIA for cross-platform analysis of gene expression data by using it to identify the main common relationships in expression profiles on a panel of 60 tumour cell lines from the National Cancer Institute (NCI) which have been subjected to microarray studies using both Affymetrix and spotted cDNA array technology. The co-ordinates of the CIA projections of the cell lines from each dataset are graphed in a bi-plot and are connected by a line, the length of which indicates the divergence between the two datasets. Thus, CIA provides graphical representation of consensus and divergence between the gene expression profiles from different microarray platforms. Secondly, the genes that define the main trends in the analysis can be easily identified. CONCLUSIONS: CIA is a robust, efficient approach to coupling of gene expression datasets. CIA provides simple graphical representations of the results making it a particularly attractive method for the identification of relationships between large datasets.

Breast Neoplasms↗

Comprehensive curation and analysis of global interaction networks in Saccharomyces cerevisiae.

BACKGROUND: The study of complex biological networks and prediction of gene function has been enabled by high-throughput (HTP) methods for detection of genetic and protein interactions. Sparse coverage in HTP datasets may, however, distort network properties and confound predictions. Although a vast number of well substantiated interactions are recorded in the scientific literature, these data have not yet been distilled into networks that enable system-level inference. RESULTS: We describe here a comprehensive database of genetic and protein interactions, and associated experimental evidence, for the budding yeast Saccharomyces cerevisiae, as manually curated from over 31,793 abstracts and online publications. This literature-curated (LC) dataset contains 33,311 interactions, on the order of all extant HTP datasets combined. Surprisingly, HTP protein-interaction datasets currently achieve only around 14% coverage of the interactions in the literature. The LC network nevertheless shares attributes with HTP networks, including scale-free connectivity and correlations between interactions, abundance, localization, and expression. We find that essential genes or proteins are enriched for interactions with other essential genes or proteins, suggesting that the global network may be functionally unified. This interconnectivity is supported by a substantial overlap of protein and genetic interactions in the LC dataset. We show that the LC dataset considerably improves the predictive power of network-analysis approaches. The full LC dataset is available at the BioGRID (http://www.thebiogrid.org) and SGD (http://www.yeastgenome.org/) databases. CONCLUSION: Comprehensive datasets of biological interactions derived from the primary literature provide critical benchmarks for HTP methods, augment functional prediction, and reveal system-level attributes of biological networks.

Computational Biology↗

Quantitative prediction of imprinting factor of molecularly imprinted polymers by artificial neural network.

Artificial neural network (ANN) implementing the back-propagation algorithm was applied for the calculation of the imprinting factors (IF) of molecularly imprinted polymers (MIP) as a function of the computed molecular descriptors of template and functional monomer molecules and mobile phase descriptors. The dataset used in our study were obtained from the literature and classified into two distinctive datasets on the basis of the polymer's morphology, irregularly sized MIP and uniformly sized MIP datasets. Results revealed that artificial neural network was able to perform well on datasets derived from uniformly sized MIP (n = 23, r = 0.946, RMS = 2.944) while performing poorly on datasets derived from irregularly sized MIP (n = 75, r = 0.382, RMS = 6.123). The superior performance of the uniformly sized MIP dataset over the irregularly sized MIP dataset could be attributed to its more predictable nature owing to the consistency of MIP particles, uniform number and association constant of binding sites, and minimal deviation of the imprinted polymers. The ability to predict the imprinting factor of imprinted polymer prior to performing actual experimental work provide great insights on the feasibility of the interaction between template-functional monomer pairs.

Neural Networks, Computer↗

Faster and more accurate global protein function assignment from protein interaction networks using the MFGO algorithm.

MOTIVATION: Predicting protein function accurately is an important issue in the post-genomic era. To achieve this goal, several approaches have been proposed deduce the function of unclassified proteins through sequence similarity, co-expression profiles, and other information. Among these methods, the global optimization method (GOM) is an interesting and powerful tool that assigns functions to unclassified proteins based on their positions in a physical interactions network [Vazquez, A., Flammini, A., Maritan, A. and Vespignani, A. (2003) Global protein function prediction from protein-protein interaction networks, Nat. Biotechnol., 21, 697-700]. To boost both the accuracy and speed of GOM, a new prediction method, MFGO (modified and faster global optimization) is presented in this paper, which employs local optimal repetition method to reduce calculation time, and takes account of topological structure information to achieve a more accurate prediction. CONCLUSION: On four proteins interaction datasets, including Vazquez dataset, YP dataset, DIP-core dataset, and SPK dataset, MFGO was tested and compared with the popular MR (majority rule) and GOM methods. Experimental results confirm MFGO's improvement on both speed and accuracy. Especially, MFGO method has a distinctive advantage in accurately predicting functions for proteins with few neighbors. Moreover, the robustness of the approach was validated both in a dataset containing a high percentage of unknown proteins and a disturbed dataset through random insertion and deletion. The analysis shows that a moderate amount of misplaced interactions do not preclude a reliable function assignment.

Algorithms↗

Richardson score predicts short-term adverse respiratory outcomes in newborns >/=34 weeks gestation.

OBJECTIVES: To develop a model to predict which newborns >/=34 weeks gestation with respiratory distress will die or will require prolonged (>3 days) assisted ventilation. METHODS: Retrospective cohort study using data from Northern California newborns >/=34 weeks gestation who presented with respiratory distress. We split the cohort into derivation and validation datasets. Bivariate and multivariate data analyses were performed on the derivation dataset. After developing a simple score on the derivation dataset, we applied it to the original as well as to a second validation dataset from Massachusetts. RESULTS: Of 2276 babies who met our initial eligibility criteria, 203 (9.3%) had the primary study outcome (assisted ventilation >3 days or death). A simple score based on gestational age, the lowest PaO 2 /FIO 2 , a variable combining lowest pH and highest PaCO 2 , and the lowest mean arterial blood pressure had excellent performance, with a c-statistic of 0.85 in the derivation dataset, 0.80 in the validation dataset, and 0.80 in the secondary validation dataset. CONCLUSIONS: A simple objective score based on routinely collected physiologic predictors can predict respiratory outcomes in infants >/=34 weeks gestation with respiratory distress.

Continuous Positive Airway Pressure↗

WHO analysis of causes of maternal death: a systematic review.

BACKGROUND: The reduction of maternal deaths is a key international development goal. Evidence-based health policies and programmes aiming to reduce maternal deaths need reliable and valid information. We undertook a systematic review to determine the distribution of causes of maternal deaths. METHODS: We selected datasets using prespecified criteria, and recorded dataset characteristics, methodological features, and causes of maternal deaths. All analyses were restricted to datasets representative of populations. We analysed joint causes of maternal deaths from datasets reporting at least four major causes (haemorrhage, hypertensive disorders, sepsis, abortion, obstructed labour, ectopic pregnancy, embolism). We examined datasets reporting individual causes of death to investigate the heterogeneity due to methodological features and geographical region and the contribution of haemorrhage, hypertensive disorders, abortion, and sepsis as causes of maternal death at the country level. FINDINGS: 34 datasets (35,197 maternal deaths) were included in the primary analysis. We recorded wide regional variation in the causes of maternal deaths. Haemorrhage was the leading cause of death in Africa (point estimate 33.9%, range 13.3-43.6; eight datasets, 4508 deaths) and in Asia (30.8%, 5.9-48.5; 11,16 089). In Latin America and the Caribbean, hypertensive disorders were responsible for the most deaths (25.7%, 7.9-52.4; ten, 11,777). Abortion deaths were the highest in Latin America and the Caribbean (12%), which can be as high as 30% of all deaths in some countries in this region. Deaths due to sepsis were higher in Africa (odds ratio 2.71), Asia (1.91), and Latin America and the Caribbean (2.06) than in developed countries. INTERPRETATION: Haemorrhage and hypertensive disorders are major contributors to maternal deaths in developing countries. These data should inform evidence-based reproductive health-care policies and programmes at regional and national levels. Capacity-strengthening efforts to improve the quality of burden-of-disease studies will further validate future estimates.

Abortion, Induced↗

Similarity searching of chemical databases using atom environment descriptors (MOLPRINT 2D): evaluation of performance.

A molecular similarity searching technique based on atom environments, information-gain-based feature selection, and the naive Bayesian classifier has been applied to a series of diverse datasets and its performance compared to those of alternative searching methods. Atom environments are count vectors of heavy atoms present at a topological distance from each heavy atom of a molecular structure. In this application, using a recently published dataset of more than 100000 molecules from the MDL Drug Data Report database, the atom environment approach appears to outperform fusion of ranking scores as well as binary kernel discrimination, which are both used in combination with Unity fingerprints. Overall retrieval rates among the top 5% of the sorted library are nearly 10% better (more than 14% better in relative numbers) than those of the second best method, Unity fingerprints and binary kernel discrimination. In 10 out of 11 sets of active compounds the combination of atom environments and the naive Bayesian classifier appears to be the superior method, while in the remaining dataset, data fusion and binary kernel discrimination in combination with Unity fingerprints is the method of choice. Binary kernel discrimination in combination with Unity fingerprints generally comes second in performance overall. The difference in performance can largely be attributed to the different molecular descriptors used. Atom environments outperform Unity fingerprints by a large margin if the combination of these descriptors with the Tanimoto coefficient is compared. The naive Bayesian classifier in combination with information-gain-based feature selection and selection of a sensible number of features performs about as well as binary kernel discrimination in experiments where these classification methods are compared. When used on a monoaminooxidase dataset, atom environments and the naive Bayesian classifier perform as well as binary kernel discrimination in the case of a 50/50 split of training and test compounds. In the case of sparse training data, binary kernel discrimination is found to be superior on this particular dataset. On a third dataset, the atom environment descriptor shows higher retrieval rates than other 2D fingerprints tested here when used in combination with the Tanimoto similarity coefficient. Feature selection is shown to be a crucial step in determining the performance of the algorithm. The representation of molecules by atom environments is found to be more effective than Unity fingerprints for the type of biological receptor similarity calculations examined here. Combining information prior to scoring and including information about inactive compounds, as in the Bayesian classifier and binary kernel discrimination, is found to be superior to posterior data fusion (in the datasets tested here).

Journal Article↗

Risk factors for musculoskeletal injuries of the lower limbs in Thoroughbred racehorses in New Zealand.

AIM: To investigate risk factors for injury to musculoskeletal structures of the lower fore- and hind-limbs of Thoroughbred horses training and racing in New Zealand. METHODS: A case-control study analysed by logistic regression was used to compare explanatory variables for musculoskeletal injuries (MSI) in racehorses. The first dataset, termed the Training dataset, involved 459 first-occurrence cases of lower-limb MSI in horses in training, and the second, the Starting dataset, comprised a subset of those horses that had started in at least one trial or race in the training preparation that ended with MSI (n=294). All training preparations for horses that did not suffer from MSI for which complete data were available were used in the analyses as controls, and provided 2,181 and 1,639 preparations for the Training and Starting datasets, respectively. Multivariate logistic regression was used to evaluate risk factors, and results were reported as odds ratios (OR) and 95% confidence intervals (CI). RESULTS: Horses aged > or =5 years were at higher risk of injury than 2-year-olds. Elevated odds of MSI occurred in horses in the Starting dataset that were training in the 1997-1998 year compared with the 1999-2000 year, and in those horses where trials comprised >20% of all starts in a preparation. Training preparations that ended in winter, and horses in their third or later training preparation, had lower odds of MSI compared with those ending in other seasons or the first preparation, respectively. Reduced odds of MSI were observed in preparations in which starts occurred compared with those that had no starts, and in the Starting dataset, preparations that included more than one start had a reduced likelihood of MSI compared with preparations that had only one start. In the Training dataset, preparations longer than 20 weeks were associated with reduced odds of MSI compared with those shorter than 20 weeks. Cumulative racing distance in the last 30 days of a training preparation was best modelled with linear and quadratic terms. Results indicated that increasing cumulative racing distances were associated with an initial reduction in the odds of MSI that then levelled out and finally appeared to increase again as the explanatory variable continued to increase. The risk of MSI varied significantly between trainers. CONCLUSION: This study identified intrinsic (age) and extrinsic risk factors for MSI in training and racing Thoroughbreds in New Zealand. The risk of MSI initially decreased, then increased, as cumulative racing distance increased. Significant variation between trainers indicated management and training methods influence the risk of MSI.

Animals↗

ANOVA-simultaneous component analysis (ASCA): a new tool for analyzing designed metabolomics data.

MOTIVATION: Datasets resulting from metabolomics or metabolic profiling experiments are becoming increasingly complex. Such datasets may contain underlying factors, such as time (time-resolved or longitudinal measurements), doses or combinations thereof. Currently used biostatistics methods do not take the structure of such complex datasets into account. However, incorporating this structure into the data analysis is important for understanding the biological information in these datasets. RESULTS: We describe ASCA, a new method that can deal with complex multivariate datasets containing an underlying experimental design, such as metabolomics datasets. It is a direct generalization of analysis of variance (ANOVA) for univariate data to the multivariate case. The method allows for easy interpretation of the variation induced by the different factors of the design. The method is illustrated with a dataset from a metabolomics experiment with time and dose factors.

Algorithms↗

Prediction of yeast protein-protein interaction network: insights from the Gene Ontology and annotations.

A map of protein-protein interactions provides valuable insight into the cellular function and machinery of a proteome. By measuring the similarity between two Gene Ontology (GO) terms with a relative specificity semantic relation, here, we proposed a new method of reconstructing a yeast protein-protein interaction map that is solely based on the GO annotations. The method was validated using high-quality interaction datasets for its effectiveness. Based on a Z-score analysis, a positive dataset and a negative dataset for protein-protein interactions were derived. Moreover, a gold standard positive (GSP) dataset with the highest level of confidence that covered 78% of the high-quality interaction dataset and a gold standard negative (GSN) dataset with the lowest level of confidence were derived. In addition, we assessed four high-throughput experimental interaction datasets using the positives and the negatives as well as GSPs and GSNs. Our predicted network reconstructed from GSPs consists of 40,753 interactions among 2259 proteins, and forms 16 connected components. We mapped all of the MIPS complexes except for homodimers onto the predicted network. As a result, approximately 35% of complexes were identified interconnected. For seven complexes, we also identified some nonmember proteins that may be functionally related to the complexes concerned. This analysis is expected to provide a new approach for predicting the protein-protein interaction maps from other completely sequenced genomes with high-quality GO-based annotations.

Databases, Genetic↗

Functional cardiac CT and MR: effects of heart rate and software applications on measurement validity.

This study sought to validate different software applications for cardiac function analysis using ECG-gated CT and MR datasets in correlation with underlying heart rate. Ten patients and a set of ventricular phantoms underwent concurrent multislice-CT and cine-MR imaging for evaluation of cardiac function. Datasets from both imaging modalities were evaluated utilizing 2 volumetric analysis tools to determine left ventricular volume and mass. Initially, intraobserver measurement variability was assessed. Detected measurement variability was correlated with underlying absolute magnitude of cardiac volumes and masses. Subsequently, results were statistically evaluated by determining significant data variability depending on imaging modality and choice of evaluation software. Finally, the data variability was correlated with underlying heart rates. This study showed that all analyzed datasets uniformly presented intraobserver variations below 2%, and variability was not related to the magnitude of measurement. Significant measurement accuracy was proven in all calculated parameters obtained from the cardiac phantoms. Acquired patient datasets and calculated functional parameters showed significant data homogeneity, with measurement variability coefficients ranging from 0.935-0.955. CT datasets showed maximal data variability at heart rates below 60 BpM. MR datasets showed maximal data variability at heart rates above 90 BpM. In conclusion, CT and MR datasets allowed an interchangeable utilization of volumetric analysis tools. However, reliable volumetric analysis was limited to an optimal range of cardiac rates for each modality, thus emphasizing the necessity of reporting volumetric measurement results in combination with heart rate to allow for consideration of this possible cause for measurement variation.

Adolescent↗

Functional candidate genes in age-related macular degeneration: significant association with VEGF, VLDLR, and LRP6.

PURPOSE: Age-related macular degeneration (AMD) is a retinal degenerative disease that is the leading cause of blindness worldwide for individuals over the age of 60. Although the etiology of AMD remains largely unknown, numerous studies have suggested that both genes and environmental risk factors significantly influence the risk of developing AMD. Identification of the underlying genes has been difficult, with both genomic screen (locational) and candidate gene (functional) approaches being used. The present study tested candidate genes for association with AMD. METHODS: Eight genes (alpha-2-macroglobulin [A2M], creatine kinase [CKB], angiotensin-converting enzyme [DCP1], interleukin-1alpha [IL1A], low-density lipoprotein receptor-related protein 6 [LRP6], microsomal glutathione-S-transferase 1 [MGST1], vascular entothelial growth factor [VEGF], and very low density lipoprotein receptor [VLDLR]) were tested for genetic linkage and allelic association, using two independent datasets: a family-based association dataset including 162 families and an independent case-control dataset with 399 cases and 159 fully evaluated controls. RESULTS: Test results suggested that genetic variation in five of these genes (IL1A, CKB, A2M, MGST1, and DCP1) is unlikely to explain a significant fraction of the risk of developing AMD in this population. LRP6 showed evidence both for linkage (heterogeneity lod [HLOD] = 1.14) in the family-based dataset and for association (P = 0.004) in the case-control dataset. VEGF showed evidence of linkage (HLOD = 1.32) and demonstrated significant independent allelic association in both the family-based (P = 0.001) and case-control (P = 0.02) datasets. VLDLR showed evidence of association in both the family based (P = 0.03) and case-control (P = 0.01) datasets. CONCLUSIONS: These data suggest that LRP6, VEGF, and VLDLR may play a role in the risk of developing AMD.

Aged↗

Analysis of survival in dairy cows with supplementary data on type scores and housing systems from a region of northwest Germany.

In survival analysis, type traits can be included as covariates to evaluate their use as predictors for survival. One problem in such an analysis is the availability of suitable data. Whereas data on the length of productive life (LPL) of individual cows can be retrieved from milk recording data, for type traits, all cows in the population must be scored for type at least once. In the present analysis, a dataset from the Osnabruck region in northwestern Germany, which fulfilled this requirement in recent years, was used. Data consisted of 169,733 cows with information on LPL for calving years 1980 to 1996 (dataset I) and of 39,233 cows with information on LPL and type for calving years 1990 to 1996 (dataset II). A further dataset (III) contained 43,116 cows from calving years 1987 to 1996 and included information on the housing system for each herd. The basic model included stage of lactation, relative production within herd, change of herd size, and year-season as time dependent effects; age at calving as a time-independent effect; and herd-year-season and sire as random effects. Other effects (information on type, housing system) were included additionally. For data-set II, the scores for 15 linear type traits were also included as corrected phenotypic values, estimated breeding values, and residuals from a previous BLUP analysis. The package Survival Kit 3.0 was used for all analyses. The results indicate a moderate heritability of 0.17 and 0.18 for true and functional LPL (dataset I). Almost all type traits analyzed (dataset II) exceeded a 0.001 level of significance in their effect on survival. The strongest relationships between survival and type were found for udder depth, fore udder attachment, and front teat placement. The main result from the comparison of housing systems (dataset III) was that bedding has a positive effect on survival.

Age Factors↗

Finlay-Wilkinson random regression for yield and yield stability prediction in cereals.

Year-to-year climate variability poses a challenge for agriculture by increasing crop yield variability; therefore, there is a need to identify genotypes that can withstand these fluctuations. With the right selection criteria, genotypes with yield stability across variable environmental conditions can be selected. Methods such as Finlay-Wilkinson random regression (FWRR) may allow us to use sparse datasets-common in plant breeding pipelines-and incorporate genomic data to leverage phenotypic information from related genotypes to predict yield stability. Our objective was to examine how the number of environments and the variance among those environments affect stability predictions. We also integrate FWRR as a genomic prediction tool for characterizing yield stability, comparing it to the traditional genomic prediction models as a reference. We used three datasets: one highly unbalanced dataset for oats (Avena sativa L.) and two completely balanced datasets with different numbers of environments for barley (Hordeum vulgare L.) and wheat (Triticum aestivum L.). We fit standard Finlay-Wilkinson (FW) and FWRR models to estimate grain yield and stability under various scenarios. We found that the estimated stability values obtained were similar using balanced datasets for FW or FWRR. FWRR also achieved moderate predictive ability for stability using unbalanced datasets under 10-fold cross-validation (CV1) with new genotypes. In terms of environmental representation, selecting the right set of environments for inclusion in the model was more important than adding more environments. Our results suggest the possibility of using FWRR to select stable genotypes earlier in line development, as well as to design resource-efficient stability-testing schemes.

Hordeum↗

Experiments to determine whether recursive partitioning (CART) or an artificial neural network overcomes theoretical limitations of Cox proportional hazards regression.

New computationally intensive tools for medical survival analyses include recursive patitioning (also called CART) and artificial neural networks. A challenge that remains is to better understand the behavior of these techniques in effort to know when they will be effective tools. Theoretically they may overcome limitations of the traditional multivariable survival technique, the Cox proportional hazards regression model. Experiments were designed to test whether the new tools would, in practice, overcome these limitations. Two datasets in which theory suggests CART and the neural network should outperform the Cox model were selected. The first was a published leukemia dataset manipulated to have a strong interaction that CART should detect. The second was a published cirrhosis dataset with pronounced nonlinear effects that a neural network should fit. Repeated sampling of 50 training and testing subsets was applied to each technique. The concordance index C was calculated as a measure of predictive accuracy by each technique on the testing dataset. In the interaction dataset, CART outperformed Cox (P < 0.05) with a C improvement of 0.1 (95% CI, 0.08 to 0.12). In the nonlinear dataset, the neural network outperformed the Cox model (P < 0.05), but by a very slight amount (0.015). As predicted by theory, CART and the neural network were able to overcome limitations of the Cox model. Experiments like these are important to increase our understanding of when one of these new techniques will outperform the standard Cox model. Further research is necessary to predict which technique will do best a priori and to assess the magnitude of superiority.

Databases, Factual↗

Empirical analyses of BOLD fMRI statistics. I. Spatially unsmoothed data collected under null-hypothesis conditions.

Temporal autocorrelation, spatial coherency, and their effects on voxel-wise parametric statistics were examined in BOLD fMRI null-hypothesis, or "noise," datasets. Seventeen normal, young subjects were scanned using BOLD fMRI while not performing any time-locked experimental behavior. Temporal autocorrelation in these datasets was described well by a 1/frequency relationship. Voxel-wise statistical analysis of these noise datasets which assumed independence (i.e., ignored temporal autocorrelation) rejected the null hypothesis at a higher rate than specified by the nominal alpha. Temporal smoothing in conjunction with the use of a modified general linear model (Worsley and Friston, 1995, NeuroImage 2: 173-182) brought the false-positive rate closer to the nominal alpha. It was also found that the noise fMRI datasets contain spatially coherent time signals. This observed spatial coherence could not be fully explained by a continuously differentiable spatial autocovariance function and was much greater for lower temporal frequencies. Its presence made voxel-wise test statistics in a given noise dataset dependent, and thus shifted their distributions to the right or left of 0. Inclusion of a "global signal" covariate in the general linear model reduced this dependence and consequently stabilized (i.e., reduced the variance of) dataset false-positive rates.

Artifacts↗

Improving data reliability using a non-compliance detection method versus using pharmacokinetic criteria.

Data from clinical trials present numerous problems for the data analyst. These include non-compliance with the prescribed dosing regimen and inaccurate recollection of dosing history by patients as well as mistakes in recording data. Several methods have been proposed to address these issues. One such technique by Lu et al. (Selecting reliable pharmacokinetic data for explanatory analyses of clinical trials in the presence of possible noncompliance. J. Pharmacokinet. Pharmacodyn. 28:343-362 (2001)) identifies occasions in pharmacokinetic (PK) data where the preceding dosing history is likely to be unreliable. We used this method, implemented in the software program NONMEM (beta) VI, to clean a dataset containing indinavir (IDV) plasma concentrations from HIV-1 infected patients. The data was also cleaned by inspection in Microsoft Excel using clinical PK criteria. A one-compartment model with first order absorption and elimination was fit to both sets of cleaned data. IDV population PK parameters obtained from these analyses were similar to those reported previously. It is established that IDV nephrotoxicity is related to high IDV exposure. However, no relationships were found between any PK parameters and nephrotoxicity in the "compliance cleaned" dataset. In the "PK cleaned" dataset, the oral clearance and apparent volume were lower by 9.1% and 6.6%, respectively in patients with any type of nephrotoxicity and the maximum IDV concentration (C(max)) was 12.1% higher. In patients suffering from nephrolithiasis in particular, C(max) was 15.5% higher. Accordingly, the use of the non-compliance detection method did not improve the reliability of our dataset over the usual method of applying clinical criteria. In fact, analyses on the compliance-cleaned dataset missed some exposure-toxicity relationships. Thus, automated methods must be tested rigorously with 'real life' datasets, used with caution, and always in conjunction with clinical reasoning to avoid overlooking a signal in noisy data.

Adult↗

Reconstruction of large phylogenetic trees: a parallel approach.

Reconstruction of phylogenetic trees for very large datasets is a known example of a computationally hard problem. In this paper, we present a parallel computing model for the widely used Multiple Instruction Multiple Data (MIMD) architecture. Following the idea of divide-and-conquer, our model adapts the recursive-DCM3 decomposition method [Roshan, U., Moret, B.M.E., Williams, T.L., Warnow, T, 2004a. Performance of suptertree methods on various dataset decompositions. In: Binida-Emonds, O.R.P. (Eds.), Phylogenetic Supertrees: Combining Information to Reveal the Tree of Life, vol. 3 of Computational Biology, Kluwer Academics, pp. 301-328; Roshan, U., Moret, B.M.E., Williams, T.L., Warnow, T., 2004b. Rec-I-DCM3: A Fast Algorithmic Technique for reconstructing large phylogenetic trees, Proceedings of the IEEE Computational Systems Bioinformatics Conference (ICSB)] to divide datasets into smaller subproblems. It distributes computation load over multiple processors so that each processor constructs subtrees on each subproblem within a batch in parallel. It finally collects the resulting trees and merges them into a supertree. The proposed model is flexible as far as methods for dividing and merging datasets are concerned. We show that our method greatly reduces the computational time of the sequential version of the program. As a case study, our parallel approach only takes 22.1h on four processors to outperform the best score to date (Found at 123.7h by the Rec-I-DCM3 program [Roshan, U., Moret, B.M.E., Williams, T.L., Warnow, T, 2004a. Performance of suptertree methods on various dataset decompositions. In: Binida-Emonds, O.R.P. (Eds.), Phylogenetic Supertrees: Combining Information to Reveal the Tree of Life, vol. 3 of Computational Biology, Kluwer Academics, pp. 301-328; Roshan, U., Moret, B.M.E., Williams, T.L., Warnow, T., 2004b. Rec-I-DCM3: A Fast Algorithmic Technique for reconstructing large phylogenetic trees, Proceedings of the IEEE Computational Systems Bioinformatics Conference (ICSB)] on one dataset. Developed with the standard message-passing library, MPI, the program can be recompiled and run on any MIMD systems.

Journal Article↗