Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Measuring extraocular muscle volume using dynamic contours.

The effect of medical treatment on extraocular muscle enlargement in thyroid associated ophthalmopathy (TAO) may be monitored by measuring the change in volume of the extraocular muscles on serial orbital MRI examinations. In theory, 3D image sets offer the opportunity to minimise errors due to poor repositioning and partial volume effects. This study describes an automated technique for estimating extraocular muscle volumes from 3D datasets. Operator input is minimal and the technique is robust. Verification of the technique on both simulated and real datasets is described. For simulated image sets, both automated segmentation and manual outlining produced estimates of volume which were on average 4% less than "true" volume. For real patient data, extraocular muscle volumes measured by the automated technique were 1.6% (SD 13%) less than volumes measured by manual outlining. Coefficient of variation for repeat outlining of the same image dataset for the automated technique was 1.0%, compared with 4% for manual outlining. The manual technique took an experienced operator approximately 20 min to perform, compared to 7 min for the automated technique. The automated method is therefore rapid, reproducible and at least as accurate as other available methods.

Algorithms↗

Vessel diameter measurements in gadolinium contrast-enhanced three-dimensional MRA of peripheral arteries.

In this study, the possibilities for quantification of vessel diameters of peripheral arteries in gadolinium contrast-enhanced magnetic resonance angiography (Gd CE MRA) were evaluated. Absolute vessel diameter measurements were assessed objectively and semi-automatically in maximum intensity projections (MIPs) of contrast-enhanced T1-weighted 3D spoiled gradient-echo datasets, studied with digital subtraction techniques. In vivo, the complete peripheral arterial bed of six patients was studied, from the aorto-iliac bifurcation down to the distal run-off. By measuring the signal intensity (SI) over the lumen of a vessel in the MIP, an SI-plot was obtained. Next, the vessel boundaries were determined using a threshold algorithm; from these boundary points individual diameter values could be obtained along the trajectory of the vessel. In an in vitro study, an optimal threshold value of 30% of the range of SI-values between the background and the maximal SI in the vessel was obtained for accurate diameter measurement in Gd CE MRA (i.e., full-width 30%-maximum). Furthermore, the relationship between the accuracy of these measurements and the scan resolution was investigated. Accuracy was found to be acceptable (i.e., less than 10% over/underestimation) for vessel sizes covering at least 3 pixels. In six patients, diameters were measured in MIPs of the total datasets (i.e., D(T)) as well as in selective MIPs of the clipped datasets (i.e., D(S)) (n = 209). D(T) and D(S) were statistically significantly correlated (p < 0.01) with a Pearson correlation coefficient rP = 0.98. Measurements in the total MIPs yielded statistically significant (p < 0.01) smaller diameter values compared with measurements in selective MIPs, with a mean difference of 0.15 mm. Diameter values from the selective MIPs of the aorto-iliac arteries were also compared with diameter values measured at corresponding anatomic positions in X-ray angiograms of these patients (i.e., D(x)) (n = 70). D(X) and D(S) were statistically significantly correlated (p < 0.01) with a Pearson correlation coefficient rP = 0.92. Diameters measured in the selective MIPs were smaller than those measured in the X-ray angiograms (mean difference 0.49 mm) and this difference was statistically significant (p < 0.01). In conclusion, diameter values can be evaluated accurately in MIPs of vessels with at least 3 pixels in diameter, using the full-width 30%-maximum criterion.

Adult↗

Influence and correction of temperature perturbations on NIR spectra during the monitoring of a polymorph conversion process prior to self-modelling mixture analysis.

The influence of temperature variations on the rank of a NIR dataset, has been investigated by comparing the results of principal component analysis (PCA) and evolving factor analysis (EFA), applied to two datasets measured at constant temperature and varying temperature. After temperature correction, the concentration profiles and spectra were obtained with PCA, SIMPLISMA and the orthogonal projection approach (OPA). The same resolution methods were used on the dataset measured at constant temperature.

Chemistry, Pharmaceutical↗

The prognostic value of serum myoglobin in patients with non-ST-segment elevation acute coronary syndromes. Results from the TIMI 11B and TACTICS-TIMI 18 studies.

OBJECTIVES: The goal of this study was to define the prognostic value of serum myoglobin in patients with non-ST-elevation acute coronary syndromes (ACS). BACKGROUND: While myoglobin is useful for the early diagnosis of myocardial infarction (MI), its role in the early risk-stratification of patients with ACS has not been established. METHODS: Myoglobin, creatine kinase-MB subfraction (CK-MB) and troponin I (cTnI) were measured at randomization in 616 patients from the Thrombolysis In Myocardial Ischemia/Infarction (TIMI) 11B study and 1,841 patients from the Treat Angina with Aggrastat and Determine Cost of Therapy with an Invasive or Conservative Therapy-Thrombolysis In Myocardial Ischemia/Infarction (TACTICS-TIMI) 18 study. The risks for death and nonfatal MI through six months of follow-up were compared between patients with and without myoglobin elevation (>110 microg/l) in each study and in a dataset combining all eligible patients from both studies (n = 2,457). RESULTS: In a multivariate model adjusting for baseline characteristics, ST changes and CK-MB and cTnI levels, an elevated baseline myoglobin was associated with increased six-month mortality in TIMI 11B (adjusted odds ratio [OR] 2.9 [95% confidence interval [CI] 1.2 to 7.1]), TACTICS-TIMI 18 (adjusted OR 3.0 [95% CI 1.5 to 5.9]) and the combined dataset (adjusted OR 3.0 [95% CI 1.8 to 5.0]). In contrast, there was no significant association between myoglobin elevation and nonfatal MI (combined dataset adjusted OR 1.55, 95% CI 0.9 to 2.6). In TACTICS-TIMI 18, patients with versus those without myoglobin elevation were more likely to have an occluded culprit artery (28% vs. 10%; p < 0.0001) and visible thrombus (49% vs. 34%; p = 0.006) and less likely to have TIMI 3 flow (53% vs. 68%; p = 0.009). CONCLUSIONS: A serum concentration of myoglobin above the MI detection threshold (>110 microg/l) is associated with an increased risk of six-month mortality, independent of baseline clinical characteristics, electrocardiographic changes and elevation in CK-MB and cTnI. These findings suggest that myoglobin may be a useful addition to cardiac biomarker panels for early risk-stratification in ACS.

Acute Disease↗

Exploring regions of interest with cluster analysis (EROICA) using a spectral peak statistic for selecting and testing the significance of fMRI activation time-series.

Much relevant information about activations and artifacts in a functional magnetic resonance imaging (fMRI) dataset can be obtained from an exploratory cluster analysis. In contrast to testing the significance of the measured experimental effect for a given model, unsupervised pattern recognition techniques, such as fuzzy clustering, often find unexpected behavior in addition to expected activations, allowing the exploitation of this element of surprise. The many artifact clusters often discovered might aid the experimenter in deciding whether the dataset is usable, whether some additional preprocessing step is required, or whether the one used has introduced spurious effects. However, clustering alone does not complete the analysis because the membership values that are generated are not indicative of the level of statistical significance with respect to the cluster activation patterns (centroids). This is of particular importance for fMRI datasets for which most time-series are "noise", with no activation patterns. We propose that an initial partition step should precede the clustering step. Only time-series that meet a certain statistical criterion (using a scaled version of Fisher's g-order statistic) are selected for clustering; this typically represents <5% of the whole brain region. The purpose of clustering is to generate a set of cluster centers that are the possible activation patterns; these are used in forming a linear model of all the time-series. The model parameter is tested for significance in both the time and frequency domains. We present a novel method of conducting these tests, which limits the number of false positives. We call the three-step process of initial partition, clustering and the two-domain significance test as exploring regions of interest with cluster analysis (EROICA).

Brain↗

Evolutionary computing for knowledge discovery in medical diagnosis.

One of the major challenges in medical domain is the extraction of comprehensible knowledge from medical diagnosis data. In this paper, a two-phase hybrid evolutionary classification technique is proposed to extract classification rules that can be used in clinical practice for better understanding and prevention of unwanted medical events. In the first phase, a hybrid evolutionary algorithm (EA) is utilized to confine the search space by evolving a pool of good candidate rules, e.g. genetic programming (GP) is applied to evolve nominal attributes for free structured rules and genetic algorithm (GA) is used to optimize the numeric attributes for concise classification rules without the need of discretization. These candidate rules are then used in the second phase to optimize the order and number of rules in the evolution for forming accurate and comprehensible rule sets. The proposed evolutionary classifier (EvoC) is validated upon hepatitis and breast cancer datasets obtained from the UCI machine-learning repository. Simulation results show that the evolutionary classifier produces comprehensible rules and good classification accuracy for the medical datasets. Results obtained from t-tests further justify its robustness and invariance to random partition of datasets.

Adolescent↗

European system for cardiac operative risk evaluation (EuroSCORE).

OBJECTIVE: To construct a scoring system for the prediction of early mortality in cardiac surgical patients in Europe on the basis of objective risk factors. METHODS: The EuroSCORE database was divided into developmental and validation subsets. In the former, risk factors deemed to be objective, credible, obtainable and difficult to falsify were weighted on the basis of regression analysis. An additive score of predicted mortality was constructed. Its calibration and discrimination characteristics were assessed in the validation dataset. Thresholds were defined to distinguish low, moderate and high risk groups. RESULTS: The developmental dataset had 13,302 patients, calibration by Hosmer Lemeshow Chi square was (8) = 8.26 (P < 0.40) and discrimination by area under ROC curve was 0.79. The validation dataset had 1479 patients, calibration Chi square (10) = 7.5, P < 0.68 and the area under the ROC curve was 0.76. The scoring system identified three groups of risk factors with their weights (additive % predicted mortality) in brackets. Patient-related factors were age over 60 (one per 5 years or part thereof), female (1), chronic pulmonary disease (1), extracardiac arteriopathy (2), neurological dysfunction (2), previous cardiac surgery (3), serum creatinine >200 micromol/l (2), active endocarditis (3) and critical preoperative state (3). Cardiac factors were unstable angina on intravenous nitrates (2), reduced left ventricular ejection fraction (30-50%: 1, <30%: 3), recent (<90 days) myocardial infarction (2) and pulmonary systolic pressure >60 mmHg (2). Operation-related factors were emergency (2), other than isolated coronary surgery (2), thoracic aorta surgery (3) and surgery for postinfarct septal rupture (4). The scoring system was then applied to three risk groups. The low risk group (EuroSCORE 1-2) had 4529 patients with 36 deaths (0.8%), 95% confidence limits for observed mortality (0.56-1.10) and for expected mortality (1.27-1.29). The medium risk group (EuroSCORE 3-5) had 5977 patients with 182 deaths (3%), observed mortality (2.62-3.51), predicted (2.90-2.94). The high risk group (EuroSCORE 6 plus) had 4293 patients with 480 deaths (11.2%) observed mortality (10.25-12.16), predicted (10.93-11.54). Overall, there were 698 deaths in 14,799 patients (4.7%), observed mortality (4.37-5.06), predicted (4.72-4.95). CONCLUSION: EuroSCORE is a simple, objective and up-to-date system for assessing heart surgery, soundly based on one of the largest, most complete and accurate databases in European cardiac surgical history. We recommend its widespread use.

Cardiac Surgical Procedures↗

Nuclear corroboration of DNA-DNA hybridization in deep phylogenies of hummingbirds, swifts, and passerines: the phylogenetic utility of ZENK (ii).

This paper documents the phylogenetic utility of ZENK at the avian intra-ordinal level using hummingbirds, swifts, and passerines as case studies. ZENK sequences (1.7 kb) were used to reconstruct separate gene trees containing the major lineages of each group, and the three trees were examined for congruence with existing DNA-DNA hybridization trees. The results indicate both that ZENK is an appropriate nuclear marker for resolving relationships deep in the avian tree, and that many relationships within these three particular groups are congruent among the different datasets. Specifically, within hummingbirds there was topological agreement that the major hummingbird lineages diverged in a graded manner from the "hermits," to the "mangoes," to the "coquettes," to the "emeralds," and finally to a sister relationship between the "mountain-gems" and the "bees." Concerning swifts, the deepest divergences were congruent: treeswifts (Hemiprocnidae) were sister to the typical swifts (Apodidae), and the subfamily Apodinae was monophyletic relative to Cypseloidinae. Within Apodinae, however, were short, unresolved branches among the swiftlets, spinetails, and more typical swifts; a finding which coincides with other datasets. Within passerine birds, there was congruent support for monophyly of sub-oscines and oscines, and within sub-oscines, for monophyly of New World groups relative to the Old World lineages. New World sub-oscines split into superfamilies Furnaroidea and Tyrannoidea, with the Tyrannoid relationships completely congruent among ZENK and DNA-DNA hybridization trees. Within Furnaroidea, however, there was some incongruence regarding the positions of Thamnophilidae and Formicariidae. Concerning oscine passerines, both datasets showed a split between Corvida and Passerida and confirmed the traditional membership of passerid superfamilies Muscicapoidea and Passeroidea. Monophyly of Sylvioidea, however, remained uncertain, as did the relationships among the superfamiles themselves. These results are strikingly similar to other recent findings and indicative of continuing uncertainty about the higher level relationships of oscine passerines.

Animals↗

Decreasing the effects of horizontal gene transfer on bacterial phylogeny: the Escherichia coli case study.

Phylogenetic reconstructions of bacterial species from DNA sequences are hampered by the existence of horizontal gene transfer. One possible way to overcome the confounding influence of such movement of genes is to identify and remove sequences which are responsible for significant character incongruence when compared to a reference dataset free of horizontal transfer (e.g., multilocus enzyme electrophoresis, restriction fragment length polymorphism, or random amplified polymorphic DNA) using the incongruence length difference (ILD) test of Farris et al. [Cladistics 10 (1995) 315]. As obtaining this "whole genome dataset" prior to the reconstruction of a phylogeny is clearly troublesome, we have tested alternative approaches allowing the release from such reference dataset, designed for a species with modest level of horizontal gene transfer, i.e., Escherichia coli. Eleven different genes available or sequenced in this work were studied in a set of 30 E. coli reference (ECOR) strains. Either using ILD to test incongruence between each gene against the all remaining (in this case 10) genes in order to remove sequences responsible for significant incongruence, or using just a simultaneous analysis without removals, gave robust phylogenies with slight topological differences. The use of the ILD test remains a suitable method for estimating the level of horizontal gene transfer in bacterial species. Supertrees also had suitable properties to extract the phylogeny of strains, because the way they summarize taxonomic congruence clearly limits the impact of individual gene transfers on the global topology. Furthermore, this work allowed a significant improvement of the accuracy of the phylogeny within E. coli.

DNA, Bacterial↗

A statewide population-based study of gender differences in trauma: validation of a prior single-institution study.

BACKGROUND: Women usually have lower mortality rates than men do at any age. This pattern is observed for most causes of death from chronic diseases. Significant controversy still exists about gender differences in outcomes in trauma. We previously reported no differences in in-hospital mortality based on gender in a large single-institution study (n= 18,892) that had a significant limitation in that it was not population based. This current study was performed to validate our earlier findings in a separate, statewide, population-based dataset of trauma victims. STUDY DESIGN: Prospective data were collected on 22,332 trauma patients (18,432 blunt, 3,900 penetrating) admitted to all trauma centers (n = 26) in Pennsylvania over 24 months (January 1996 to December 1997). Gender differences in in-hospital mortality were determined for the entire dataset and for the subsets of blunt and penetrating injury patients. A second analysis examined all blunt injury patients and excluded all patients with a hospital length of stay of less than 24 hours, eliminating patients who expired soon after admission. The null hypothesis was that female gender is protective in trauma outcomes. RESULTS: Multiple logistic regression analysis identified age (odds ratio [OR] 1.03, confidence interval [CI] 1.02 to 1.03), Injury Severity Score (OR 1.06, CI 1.05 to 1.06), non-Caucasian race (OR 1.72, CI 1.39 to 2.15), blunt injury type (OR 0.327, CI 0.26 to 0.41), and Revised Trauma Score (OR 0.44, CI 0.41 to 0.47) as independent predictors of in-hospital mortality in trauma. Preexisting diseases, including cardiac disease (OR 1.53, CI 1.12 to 2.09) and malignancy (OR 4.08, CI 1.64 to 10.17), were also identified as independent predictors of in-hospital mortality in trauma. Female gender was not associated with decreased mortality (OR 0.83, CI 0.67 to 1.03, p = 0.093). A second multiple regression analysis in blunt trauma patients admitted for longer than 24 hours (which eliminated early deaths and patients with minor injuries) determined that in-hospital mortality was not significantly different in male or female blunt trauma patients stratified by Injury Severity Score and age. The same factors that were predictive of in-hospital mortality in the total dataset were also significant in this secondary analysis. CONCLUSIONS: These population-based data confirm that female gender does not adversely affect in-hospital mortality in trauma when patients are appropriately stratified for other variables, including Injury Severity Score and age, that do significantly affect outcomes.

Adult↗

A graphic representation of protein sequence and predicting the subcellular locations of prokaryotic proteins.

Zp curve, a three-dimensional space curve representation of protein primary sequence based on the hydrophobicity and charged properties of amino acid residues along the primary sequence is suggested. Relying on the Zp parameters extracted from the three components of the Zp curve and the Bayes discriminant algorithm, the subcellular locations of prokaryotic proteins were predicted. Consequently, an accuracy of 81.5% in the cross-validation test has been achieved using 13 parameters extracted from the curve for the database of 997 prokaryotic proteins. The result is slightly better than that of using the neural network method (80.9%) based on the amino acid composition for the same database. By jointing the amino acid composition and the Zp parameters, the overall predictive accuracy 89.6% can be achieved. It is about 3% higher than that of the Bayes discriminant algorithm based merely on the amino acid composition for the same database. The prediction is also performed with a larger dataset derived from the version 39 SWISS-PROT databank and two datasets with different sequence similarity. Even for the dataset of non-sequence similarity, the improvement can be of 4.4% in the cross-validation test. The results indicate that the Zp parameters are effective in representing the information within a protein primary sequence. The method of extracting information from the primary structure may be useful for other areas of protein studies.

Algorithms↗

Clustering gene expression pattern and extracting relationship in gene network based on artificial neural networks.

Massive datasets such as gene expression profiles are accumulating along with the development of DNA microarray technologies. In this paper, we focus on mining biological relevant information such as typical expression patterns and the interconnections of gene networks from massive datasets. At first, the algorithm of a self-organizing map (SOM) was used to cluster gene expression data. Then, for the typical patterns extracted by the SOM, a three-layer artificial neural network (ANN) model was used to extract the relationships between the expression patterns. In order to evaluate the clustering analysis based on the SOM, biological and statistical indices were introduced. To validate the efficiency of the scheme proposed for extracting the relationships between the expression patterns with the ANN, a test dataset was created and used for the test. Finally, the interconnections of a typical pattern of early G1, late G1, S, G2, and M phases in a yeast cell cycle were extracted and visualized.

Journal Article↗

Computational models for prediction of IVF/ICSI outcomes with surgically retrieved spermatozoa.

IVF/intracytoplasmic sperm injection (ICSI) using surgically retrieved spermatozoa (SRS) is a key option in the treatment of severe male infertility. It was aimed to develop a computational model for the prediction of this modality's outcome. A dataset of 113 exemplars, derived from patients who underwent IVF/ICSI with SRS, was retrospectively analysed. The dataset, containing input features maternal age, sperm retrieval technique, type of spermatozoa used, type of male factor and output intrauterine pregnancy, was randomized into a modelling ('training') set of 83 and cross-validation ('test') set of 30. neUROn++, a set of C++ programs, was used to model the dataset using linear and quadratic discriminant function analysis, logistic regression, and neural computation. A 4-hidden node neural network was found to have the highest accuracy, with a test set receiver operator characteristic (ROC) curve area of 0.783. Reverse regression of this neural network showed maternal age to be the most significant feature in predicting pregnancy (P = 0.025), followed by sperm type (P = 0.076). Type of male factor (P = 0.47) and sperm retrieval technique (P = 0.88) did not predict outcome. In summary, a neural network of clinical relevance was found to be superior in terms of IVF/ICSI outcome prediction. Future media deployment is planned.

Adult↗

Peak capacity of ion mobility mass spectrometry: separation of peptides in helium buffer gas.

Advances in the field of proteomics depend upon the development of high-throughput separation methods. Ion mobility-mass spectrometry is a fast separation method (separations on the millisecond time-scale), which has potential for peptide complex mixture analysis. Possible disadvantages of this technique center around the lack of orthogonality between separation based on ion mobility and separation based on mass. In order to examine the utility of ion mobility-mass spectrometry, the peak capacity (phi) of the technique was estimated by subjecting a large dataset of peptides to linear regression analysis to determine an average trend for tryptic peptides. This trend-line, along with the deviation from a linear relationship observed for this dataset, was used to define the separation space for ion mobility-mass spectrometry. Using the maximum deviation found in the dataset (+/-11%) the peak capacity of ion mobility-mass spectrometry is approximately 2600 peptides. These results are discussed in light of other factors that may increase the peak capacity of ion mobility-mass spectrometry (i.e. multiple trends in the data resulting from multiple classes of compounds present in a sample) and current liquid chromatography approaches to complex peptide mixture analysis.

Helium↗

Identification of ABC transporters in Sarcoptes scabiei.

We have identified and partially sequenced 8 ABC transporters from an EST dataset of Sarcoptes scabiei var. hominis, the causative agent of scabies. Analysis confirmed that most of the known ABC subfamilies are represented in the EST dataset including several members of the multidrug resistance protein subfamily (ABC-C). Although P-glycoprotein (ABC-B) sequences were not found in the EST dataset, a partial P-glycoprotein sequence was subsequently obtained using a degenerate PCR strategy and library screening. Thus a total of 9 potential S. scabiei ABC transporters representing the subfamilies A, B, C, E, F and H have been identified. Ivermectin is currently used in the treatment of hyper-infested (crusted) scabies, and has also been identified as a potentially effective acaricide for mass treatment programmes in scabies-endemic communities. The observation of clinical and in vitro ivermectin resistance in 2 crusted scabies patients who received multiple treatments has raised serious concerns regarding the sustainability of such programmes. One possible mechanism for ivermectin resistance is through ABC transporters such as P-glycoprotein. This work forms an important foundation for further studies to elucidate the potential role of ABC transporters in ivermectin resistance of S. scabiei.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Shared differential factors underlying individual spontaneous neural activity abnormalities in major depressive disorder.

BACKGROUND: In contemporary neuroimaging studies, it has been observed that patients with major depressive disorder (MDD) exhibit aberrant spontaneous neural activity, commonly quantified through the amplitude of low-frequency fluctuations (ALFF). However, the substantial individual heterogeneity among patients poses a challenge to reaching a unified conclusion. METHODS: To address this variability, our study adopts a novel framework to parse individualized ALFF abnormalities. We hypothesize that individualized ALFF abnormalities can be portrayed as a unique linear combination of shared differential factors. Our study involved two large multi-center datasets, comprising 2424 patients with MDD and 2183 healthy controls. In patients, individualized ALFF abnormalities were derived through normative modeling and further deconstructed into differential factors using non-negative matrix factorization. RESULTS: Two positive and two negative factors were identified. These factors were closely linked to clinical characteristics and explained group-level ALFF abnormalities in the two datasets. Moreover, these factors exhibited distinct associations with the distribution of neurotransmitter receptors/transporters, transcriptional profiles of inflammation-related genes, and connectome-informed epicenters, underscoring their neurobiological relevance. Additionally, factor compositions facilitated the identification of four distinct depressive subtypes, each characterized by unique abnormal ALFF patterns and clinical features. Importantly, these findings were successfully replicated in another dataset with different acquisition equipment, protocols, preprocessing strategies, and medication statuses, validating their robustness and generalizability. CONCLUSIONS: This research identifies shared differential factors underlying individual spontaneous neural activity abnormalities in MDD and contributes novel insights into the heterogeneity of spontaneous neural activity abnormalities in MDD.

Humans↗

An efficient projection protocol for chemical databases: singular value decomposition combined with truncated-newton minimization.

A rapid algorithm for visualizing large chemical databases in a low-dimensional space (2D or 3D) is presented as a first step in database analysis and design applications. The projection mapping of the compound database (described as vectors in the high-dimensional space of chemical descriptors) is based on the singular value decomposition (SVD) combined with a minimization procedure implemented with the efficient truncated-Newton program package (TNPACK). Numerical experiments on four chemical datasets with real-valued descriptors (ranging from 58 to 27 255 compounds) show that the SVD/TNPACK projection duo achieves a reasonable accuracy in 2D, varying from 30% to about 100% of pairwise distance segments that lie within 10% of the original distances. The lowest percentages, corresponding to scaled datasets, can be made close to 100% with projections onto a 10-dimensional space. We also show that the SVD/TNPACK duo is efficient for minimizing the distance error objective function (especially for scaled datasets), and that TNPACK is much more efficient than a current popular approach of steepest descent minimization in this application context. Applications of our projection technique to similarity and diversity sampling in drug design can be envisioned.

Database Management Systems↗

Application of validated QSAR models of D1 dopaminergic antagonists for database mining.

Rigorously validated quantitative structure-activity relationship (QSAR) models have been developed for 48 antagonists of the dopamine D1 receptor and applied to mining chemical datasets to discover novel potential antagonists. Several QSAR methods have been employed, including comparative molecular field analysis (CoMFA), simulated annealing-partial least squares (SA-PLS), k-nearest neighbor (kNN), and support vector machines (SVM). With the exception of CoMFA, these approaches employed 2D topological descriptors generated with the MolConnZ software package (EduSoft, LLC. MolconnZ, version 4.05; http://www.eslc.vabiotech.com/ [4.05], 2003). The original dataset was split into training and test sets to allow for external validation of each training set model. The resulting models were characterized by cross-validated R2 (q2) for the training set and predictive R2 values for the test set of (q2/R2) 0.51/0.47 for CoMFA, 0.7/0.76 for kNN, R2 for the training and test sets of 0.74/0.71 for SVM, and training set fitness and test set R2 values of 0.68/0.63 for SA-PLS. Validated QSAR models with R2 > 0.7, (i.e., kNN and SVM) were used to mine three publicly available chemical databases: the National Cancer Institute (NCI) database of ca. 250,000 compounds, the Maybridge Database of ca. 56,000 compounds, and the ChemDiv Database of ca. 450,000 compounds. These searches resulted in only 54 consensus hits (i.e., predicted active by all models); five of them were previously characterized as dopamine D1 ligands, but were not present in the original dataset. A small fraction of the purported D1 ligands did not contain a catechol ring found in all known dopamine full agonist ligands, suggesting that they may be novel structural antagonist leads. This study illustrates that the combined application of predictive QSAR modeling and database mining may provide an important avenue for rational computer-aided drug discovery.

Databases, Factual↗