Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Knowledge management in healthcare: towards 'knowledge-driven' decision-support services.

In this paper, we highlight the involvement of Knowledge Management in a healthcare enterprise. We argue that the 'knowledge quotient' of a healthcare enterprise can be enhanced by procuring diverse facets of knowledge from the seemingly placid healthcare data repositories, and subsequently operationalising the procured knowledge to derive a suite of Strategic Healthcare Decision-Support Services that can impact strategic decision-making, planning and management of the healthcare enterprise. In this paper, we firstly present a reference Knowledge Management environment-a Healthcare Enterprise Memory-with the functionality to acquire, share and operationalise the various modalities of healthcare knowledge. Next, we present the functional and architectural specification of a Strategic Healthcare Decision-Support Services Info-structure, which effectuates a synergy between knowledge procurement (vis-à-vis Data Mining) and knowledge operationalisation (vis-à-vis Knowledge Management) techniques to generate a suite of strategic knowledge-driven decision-support services. In conclusion, we argue that the proposed Healthcare Enterprise Memory is an attempt to rethink the possible sources of leverage to improve healthcare delivery, hereby providing a valuable strategic planning and management resource to healthcare policy makers.

Clinical Competence↗

Decision support in medicine: lessons from the HELP system.

PURPOSE: This report describes an ongoing transition from the HELP Hospital Information System to HELP II, a replacement Health Information System built to manage clinical information captured in a variety of medical settings. The focus of the article is on the medical decision support provided by this system and studied by researchers at the University of Utah and Intermountain Health Care (IHC), a large health care organization in Utah, for many years. METHODS: Select success features of the original HELP system's decision support environment are identified and lessons learned are related. Plans for transferring these features to HELP II are discussed. RESULTS: The article focuses on four features: (1) the importance of easy access to patient data essential for decision support, (2) the commitment to continued measurement and revision of both the logic and the interventional strategy in a decision support application, (3) experience with data mining as a tool for developing decision support tools, and (4) the role of clinical reports in supporting the decision making process.

Decision Support Systems, Clinical↗

NewYork-Presbyterian Hospital: translating innovation into practice.

BACKGROUND: NewYork-Presbyterian (NYP) Hospital, a 2,242-bed not-for-profit academic medical center, was formed by a merger of The New York Hospital and The Presbyterian Hospital in the City of New York. It is also the flagship for the NewYork-Presbyterian Healthcare System, with 37 acute care facilities and 18 others. OVERALL APPROACH TO QUALITY AND SAFETY: The hospital embeds safety in the culture through strategic initiatives and enhances service and efficiency using Six Sigma and other techniques to drive adoption of improvements. Goals are selected in alignment with the annual strategic initiatives, which are chosen on the basis of satisfaction surveys, patient and family complaints, community advisory groups, and performance measures, among other sources. USE OF INFORMATION TO SET AND EVALUATE QUALITY GOALS AND PRIORITIZE INITIATIVES: A new business intelligence system enables online, dynamic analysis of performance results, replacing static paper reports. Advanced features in the clinical information systems include computerized physician order entry; interactive clinical alerts for decision support; a real-time infection control tracking system; and a clinical data warehouse supporting data mining and analysis for quality improvement, decision making, and education. APPROACH TO ADDRESSING THE SIX IOM QUALITY AIMS: To achieve clinical, service, and operational excellence, NYP focuses on all Institute of Medicine quality aims.

Hospital Bed Capacity, 500 and over↗

High-throughput evaluation of olefin copolymer composition by means of attenuated total reflection Fourier tranform infrared spectroscopy.

As a consequence of developing fully automated reactors for organic and organometallic synthesis and polymerizations combined with rapid on-line analysis, databases, and data mining, the analysis of polymers with respect to composition and properties has been speeded up. High-throughput evaluation of olefin copolymers requires fast measurements and high accuracy without tedious sample preparation such as pressing KBr pellets. This has been achieved by using attenuated total reflection Fourier transform infrared spectroscopy (ATR-FTIR spectroscopy) in conjunction with multivariate calibration in order to determine the composition of olefin copolymers such as ethene/propene, ethene/1-hexene and ethene/1-octene copolymers.

Journal Article↗

High-throughput screening and optimization of photoembossed relief structures.

A methodology for the rapid design, screening, and optimization of coating systems with surface relief structures, using a combination of statistical experimental design, high-throughput experimentation, data mining, and graphical and mathematical optimization routines was developed. The methodology was applied to photopolymers used in photoembossing applications. A library of 72 films was prepared by dispensing a given amount of sample onto a chemically patterned substrate consisting of hydrophilic areas separated by fluorinated hydrophobic barriers. Film composition and film processing conditions were determined using statistical experimental design. The surface topology of the films was characterized by automated AFM. Subsequently, models explaining the dependence of surface topologies on sample composition and processing parameters were developed and used for screening a virtual 4000-membered in silico library of photopolymer lacquers. Simple graphical optimization or Pareto algorithms were subsequently used to find an ensemble of formulations, which were optimal with respect to a predefined set of properties, such as aspect ratio and shape of the relief structures.

Automation↗

Structural analysis of transition metal beta-X substituent interactions. Toward the use of soft computing methods for catalyst modeling

Fuzzy logic and neural network techniques are used to classify intramolecular interactions between transition metals (M) and beta-X substituents in the following structural motif (LnMC(alpha)(A1)(A2)-C(beta)(B1)(B2)X). These interactions are relevant to the direct polymerization of functionalized olefins by Ziegler-Natta (ZN) catalysis. The efficiency and effectiveness of different soft computing techniques are compared. These methods give not only encouraging results with respect to general data mining issues but also insight into the factors that effect interactions between transition metals and beta-X substituents.

Journal Article↗

SIRS-SS: a system for simulating IR/Raman spectra. 1. Substructure/subspectrum correlation.

An IR/RAMAN spectra simulation system is reported. The development of this software was based on the substructure/subspectrum relationships established for four different structural classes: small molecules, special fragments, atom-centered FRELs, and bond-centered FRELs (FREL: Fragment centered on an Environment which is Limited). Four corresponding knowledge-bases (now, at a pilot stage) are constructed from usual correlation charts or data analyses of large populations of compounds using data mining techniques.

Journal Article↗

Analysis of calcium, oxalate, and citrate interaction in idiopathic calcium urolithiasis in children.

The majority of urinary stones in children are composed of calcium oxalate. To investigate the interaction between urinary calcium, oxalate, and citrate as major risk factors for calcium stones formation, their 24-h urinary excretion was determined in 30 children with urolithiasis and 15 normal healthy children. The cutoff points between children with urolithiasis and healthy children, accuracy, sensitivity, and specificity for each risk factor alone as well as for all three taken together were determined. OneR and J4.8 classifiers as parts of the larger data mining software Weka, based on machine learning algorithms, were used for the determination of the cutoff points for differentiation of the children. The decision tree based on J4.8 classifier analysis of all three risk factors together proved to be the best for differentiating stone formers from normal children. In comparison to the accuracy of the differentiation after calcium and oxalate of 80% and 75.6%, respectively, the decision tree showed an accuracy of 97.8%. Even when its stability was tested by the leave-one-out cross-validation procedure, the accuracy remained at a very acceptable percentage of 93.2% correctly classified patients. J4.8 classifier analysis gave a look inside urinary calcium, oxalate, and citrate interaction. Urinary calcium excretion was shown as the most informative in discrimination of the children with urolithiasis from healthy children. However, it was shown that oxalate and citrate excretions might influence the stone formation in a subpopulation of the stone formers. In patients with low urinary calcium, a major role in lithogenesis belongs to oxalate, in some of them alone and in others in conjunction with citrate. Decreased urinary citrate excretion in the presence of increased oxalate excretion may lead to stone formation.

Adolescent↗

A wavelet, fourier, and PCA data analysis pipeline: application to distinguishing mixtures of liquids.

Using a new optical engineering technique for the "fingerprinting" of beverages and other liquids, we study and evaluate a range of features. The features are based on resolution scale, invariant frequency information, entropy, and energy. They allow mixtures of beverages to be very precisely placed in principal component plots used for the data analysis. To show this we make use of data sets resulting from optical/near-infrared and ultrasound sensors. Our liquid "fingerprinting" is a relatively open analysis framework in order to cater for different practical applications, in particular, on one hand, discrimination and best fit between fingerprints, and, on the other hand, more exploratory and open-ended data mining.

Journal Article↗

Calculating similarities between biological activities in the MDL Drug Data Report database.

There are a number of licensed databases that assign biological activities to druglike compounds. The MDL Drug Data Report (MDDR), compiled from the patent literature, is a popular example. It contains several hundred distinct activities, some of which are therapeutic areas (e.g., Antihypertensive) and some of which are related to specific enzymes or receptors (e.g., ACE inhibitor). There are several data mining applications where it would be useful to calculate a similarity between any two activities. Two distinct activity labels can have a significant similarity for a number of reasons: two activities can be nearly synonymous (e.g., CCK B antagonist vs Gastrin antagonist), one activity may be a subset of another (e.g., Dopamine (D2) agonist vs Dopamine agonist), or an activity can be the mechanism by which another activity works (e.g., ACE inhibitor vs Antihypertensive), etc. In an ideal world, similarities for two activities could be calculated simply by comparing the compounds they have in common, but in hand-curated databases such as the MDDR the assignment of activities to compounds are inevitably inconsistent and incomplete. We propose a number of methods of calculating activity-activity similarities that hopefully compensate for errors in hand-curation. Two of these, TIMI and trend vector, show promise. Soft clustering of the activities using a union of similarity methods shows a reasonable association of therapeutic areas with their mechanisms.

Algorithms↗

Distribution-based descriptors of the molecular shape.

A rational design of economically cost-effective chemical libraries as well as successful data mining during a process of drug discovery employs a vast array of the molecular descriptors. Despite the huge importance of this area of the research there is still a need for the further development of the simple, intuitive, easily calculable, specific and size-invariant parameters of the molecular shape. Here we present ab initio calculation of the molecular volumes and expectations for the molecular areas of projection. These molecular size parameters were used as a basis to define a group of novel descriptors of the molecular shape. A set of molecular descriptors was developed: ovality, roughness, size-corrected parameters, and the parameters derived from the higher central momenta of the distributions of the size descriptors-skewness and kurtosis. The rationale for the construction of the descriptors was first to calculate the descriptors of the molecular size along many directions in the space and second to use the statistical parameters of the distribution of those descriptors as the shape descriptors. The size descriptors well suited for the above purpose and discussed in this paper are generalized molecular radii and the molecular areas of projection. Molecular volume and projection area-derived descriptors were calculated and their applicability as shape descriptors was illustrated using exploratory methods (factor analysis and hierarchical cluster analysis). The shape descriptors appear to be promising in their ability to discriminate and classify the molecular shapes e.g. spheroids, disklike, rodlike, starlike, crosslike, anglelike, etc.

Journal Article↗

In silico prediction of buffer solubility based on quantum-mechanical and HQSAR- and topology-based descriptors.

We present an artificial neural network (ANN) model for the prediction of solubility of organic compounds in buffer at pH 6.5, thus mimicking the medium in the human gastrointestinal tract. The model was derived from consistently performed solubility measurements of about 5000 compounds. Semiempirical VAMP/AM1 quantum-chemical wave function derived, HQSAR-derived logP, and topology-based descriptors were employed after preselection of significant contributors by statistical and data mining approaches. Ten ANNs were trained each with 90% as a training set and 10% as a test set, and deterministic analysis of prediction quality was used in an iterative manner to optimize ANN architecture and descriptor space, based on Corina 3D molecular structure and AM1/COSMO single point wave function. In production mode, a mean prediction value of the 10 ANNs is created, as is a standard deviation based quality parameter. The productive ANN based on Corina geometries and AM1/COSMO wave function gives an r2cv of 0.50 and a root-mean-square error of 0.71 log units, with 87 and 96% of the compounds having an error of less than 1 and 1.5 log units, respectively. The model is able to predict permanently charged species, e.g. zwitterions or quaternary amines, and problematic structures such as tautomers and unresolved diastereomers almost as well as neutral compounds.

Buffers↗

Generation of a focused set of GSK compounds biased toward ligand-gated ion-channel ligands.

A "data mining" methodology based on substructural analysis and standard 1024 Daylight fingerprints as descriptors was applied to a set of known antagonists of a subfamily of ligand-gated ion channels comprising nicotinic acetylcholine receptors (nAChR's), 5-hydroxytryptamine, gamma-amino butyric acid-A, and glycine receptors. The derived scoring function was used to generate a focused set that was screened for alpha7 nAChR, resulting in the identification of novel alpha7 ligands easily amenable to chemical modification. Finally, the same scoring function was applied retrospectively to other in-house sets screened for the same target in the same assay. The results and performance of the method are described in detail.

Algorithms↗

A searchable database for comparing protein-ligand binding sites for the analysis of structure-function relationships.

The rapid expansion of structural information for protein-ligand binding sites is potentially an important source of information in structure-based drug design and in understanding ligand cross reactivity and toxicity. We have developed a large database of ligand binding sites extracted automatically from the Protein Data Bank. This has been combined with a method for calculating binding site similarity based on geometric hashing to create a relational database for the retrieval of site similarity and binding site superposition. It contains an all-against-all comparison of binding sites and holds known protein-ligand binding sites, which are made accessible to data mining. Here we demonstrate its utility in two structure-based applications: in determining site similarity and in aiding the derivation of a receptor-based pharmacophore model. The database is available from http://www.bioinformatics.leeds.ac.uk/sb/.

Binding Sites↗

SMIREP: predicting chemical activity from SMILES.

Most approaches to structure-activity-relationship (SAR) prediction proceed in two steps. In the first step, a typically large set of fingerprints, or fragments of interest, is constructed (either by hand or by some recent data mining techniques). In the second step, machine learning techniques are applied to obtain a predictive model. The result is often not only a highly accurate but also hard to interpret model. In this paper, we demonstrate the capabilities of a novel SAR algorithm, SMIREP, which tightly integrates the fragment and model generation steps and which yields simple models in the form of a small set of IF-THEN rules. These rules contain SMILES fragments, which are easy to understand to the computational chemist. SMIREP combines ideas from the well-known IREP rule learner with a novel fragmentation algorithm for SMILES strings. SMIREP has been evaluated on three problems: the prediction of binding activities for the estrogen receptor (Environmental Protection Agency's (EPA's) Distributed Structure-Searchable Toxicity (DSSTox) National Center for Toxicological Research estrogen receptor (NCTRER) Database), the prediction of mutagenicity using the carcinogenic potency database (CPDB), and the prediction of biodegradability on a subset of the Environmental Fate Database (EFDB). In these applications, SMIREP has the advantage of producing easily interpretable rules while having predictive accuracies that are comparable to those of alternative state-of-the-art techniques.

Algorithms↗

Rational design and HT techniques allow the synthesis of new IWR zeolite polymorphs.

By combining a rational design of structure directing agents and high throughput and data mining techniques, it has been possible to obtain Ge-free ITQ-24 as pure silica as well as borosilicate polymorphs up to a Si/TIII ratio of 10. Al can be exchanged by B giving strong acid materials. Also, the availability of the pure silica material allows one to determine more precisely the crystal symmetry. This work opens the possibility to synthesize the Ge-free polymorphs of a large number of new germanosilicate structures reported in the last five years. This certainly will increase their possibilities for industrial application.

Germanium↗

Dimerization of G-protein-coupled receptors.

The evolutionary trace (ET) method, a data mining approach for determining significant levels of amino acid conservation, has been applied to over 700 aligned G-protein-coupled receptor (GPCR) sequences. The method predicted the occurrence of functionally important clusters of residues on the external faces of helices 5 and 6 for each family or subfamily of receptors; similar clusters were observed on helices 2 and 3. The probability that these clusters are not random was determined using Monte Carlo techniques. The cluster on helices 5 and 6 is consistent with both 5,6-contact and 5,6-domain swapped dimer formation; the possible equivalence of these two types of dimer is discussed because this relates to activation by homo- and heterodimers. The observation of a functionally important cluster of residues on helices 2 and 3 is novel, and some possible interpretations are given, including heterodimerization and oligomerization. The application of the evolutionary trace method to 113 aligned G-protein sequences resulted in the identification of two functional sites. One large, well-defined site is clearly identified with adenyl cyclase, beta/gamma and regulator of G-protein signaling (RGS) binding. The other G-protein functional site, which extends from the ras-like domain onto the helical domain, has the correct size and electrostatic properties for GPCR dimer binding. The implications of these results are discussed in terms of the conformational changes required in the G-protein for activation by a receptor dimer. Further, the implications of GPCR dimerization for medicinal chemistry are discussed in the context of these ET results.

Amino Acid Sequence↗

Prediction of aqueous solubility of a diverse set of compounds using quantitative structure-property relationships.

"Fail early and fail fast" is the current paradigm that the pharmaceutical industry has adopted widely. Removing non-drug-like compounds from the drug discovery lifecycle in the early stages can lead to tremendous savings of resources. Thus, fast screening methods are needed to profile the large collection of synthesized and virtual libraries involved in the early stage. Solubility is one of the filters that are applied extensively to ensure that the compounds are reasonably soluble so that synthesis of the compounds and assay studies of pharmacokinetics and toxicity are feasible. To address this need, we have developed a fast quantitative structure-property relationship (QSPR) model for the prediction of aqueous solubility (at 298 K, unbuffered solution) from the molecular structures. Multiple linear regressions and genetic algorithms were used to develop the models. The model was based on a set of diverse compounds including small organic molecules and drug and drug-like species. The predicted solubility for the training and test sets agrees well with the experimental values. The coefficient of determination is R(2) = 0.84 for the training set of 775 compounds and the RMS error = 0.87. This model was validated on four sets of compounds. The RMS error for the 1665 compounds from the four validation data sets (including compounds from the Physician's Desk References and Comprehensive Medicinal Chemistry databases) is 1 log unit and the unsigned error is 0.77. This model does not require 3-D structure generation which is rather time-consuming. Using 2-D structure as input, this model is able to compute solubility for 90 000-700 000 compounds/h on a SGI Origin 2000 workstation. This kind of fast calculation allows the model to be used in data mining and screening of large synthesized or virtual libraries.

Databases, Factual↗