Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Molecular classification of breast carcinomas by comparative genomic hybridization: a specific somatic genetic profile for BRCA1 tumors.

In approximately 70% of the families with a high frequency of early-onset breast and/or ovarian cancer, BRCA1 or BRCA2 germline mutations cannot be identified with the current screening regime. Therefore, we used data mining to identify a somatic genetic signature to differentiate BRCA1 mutation carriers from non-BRCA1 carriers based on the genetic characteristics of their breast carcinomas. For this purpose, we developed a molecular classifier, which assigns a given tumor to either the BRCA1 or control group based on somatic genetic profiles as revealed by comparative genomic hybridization. This was performed on breast tumors selected from two groups of patients: 28 proven BRCA1 germline mutation carriers; and a control group consisting of 42 breast tumors from patients with unknown BRCA1 or BRCA2 status. We show that BRCA1 breast carcinomas exhibit specific somatic genetic aberrations and can be distinguished from control tumors with an accuracy of 84% (sensitivity of 96% and specificity of 76%). Chromosomal bands used by this classifier include regions on chromosomes 3p, 3q, and 5q. The classifier miss-assigned one patient with a BRCA1 mutation to the non-BRCA1 class. The germline mutation in this patient is a 62bp deletion in the last exon of BRCA1 (5622del62). Possibly, this mutation may give a different phenotypic effect than do mutations in other regions of the gene. Validation on an independent set of BRCA1 and sporadic tumors showed that the BRCA1 classifier correctly identified all 6 BRCA1 tumors and assigned 4 of the 19 control patients to the BRCA1 class. The resulting accuracy on the validation set is 84%.

Breast Neoplasms↗

Establishing connections between microarray expression data and chemotherapeutic cancer pharmacology.

We have investigated three different microarray datasets of approximately 6 K gene expressions across the National Cancer Institute's panel of 60 tumor cell lines. Initial assessments of reproducibility for gene expressions within each dataset, as derived from sequence analysis of full-length sequences as well as expressed sequence tags (EST), found statistically significant results for no more than 36% of those cases where at least one replicate of a gene appears on the array. Filtering the data based only on pairwise comparisons among these three datasets creates a list of approximately 400 significant concordant expression patterns. The expression profiles of these smaller sets of genes were used to locate similar expression profiles of synthetic agents screened against these same 60 tumor cell lines. A correspondence was found between mRNA expression patterns and 50% growth inhibition response patterns of screened agents for 11 cases that were subsequently verifiable from ligand-target crystallographic data. Notable amongst these cases are genes encoding a variety of kinases, which were also found to be targets of small drug-like molecules within the database of protein structures. These 11 cases lend support to the premise that similarities between expression patterns and chemical responses for the National Cancer Institute's tumor panel can be related to known cases of molecular structure and putative cellular function. The details of the 11 verifiable cases and the concordant gene subsets are provided. Discussions about the prospects of using this approach as a data mining tool are included.

Algorithms↗

Ontologies for molecular biology and bioinformatics.

About five years ago, ontology was almost unknown in bioinformatics, even more so in molecular biology. Nowadays, many bioinformatics articles mention it in connection with text mining, data integration or as a metaphysical cure for problems in standardisation of nomenclature and other applications. This article attempts to give an account of what concept ontologies in the domain of biology and bioinformatics are; what they are not; how they can be constructed; how they can be used; and some fallacies and pitfalls creators and users should be aware of.

Computational Biology↗

PESI--an intelligent system for prediction of enzyme-substrate interactions based on experimental constraints.

We present a system for predicting protein-protein modifications, and demonstrate its usefulness in the field of signal transduction research. Signal transduction is one of the most important areas of investigation in biological research. One of the major mechanisms frequently employed by cells to regulate signal transduction processes involves protein phosphorylation by various kinases. As many as 1,000 protein kinases and 500 protein phosphatases in the human genome are thought to be involved in phosphorylation processes which regulate all aspects of cell function. The complexity of such interactions stems from the enormous number of factors and interactions, which makes the identification of putative substrates for any given enzyme by straightforward experimentation increasingly difficult. We present here a data mining algorithm, based on the similarity between the modifier proteins and between the modified proteins, and on experimental constraints. The application presented here (PESI) focuses on substrate phosphorylation by various enzymes. This algorithm reduces the number of substrate candidates for experimental study by about two orders of magnitude. Moreover, this algorithm has already yielded predictions for previously unknown substrates of the enzymes PKCdelta and PKCeta, which we have confirmed experimentally.

Algorithms↗

Molecular description of evolving paclitaxel resistance in the SKOV-3 human ovarian carcinoma cell line.

Ovarian cancer is currently the most lethal gynecological malignancy in the United States. Although effective therapies exist, the acquisition of multidrug resistance within persisting tumor cells renders curative therapies elusive for the majority of women with ovarian cancer. In an attempt to better define the evolution of paclitaxel resistance, three SKOV-3 sublines were selected during successive rounds of exposure to increasing paclitaxel concentrations. The sublines were selected to represent early (0.003 micro M), intermediate (0.03 micro M), and late (0.3 micro M) paclitaxel resistance. RNA from these cell lines, SKOV-3(0.003TR), SKOV-3(0.03TR), and SKOV-3(0.3TR), as well as the parent cell line SKOV-3, was analyzed by cDNA array to evaluate transcript expression profiles. Arrays were performed using Affymetrix HG-U95Av2 arrays, which contain probes for approximately 9600 known human genes. Signal intensities were calculated by Microarray Suite 5.0 (Affymetrix, Santa Clara, CA). Expression patterns were analyzed by Affymetrix Data Mining Tool 3.0 with filtering of expression patterns for fold change in expression (maximum divided by minimum expression value/gene) and for variation of expression (maximum minus minimum expression value/gene). This analysis dismissed approximately 11,000 of approximately 12,000 expression patterns. The remaining approximately 1000 expression patterns were normalized and segregated into 20 partitions of a self-organizing map (SOM). The resulting SOM discriminates between genes, which are differentially expressed in early versus intermediate versus late paclitaxel resistance. For example, multidrug resistance 1 transcript expression is not elevated in SKOV-3(0.003TR) as compared with parental SKOV-3 but demonstrates elevated expression in SKOV-3(0.03TR) and SKOV-3(0.3TR). In contrast, SOM analysis demonstrates early (SKOV-3(0.003TR)) transcriptional changes in a wide variety of genes, including gene families involved in cell growth/maintenance, cell structure, signal transduction, and inflammatory response. The use of array analysis with SOMs in sublines with progressive paclitaxel resistance can successfully define an evolution of resistance. Such an analysis may be useful at defining candidate gene families involved in the early-drug resistance phenotype.

Algorithms↗

Informatics united: exemplary studies combining medical informatics, neuroinformatics and bioinformatics.

OBJECTIVES: Medical informatics, neuroinformatics and bioinformatics provide a wide spectrum of research. Here, we show the great potential of synergies between these research areas on the basis of four exemplary studies where techniques are transferred from one of the disciplines to the other. METHODS: Reviewing and analyzing exemplary and specific projects at the intersection of medical informatics, neuroinformatics, and bioinformatics from our experience in an interdisciplinary research group. RESULTS: Synergy emerges when techniques and solutions from medical informatics, bioinformatics, or neuroinformatics are successfully applied in one of the other disciplines. Synergy was found in 1. the modeling of neurophysiological systems for medical therapy development, 2. the use of image processing techniques from medical computer vision for the analysis of the dynamics of cell nuclei, and 3. the application of neuroinformatics tools for data mining in bioinformatics and as classifiers in clinical oncology. CONCLUSIONS: Each of the three different disciplines have delivered technologies that are readily applicable in the other disciplines. The mutual transfer of knowledge and techniques proved to increase efficiency and accuracy in a manifold of applications. In particular, we expect that clinical decision support systems based on techniques derived from neuro- and bioinformatics have the potential to improve medical diagnostics and will finally lead to a personalized delivery of healthcare.

Computational Biology↗

Beyond the genome--SMi Conference 27-28 January 2003, London, UK.

Over a decade of astonishing developments, genomics and proteomics have promised a fundamentally new approach to drug discovery. Although there has been an undeniable increase in the range of potential targets available, this has not led to an increased output of the drug discovery pipeline into the clinic. With tighter markets and increasing competition, the major pharmaceutical companies are under intense pressure to achieve rapid, concrete delivery of those early promises, but there remain acute problems in the genes-to-drugs pipeline. This meeting showcased a range of novel approaches from proteomics and bioinformatics to address these problems. A common theme in the range of proteomics offerings was the prioritization of potential novel targets on the basis of their accessibility to drugs and their functional link to disease phenotypes. Informatics and in silico offerings also concentrated on fast, accurate, drug-focused workflows built on large integrative databases and novel data-mined algorithms.

Animals↗

New terminology services based on term comparison using semantic definitions and similarity computation.

As medical information can be encoded within different terminological systems, terms comparison is an important issue to allow communication between applications. In description logics, terms are compared by the means of semantic definitions and subsumption relations. Similarity is also a convenient method for term comparison but is not supported by terminology servers which implement subsumption relations. We present new terminology services built on a semantic distance that could help for semantic mediation between medical applications ranging from semi-automatic encoders to data mining tools. These services are 1) comparison between terms 2) k-nearest-neighbors 3) support for concept coding 4) automatic generation of similarity tables 5) distance based queries 6) support for clustering.

Artificial Intelligence↗

MIMIC II: a massive temporal ICU patient database to support research in intelligent patient monitoring.

Development and evaluation of Intensive Care Unit (ICU) decision-support systems would be greatly facilitated by the availability of a large-scale ICU patient database. Following our previous efforts with the MIMIC (Multi-parameter Intelligent Monitoring for Intensive Care) Database, we have leveraged advances in networking and storage technologies to develop a far more massive temporal database, MIMIC II. MIMIC II is an ongoing effort: data is continuously and prospectively archived from all ICU patients in our hospital. MIMIC II now consists of over 800 ICU patient records including over 120 gigabytes of data and is growing. A customized archiving system was used to store continuously up to four waveforms and 30 different parameters from ICU patient monitors. An integrated user-friendly relational database was developed for browsing of patients' clinical information (lab results, fluid balance, medications, nurses' progress notes). Based upon its unprecedented size and scope, MIMIC II will prove to be an important resource for intelligent patient monitoring research, and will support efforts in medical data mining and knowledge-discovery.

Artificial Intelligence↗

Multiscale analysis of long time-series medical databases.

Data mining in time-series medical databases has been receiving considerable attention since it provides a way of revealing useful information hidden in the database; for example relationships between the temporal course of examination results and onset time of diseases. This paper presents a new method for finding similar patterns in temporal sequences based on multiscale matching. Multiscale matching enables us the cross-scale comparison of sequences, namely, it enable us to compare temporal patterns by partially changing observation scales. We examined the usefulness of the method on the chronic hepatitis dataset and found some interesting patterns. On GPT sequences, we found patterns that may represent the effectiveness of interferon (IFN) treatment. On platelet count sequences, we found that, if IFN treatment was ineffective, platelet count kept decreasing following the progress of liver fibrosis, while it started increasing if the treatment was effective.

Antiviral Agents↗

Complete blood count reference interval diagrams derived from NHANES III: stratification by age, sex, and race.

BACKGROUND: Comprehensive, up-to-date "health-associated" reference interval studies of North American populations are uncommon. The third US National Health and Nutrition Examination Survey (NHANES III) was concluded in 1994 and yielded important reference interval data. OBJECTIVE: To obtain health-associated Coulter counter reference interval data from NHANES III according to age, sex, and race. METHODS: Of the 29,314 civilian noninstitutionalized US citizens who participated in NHANES III, approximately 25,000 had a complete blood count, red cell distribution width (RDW), platelet count, and automated white blood cell (WBC) differential determined on a Coulter S-Plus Jr. To determine health-associated reference intervals, we used the following exclusion criteria: pregnancy, breast feeding, obesity (body mass index [BMI] >40 and >35 for females and males, respectively), diastolic blood pressure >100 mm Hg, any smoking, any drinking of alcohol, recent treatment for anemia, creatinine level >2.5 mg/dL, glucose level >126 mg/dL, excessive thinness (BMI <8), recent surgery or hospitalization, or having antibodies to hepatitis viruses A, B, or C. The Coulter counter data (hemoglobin, hematocrit, red blood cell count, mean corpuscular volume (MCV), mean cell hemoglobin concentration (MCHC), MCH, WBC count, platelet count, granulocyte count, monocyte count, lymphocyte count, RDW, platelet distribution width, and mean platelet volume) were separated into 6 sex/racial categories (female non-Hispanic white, female non-Hispanic black, female Mexican American, male non-Hispanic white, male non-Hispanic black, and male Mexican American) and 9 age groupings (10-14, 14-18, 18-25, 25-35, 35-45, 45-55, 55-65, 65-75, and >75 years). RESULTS: There was a high exclusion rate; for example, of the 20,685 individuals with measured hemoglobin levels, 12,688 (61.3%) were excluded. Percentile estimates could be derived accurately for almost all of the female age/sex categories. A few of the male Mexican American and non-Hispanic black categories contained observations for ages 45 to 75 years. CONCLUSIONS: There are age-dependent trends for many of the tests, notably in RDW, MCMV, platelet count, and granulocyte and lymphocyte percentages. Sex-dependent changes involved hemoglobin values, and race-related trends centered around mononuclear and lymphocyte percentages, hematocrit, MCHC, MCH, and hemoglobin. This study reveals the potential for using data mining of large samples to yield potentially useful reference ranges.

Adolescent↗

Ideal discrimination of discrete clinical endpoints using multilocus genotypes.

Multifactor Dimensionality Reduction (MDR) is a method for the classification and prediction of discrete clinical endpoints using attributes constructed from multilocus genotype data. Empirical studies with both real and simulated data suggest that MDR has good power for detecting gene-gene interactions in the absence of independent main effects. The purpose of this study is to develop an objective, theory-driven approach to evaluate the strengths and limitations of MDR. To accomplish this goal, we borrow concepts from ideal observer analysis used in visual perception to evaluate the theoretical limits of classifying and predicting discrete clinical endpoints using multilocus genotype data. We conclude that MDR ideally discriminates between low risk and high risk subjects using attributes constructed from multilocus genotype data. We also how that the classification approach used once a multilocus attribute is constructed is similar to that of a naive Bayes classifier. This study provides a theoretical foundation for the continued development, evaluation, and application of the MDR as a data mining tool in the domain of statistical genetics and genetic epidemiology.

Animals↗

Parametric brain MR atlases: standardization for imaging informatics.

This paper is focused on the development of normal MR brain atlases of intrinsic MR parameters. These parameters permit quantitative comparisons across imaging studies (as opposed to raw image intensity values) and are important markers of neurological diseases. The development includes fast sequences to generate three parameters (T1: spin-lattice relaxation, T2: spin-spin relaxation, and Diffusion Tensor) covering the whole brain with isotropic and high-resolution images. The analysis of raw data to generate the parametric images is followed by registration algorithms to bring the image studies acquired on normal subjects aligned to a common frame of reference. The registration method includes both linear and non-linear algorithms. Two atlas schemes are discussed: an average atlas and a probabilistic atlas. Initial results on sequence development and registration are presented. The atlases are envisaged as an integral part of an imaging informatics infrastructure that enables image analysis across imaging studies to perform automated image data mining.

Algorithms↗

Potential impact of advanced clinical information technology on healthcare in 2015.

Clinical information technologies now sporadically available will soon be in routine clinical use, bringing many changes to healthcare. For example, 1) The next generation Internet; 2) Real-time clinical decision support systems; 3) Off-line, population-based systems; 4) Large, integrated, individual patient-level phenotypic and genotypic databases with intelligent data mining capabilities; 5) Wireless, invasive and non-invasive physiologic monitoring devices; 6) Natural Language Processing (NLP) systems; and 7) Mathematical models of complex biological systems have the potential to impact significantly the future healthcare delivery system. While new information management and communication techniques and technologies will reduce many of the inefficiencies and inaccuracies of our present systems, there will be an equal, and potentially far more dangerous, set of unintended consequences. Informatics investigators and health system administrators must focus on the study of what is working and what is not, as well as, on development and testing of the new clinical information management and communication technologies, if we are to be ready for the future.

Databases as Topic↗

Leveraging intelligent agents for knowledge discovery from heterogeneous healthcare data repositories.

This paper presents a case for an intelligent agent based framework for knowledge discovery in a distributed healthcare environment comprising multiple heterogeneous healthcare data repositories. Data-mediated knowledge discovery, especially from multiple heterogeneous data resources, is a tedious process and imposes significant operational constraints on end-users. We demonstrate that autonomous, reactive and proactive intelligent agents provide an opportunity to generate end-user oriented, packaged, value-added decision-support/strategic planning services for healthcare professionals, manages and policy makers, without the need for a priori technical knowledge. Since effective healthcare is grounded in good communication, experience sharing, continuous learning and proactive actions, we use intelligent agents to implement an Agent based Data Mining Infostructure that provides a suite of healthcare-oriented decision-support/strategic planning services.

Efficiency, Organizational↗

Automatic diagnosis classification of patient discharge letters.

CAIRN (Computer Assisted Medical Information Resource Navigation) is a prototyping System that allows flexible medical data storage and retrieval supporting medical informatics research. In this paper methods that automate the selection of ICD-9 diagnosis (International Classification of Diseases and Diagnoses, 9th Revision) are investigated. We present the Text Data Mining module extension of CAIRN and its application in order to organize in a systematic way uncontrolled terms, to propose relationships between uncontrolled terms and finally aid the diagnosis classification.

Disease↗

[The in silico elongation and analysis of the EST from Schistosoma japonicum].

OBJECTIVE: To construct a platform for in silico elongation and batch analysis of Schistosoma japonicum (Sj) ESTs, acquire the potential novel genes and research the expression profile of the genes. METHODS: On the basis of Linux operating system and local ESTs database of Sj, the BLAST and PHRAP softwares were used to construct a program to achieve the elongation of ESTs. Stand-alone BLAST search against the nr database helped analyze the elongated sequence. After finishing the batch analysis script, the platform was used to research the Sj gene expression profile and acquire the potential novel genes. RESULTS: The platform showed satisfactory efficiency and fidelity. 487 elongated sequences obtained from 552 and 307 elongated sequences showed high homology within the nr database downloaded from NCBI. Furthermore, 104 elongated sequences displayed significant homology but showed no homology before elongated. 27 potential novel genes were filtered out. CONCLUSION: An effective platform for Sj ESTs data mining was accomplished and further information on the potential novel genes was acquired.

Animals↗

[An introduction of several programs used in genomic analysis].

Genomics is a novel subject that has been developed accompanying with the progress of human genome project. Genomics deals with the chemistry component, structure organization and evolution of genome at global level. As genomics associated with huge data, bioinformatics plays an important role in these processes of data production, data management and data mining. At present, many reliable programs have been used in genomic research successfully, which are usually accessible and downloaded freely. We address here the principles of some programs used wildly in genomics such as sequence alignment, sequence assembly, repeat identification and gene prediction, which are exemplified with typical programs respectively.

English Abstract↗