Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

The value of urine citrate/calcium ratio in the estimation of risk of urolithiasis.

The urine saturation is considered as the better parameter for the estimation of risk of urolithiasis than any single urinary constituent. However, the determination of urine saturation is unsuitable for routine clinical practice. To evaluate a simpler and cheaper test than urine saturation for distinguishing stone formers from healthy individuals, urinary citrate/calcium ratio was determined in 30 children with urolithiasis, 36 children with isolated hematuria, and 15 healthy control children. The ratio was significantly lower in urolithiasis group comparing to controls, and significantly higher in hematuria than in urolithiasis group. The cut-off points between normal children and children with urolithiasis, accuracy, specificity and sensitivity were determined and compared with those of the urine saturation calculated with the computer program EQUIL 2. The data mining Weka software was used for the determination of the cut-off points. Children with urolithiasis had citrate/calcium ratio below 1.38 and urine saturation above 5.285. The citrate/calcium ratio showed in comparison to urine saturation similar high accuracy (91.11 vs. 88.89%), somewhat lesser specificity (73.33% vs. 93.33%) and much better sensitivity (100% vs. 86.89%) in discrimination of stone formers from normal children. The advantage in comparison to urine saturation is that it can be easily performed in clinical practice.

Analysis of Variance↗

GEDA: new knowledge base of gene expression in drug addiction.

Abuse of drugs can elicit compulsive drug seeking behaviors upon repeated administration, and ultimately leads to the phenomenon of addiction. We developed a procedure for the standardization of microarray gene expression data of rat brain in drug addiction and stored them in a single integrated database system, focusing on more effective data processing and interpretation. Another characteristic of the present database is that it has a systematic flexibility for statistical analysis and linking with other databases. Basically, we adopt an intelligent SQL querying system, as the foundation of our DB, in order to set up an interactive module which can automatically read the raw gene expression data in the standardized format. We maximize the usability of this DB, helping users study significant gene expression and identify biological function of the genes through integrated up-to-date gene information such as GO annotation and metabolic pathway. For collecting the latest information of selected gene from the database, we also set up the local BLAST search engine and nonredundant sequence database updated by NCBI server on a daily basis. We find that the present database is a useful query interface and data-mining tool, specifically for finding out the genes related to drug addiction. We apply this system to the identification and characterization of methamphetamine-induced genes' behavior in rat brain.

Animals↗

Technology that will initiate future revolutionary changes in healthcare and the clinical laboratory.

Future trends in healthcare delivery will focus on preventive medicine that will require integration of molecular medicine with advanced information and computer technology. Molecular medicine will shift from costly intervention and treatment of established diseases to proactive prediction and prevention of disease risks. This approach will require new informatic systems that will link large scale databanks and special programs for data mining and retrieval in bioinformatics, cheminformatics, and population genetics. The clinical laboratory will soon be able to provide powerful new molecular diagnostic tools along with multianalytic assays for expression of genes and proteins in different patterns of diseases, disease progression, and predisposition to diseases.

Biotechnology↗

Train case managers in information technology.

At Hinsdale (IL) Hospital, health information specialists and case managers have created an innovative training and orientation program to get new case managers up to speed on the latest tools and techniques in information management. Established to help clinical case managers better comes to terms with complicated financial data, the program has grown into a comprehensive training system that acquaints new hires and old hands alike with the hospital's databases and teaches them appropriate methods of data mining, collection and usage. The program consists of a one-and-a-half hour initial orientation, a two-hour follow-up, and then less formal weekly and monthly meetings held on an ongoing basis.

Case Management↗

Cancer gene discovery using digital differential display.

The Cancer Gene Anatomy Project database of the National Cancer Institute has thousands of expressed sequences, both known and novel, in the form of expressed sequence tags (ESTs). These ESTs, derived from diverse normal and tumor cDNA libraries, offer an attractive starting point for cancer gene discovery. Using a data-mining tool called Digital Differential Display (DDD) from the Cancer Gene Anatomy Project database, ESTs from six different solid tumor types (breast, colon, lung, ovary, pancreas, and prostate) were analyzed for differential expression. An electronic expression profile and chromosomal map position of these hits were generated from the Unigene database. The hits were categorized into major classes of genes including ribosomal proteins, enzymes, cell surface molecules, secretory proteins, adhesion molecules, and immunoglobulins and were found to be differentially expressed in these tumorderived libraries. Genes known to be up-regulated in prostate, breast, and pancreatic carcinomas were discovered by DDD, demonstrating the utility of this technique. Two hundred known genes and 500 novel sequences were discovered to be differentially expressed in these select tumor-derived libraries. Test genes were validated for expression specificity by reverse transcription-PCR, providing a proof of concept for gene discovery by DDD. A comprehensive database of hits can be accessed at http:// www.fau.edu/cmbb/publications/cancergenes. htm. This solid tumor DDD database should facilitate target identification for cancer diagnostics and therapeutics.

Biological Specimen Banks↗

Preparing for a decision support system.

The increasing pressure to reduce costs and improve outcomes is driving the health care industry to view information as a competitive advantage. Timely information is required to help reduce inefficiencies and improve patient care. Numerous disparate operational or transactional information systems with inconsistent and often conflicting data are no longer adequate to meet the information needs of integrated care delivery systems and networks in competitive managed care environments. This article reviews decision support system characteristics and describes a process to assess the preparedness of an organization to implement and use decision support systems to achieve a more effective, information-based decision process. Decision support tools included in this article range from reports to data mining.

Decision Support Systems, Management↗

High-tech software sleuthing. New computer tools give government tighter handle on hard-to-track healthcare fraud.

After developing new data-mining techniques, the government no longer needs to rely on outdated methods when investigating allegations of hard-to-track healthcare fraud. Today, federal officials, such as U.S. Deputy Attorney General Eric Holder Jr. (left), have an enhanced ability to collect and analyze massive amounts of financial and medical information necessary to successfully prosecute perpetrators of crimes that drain the Medicare program.

Aged↗

Information systems: the key to evidence-based health practice.

Increasing prominence is being given to the use of best current evidence in clinical practice and health services and programme management decision-making. The role of information in evidence-based practice (EBP) is discussed, together with questions of how advanced information systems and technology (IS&T) can contribute to the establishment of a broader perspective for EBP. The author examines the development, validation and use of a variety of sources of evidence and knowledge that go beyond the well-established paradigm of research, clinical trials, and systematic literature review. Opportunities and challenges in the implementation and use of IS&T and knowledge management tools are examined for six application areas: reference databases, contextual data, clinical data repositories, administrative data repositories, decision support software, and Internet-based interactive health information and communication. Computerized and telecommunications applications that support EBP follow a hierarchy in which systems, tasks and complexity range from reference retrieval and the processing of relatively routine transactions, to complex "data mining" and rule-driven decision support systems.

Decision Support Systems, Clinical↗

Artificial neural networks for classifying olfactory signals.

For practical applications, artificial neural networks have to meet several requirements: Mainly they should learn quick, classify accurate and behave robust. Programs should be user-friendly and should not need the presence of an expert for fine tuning diverse learning parameters. The present paper demonstrates an approach using an oversized network topology, adaptive propagation (APROP), a modified error function, and averaging outputs of four networks described for the first time. As an example, signals from different semiconductor gas sensors of an electronic nose were classified. The electronic nose smelt different types of edible oil with extremely different a-priori-probabilities. The fully-specified neural network classifier fulfilled the above mentioned demands. The new approach will be helpful not only for classifying olfactory signals automatically but also in many other fields in medicine, e.g. in data mining from medical databases.

Algorithms↗

[Cancer genome or the development of molecular portraits of tumors].

The rapid development of cancer genomics is due to important progresses in oncogenesis, human genome sequencing and emergence of new technologies in genome and transcriptome analysis. In this context, the aim of the French program 'Cartes d'Identites des Tumeurs--Molecular Portraits of Tumors' is to build a public data base containing a pan genome assessment of genome and transcriptome alterations in the major types of tumors as well as in relevant normal cells and experimental models. Data mining is done in the context of genome annotations and clinical and biological informations attached to the enrolled samples. The goal of the program is to define new tests useful for diagnostic procedures in clinical laboratories and new targets for biological treatments of tumors.

France↗

BioProspector: discovering conserved DNA motifs in upstream regulatory regions of co-expressed genes.

The development of genome sequencing and DNA microarray analysis of gene expression gives rise to the demand for data-mining tools. BioProspector, a C program using a Gibbs sampling strategy, examines the upstream region of genes in the same gene expression pattern group and looks for regulatory sequence motifs. BioProspector uses zero to third-order Markov background models whose parameters are either given by the user or estimated from a specified sequence file. The significance of each motif found is judged based on a motif score distribution estimated by a Monte Carlo method. In addition, BioProspector modifies the motif model used in the earlier Gibbs samplers to allow for the modeling of gapped motifs and motifs with palindromic patterns. All these modifications greatly improve the performance of the program. Although testing and development are still in progress, the program has shown preliminary success in finding the binding motifs for Saccharomyces cerevisiae RAP1, Bacillus subtilis RNA polymerase, and Escherichia coli CRP. We are currently working on combining BioProspector with a clustering program to explore gene expression networks and regulatory mechanisms.

Bacillus subtilis↗

[Computer system "Gene Discovery" for searching for regularities in organization of eukaryotic regulatory sequences].

A method is proposed to automatically search for patterns in the mutual location of context signals in regulatory DNA sequences. The procedure is based on the methods of Data Mining and Knowledge Discovery software, implemented in a computer system Gene Discovery. This system was used to study erythroid-specific promoters and promoters of the endocrine-system genes from TRRD. We detected some trends in occurrence and localization of specific oligonucleotide groups.

Computer Systems↗

Genomic approaches to elucidating the pathophysiology of renal diseases.

The physiological and pathological processes of the kidney as a whole can now be analyzed with a molecular precision at a genomic-scale. Using massively parallel cDNA microarray technology, the mRNA expression of thousands of genes can be quantified simultaneously. The advantages of microarray analyses include the ability to examine the interaction of several genes or the entire genome in a single experiment. Bioinformatics approaches such as data mining through mathematical condensation of the massive gene expression profiles are essential for elucidating molecular and biological logic underlying gene expression programs. Genes that encode similar protein components are often coordinately regulated. Recent application of gene expression profiling to the normal human renal cortical tissue, experimental in vitro and in vivo models has shown that cellular activation is accompanied by changes of hundreds of genes in parallel. The databases of gene expression emerging from these studies will be used to interpret the pathological changes in gene expression that accompany a variety of human renal diseases.

Gene Expression Profiling↗

A knowledge model for the interpretation and visualization of NLP-parsed discharged summaries.

At our institution, a Natural Language Processing (NLP) tool called MedLEE is used on a daily basis to parse medical texts including complete discharge summaries. MedLEE transforms written text into a generic structured format, which preserves the richness of the underlying natural language expressions by the use of concept modifiers (like change, certainty, degree and status). As a tradeoff, extraction of application-specific medical information is difficult without a clear understanding of how these modifiers combine. We report on a knowledge model for MedLEE modifiers that is helpful for a high level interpretation of NLP data and is used for the generation of two distinct views on NLP-parsed discharge summaries: A physician view offering a condensed overview of the severity of patient problems and a data mining view featuring binary problem states useful for machine learning.

Artificial Intelligence↗

Theory, abstraction and design in medical informatics.

OBJECTIVE: To analyze the scientific and engineering components of Medical Informatics. A clear characterization of these components should be undertaken to categorize different areas of Medical Informatics and create a research agenda for the future. METHODS: We have adapted a classical ACM and IEEE report on computing to analyze Medical Informatics from three different viewpoints: Theory, Abstraction, and Design. RESULTS: We suggest that Medical Informatics can be considered from these three perspectives: (1) Theory, from which medical informaticians formally characterize the properties of the objects of study, creating new theories or using and adapting existing theories (e.g., from mathematics), (2) Abstraction, from which medical informaticians deal with all aspects of medical information and create new abstractions, methods, and technology-independent models, which can be experimentally verified, and (3) Design, from which medical informaticians develop systems or act as information brokers or advisors between medical and technology professionals, to improve the quality of computer applications in medicine. CONCLUSION: Based on this framework, we suggest that Medical Informatics has an independent scientific character, different from other applied informatics areas. Finally, we analyze these three perspectives using data mining in medicine.

Medical Informatics↗

Fragment analysis in small molecule discovery.

Cheminformatics is playing an ever-increasing role in small molecule drug discovery. The widespread use of high-throughput screening (HTS) and combinatorial chemistry techniques has led to the generation of large amounts of pharmacological data which, in turn, has catalyzed the development of computational methods designed to reduce the time and cost in identifying molecules suitable for pharmaceutical development. This review focuses on recent advances in the field of substructure analysis, an increasingly popular data mining technique with applications at many levels of the discovery process, including HTS, compound library design, virtual screening and the prediction of biological activity.

Animals↗

A humanist's legacy in medical informatics: visions and accomplishments of Professor Jean-Raoul Scherrer.

OBJECTIVE: To report about the work of Prof. Jean-Raoul Scherrer, and show how his humanist vision, his medical skills and his scientific background have enabled and shaped the development of medical informatics over the last 30 years. RESULTS: Starting with the mainframe-based patient-centered hospital information system DIOGENE in the 70s, Prof. Scherrer developed, implemented and evolved innovative concepts of man-machine interfaces, distributed and federated environments, leading the way with information systems that obstinately focused on the support of care providers and patients. Through a rigorous design of terminologies and ontologies, the DIOGENE data would then serve as a basis for the development of clinical research, data mining, and lead to innovative natural language processing techniques. In parallel, Prof. Scherrer supported the development of medical image management, ranging from a distributed picture archiving and communication systems (PACS) to molecular imaging of protein electrophoreses. Recognizing the need for improving the quality and trustworthiness of medical information on the Web, Prof. Scherrer created the Health-On-the-Net (HON) foundation. CONCLUSIONS: These achievements, made possible thanks to his visionary mind, deep humanism, creativity, generosity and determination, have made of Prof. Scherrer a true pioneer and leader of the human-centered, patient-oriented application of information technology for improving healthcare.

History, 20th Century↗

Multiple splice variants of lactate dehydrogenase C selectively expressed in human cancer.

We applied a combined data mining and experimental validation Approach for the discovery of germ cell-specific genes aberrantly expressed in cancer. Six of 21 genes with confirmed germ cell specificity were detected in tumors, indicating that ectopic activation of testis-specific genes in cancer is a frequent phenomenon. Most surprisingly one of the genes represented lactate dehydrogenase C (LDHC), the germ cell-specific member of the lactate dehydrogenase family. LDHC escapes from transcriptional repression, resulting in significant expression levels in virtually all tumor types tested. Moreover, we discovered aberrant splicing of LDHC restricted to cancer cells, resulting in four novel tumor-specific variants displaying structural alterations of the catalytic domain. Expression of LDHC in tumors is neither mediated by gene promotor demethylation, as previously described for other germ cell-specific genes activated in cancer, nor induced by hypoxia as demonstrated for enzymes of the glycolytic pathway. LDHC represents the first lactate dehydrogenase isoform with restriction to tumor cells. In contrast to other LDH isoenzymes, LDHC has a preference for lactate as a substrate. Thus LDHC activation in cancer may provide a metabolic rescue pathway in tumor cells by exploiting lactate for ATP delivery.

Alternative Splicing↗