Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Bioinformatics and food allergens.

Bioinformatics can play an important role in developing improved technology for the detection and characterization of food allergens. However, the full realization of this potential will depend on the development of allergen-specific databases as well as improved methods for data mining within these databases. Examples of existing allergen databases and analysis tools are described, as are the most important issues that need to be addressed in the next stage of database development.

Allergens↗

Towards creating an informatics infrastructure in home health care.

Although information technology is utilized successfully in many industries, its use in health care-and home health care in particular--continues to lag. This column summarizes a recent article by Bakken and Hripcsak (2004) examining the potential for informatics to improve patient care quality in home health care by supporting evidence-based practices and patient safety. The authors provide definitions of the basic components of an informatics infrastructure e.g., data mining, digital sources of evidence, etc.--and recommend how to make an informatics infrastructure for the home health care industry a reality. Suggestions include: (1) integrating informatics into education and training; (2) creating public/private partnerships among government agencies, vendors, and industry associations; and (3) performing cost-effective analyses to determine the optimal uses of specific technologies.

Cooperative Behavior↗

Web-based health services and clinical decision support.

The purpose of this study was the development of a Web-based e-health service for comprehensive assistance and clinical decision support. The service structure consists of a Web server, a PHP-based Web interface linked to a clinical SQL database, Java applets for interactive manipulation and visualization of signals and a Matlab server linked with signal and data processing algorithms implemented by Matlab programs. The service ensures diagnostic signal- and image analysis-sbased clinical decision support. By using the discussed methodology, a pilot service for pathology specialists for automatic calculation of the proliferation index has been developed. Physicians use a simple Web interface for uploading the pictures under investigation to the server; subsequently a Java applet interface is used for outlining the region of interest and, after processing on the server, the requested proliferation index value is calculated. There is also an "expert corner", where experts can submit their index estimates and comments on particular images, which is especially important for system developers. These expert evaluations are used for optimization and verification of automatic analysis algorithms. Decision support trials have been conducted for ECG and ophthalmology ultrasonic investigations of intraocular tumor differentiation. Data mining algorithms have been applied and decision support trees constructed. These services are under implementation by a Web-based system too. The study has shown that the Web-based structure ensures more effective, flexible and accessible services compared with standalone programs and is very convenient for biomedical engineers and physicians, especially in the development phase.

Adult↗

Applying informatics in tissue engineering.

OBJECTIVE: To facilitate tissue engineering strategies determination with informatics tools. METHODS: Firstly, tissue engineering experimental data were standardized and integrated into a centralized database; secondly, we used data mining tools (e.g. artificial neural networks and decision trees) to predict the outcomes of tissue engineering strategies; thirdly, a strategy design algorithm was developed, and its efficacy was validated with animal experiments; lastly, we constructed an online database and a decision support system for tissue engineering. RESULTS: The artificial neural networks and the decision trees respectively predicted the outcomes of tissue engineering strategies with the predictive accuracy of 95.14% and 85.26%. Following the strategies generated by computer, we cured 18 of the 20 experimental animals with a significantly lower cost than usual. CONCLUSION: Informatics is beneficial for realizing safe, effective and economical tissue engineering.

Artificial Intelligence↗

Protein-based analysis of alternative splicing in the human genome.

Understanding the functional significance of alternative splicing and other mechanisms that generate RNA transcript diversity is an important challenge facing modern-day molecular biology. Using homology-based, protein sequence analysis methods, it should be possible to investigate how transcript diversity impacts protein structure and function. To test this, a data mining technique ("DiffHit") was developed to identify and catalog genes producing protein isoforms which exhibit distinct profiles of conserved protein motifs. We found that out of a test set of over 1,300 alternatively spliced genes with solved genomic structure, over 30% exhibited a differential profile of conserved InterPro and/or Blocks protein motifs across distinct isoforms. These results suggest that motif databases such as Blocks and InterPro are potentially useful tools for investigating how alternative transcript structure affects gene function.

Algorithms↗

[Construction and application of an in silico elongation system of nucleic acid sequence].

Normally it is difficult to obtain full-length cDNA sequence of novel genes. More and more expressed sequence tags(ESTs) have been obtained since the start-up of human genome project. Powerful system is badly needed for data mining on these EST sequences. Based on a personal computer coupled with Linux operating system and EST database, the Blast software and Phrap software were used to construct a platform for in silico elongation of ESTs in our lab. The performance was tested using 11386 EST sequences and 511 partial-length cDNA sequences. Results demonstrated that 8373 EST and 389 cDNA sequence were elongated using this system. Thus the platform seems to be a fast way for full-length cDNA sequence cloning of new genes.

English Abstract↗

Healthcare, molecular tools and applied genome research.

Biotechnology 2000 offered a rare opportunity for scientists from academia and industry to present and discuss data in fields as diverse as environmental biotechnology and applied genome research. The healthcare section of the meeting encompassed a number of gene therapy delivery systems that are successfully treating genetic disorders. Beta-thalassemia is being corrected in mice by continous erythropoeitin delivery from engineered muscles cells, and from naked DNA electrotransfer into muscles, as described by Dr JM Heard (Institut Pasteur, Paris, France). Dr Reszka (Max-Delbrueck-Centrum fuer Molekulare Medizin, Berlin, Germany), meanwhile, described a treatment for liver metastasis in the form of a drug carrier emolization system, DCES (Max-Delbrueck-Centrum fuer Molekulare Medizin), composed of surface modified liposomes and a substance for chemo-occlusion, which drastically reduces the blood supply to the tumor and promotes apoptosis, necrosis and antiangiogenesis. In the molecular tools section, Willem Stemmer (Maxygen Inc, Redwood City, CA, USA) gave an insight into the importance that techniques, such as molecular breeding (DNA shuffling), have in the evolution of molecules with improved function, over a range of fields including pharmaceuticals, vaccines, agriculture and chemicals. Technologies, such as ribosome display, which can incorporate the evolution and the specific enrichment of proteins/peptides in cycles of selection, could play an enormous role in the production of novel therapeutics and diagnostics in future years, as explained by Andreas Plückthun (Institute of Biochemistry, University of Zurich, Switzerland). Applied genome research offered technologies, such as the 'in vitro expression cloning', described by Dr Zwick (Promega Corp, Madison, WI, USA), are providing a functional analysis for the overwhelming flow of data emerging from high-throughput sequencing of genomes and from high-density gene expression microarrays (DNA chips). The importance of bioinformatics was stressed throughout the conference. In particular, its applications in the storage, analysis and visualization of data, but also the linkage of databases and data mining tools. It is the storage and linkage that generates insight into disease processes, enabling the nomination of candidate drug targets.

Journal Article↗

[Screening of differentially expressed genes in colorectal cancer using human whole genomic oligonucleotide microarrays].

OBJECTIVE: To screen the differentially expressed genes in human colorectal cancer (CRC) tissue. METHODS: Affymetrix oligonucleotide microarrays HG-U133 representing 32,264 human genes including 19,308 known genes and 12,956 expressed sequence tags (ESTs) were used to detect the gene expressions of CRC tissue paired with normal mucosa tissue. The microarray findings were confirmed by real-time quantitative reverse transcriptase-polymerase chain reaction (FQ-PCR). The gene expression profiles were analyzed by intersection and complement, rank sum test and t test. RESULTS: Totally 3,125 genes and ESTs expressed differentially were detected in normal and cancer tissues, consisting of 974 up-regulated and 2,151 down-regulated genes with 247 ESTs present in CRC tissue and absent in normal mucosa and 162 ESTs absent in CRC tissue but present in normal mucosa. A percent of 80.1% of the differentially expressed genes were not reported in the literatures. CONCLUSION: The strategy of data mining provides a foundation for filtering molecular markers and interpreting molecular carcinogenesis of CRC.

Colorectal Neoplasms↗

[The application of biomed-informatics in cardiovascular research--data and knowledge].

With the development of biotechnology, especially the projects which lead to high throughout experimental results, lots of data have been jammed in every related fields, and broke the balance between data and knowledge in these fields. It is urgent to refine these data, and let them be more productive. To avoid these data being decaying into rubbish, technology and theory such as database, statistics, data mining, knowledge management, and artificial intelligence had been applied into biologic and medical study fields. So, a new standalone research subject--biomed-informatics has been formed. This paper reviewed the application of biomed-informatics in cardiovascular research, and gave the view for the future development of this subject.

Animals↗

[Determining factors of intention to actual use of charged long-term care services for the aged].

OBJECTIVES: To help develop strategies to cope with the changes arising from the rapid aging process by predicting the determining factors of intention to actual use of the charged long-term care services for elderly as perceived by the middle aged who play the major role of supports. METHODS: Subjects were the parents (men 177, women 507) in their 40s of the students selected from a university of Busan city. A questionnaire survey was conducted for 4 weeks in October 2003 about the knowledge for long-term care service, the intention of actual use, and the preferences about the type of service suppliers. Data analysis was performed with frequency, chi-square test, and t-test using SPSS program (ver 10.0K), along with data mining using decision tree of Enterprise Miner V8.2 by SAS. RESULTS: About half of the subjects (53.7%) had the actual experiences of elderly supports. Intentions to use the charged services were relatively high in home visiting nursing care service (40.1%) and long-term care facilities service (40.4%), and were influenced by previous knowledge about the services. The intentions were stronger in women, those with higher education, and those with greater income levels. Actual elderly supports were mc (80%) done by women, and the perceived burdens for supports were bigger in women and those of lower s economic level. Desired charges were about 10,000 for the bath service, 20,000 won for the rests services day, and about 500,000 won for the long-term care facil service per month. From the result of decision analysis, the job professionalism was the most impol determining factor of intention to actual use of the serv with validation as 63-71%. Health and welfare mixed facilities were preferred, and the most impor consideration was the level of professionalism. CONCLUSIONS: Intention to actual use of the chad services was largely determined by the aspects of time cost. Polices to increase the number of service supp and to decrease the burdens perceived by ac supporters were strongly recommended.

Adult↗

Novel strategies to identify relevant molecular signatures for complex human diseases based on data of identical-by-decent profiles and genomic context.

OBJECTIVE: To develop novel strategies to identify relevant molecular signatures for complex human diseases based on data of identical-by-decent profiles and genomic context. METHODS: In the proposed strategies, we define four relevancy criteria for mapping SNP-phenotype relationships-point-wise IBD mean difference, averaged IBD difference for window, Z curve and averaged slope for window. RESULTS: Application of these criteria and permutation test to 100 simulated replicates for two hypothetical American populations to extract the relevant SNPs for alcoholism based on sib-pair IBD profiles of pedigrees demonstrates that the proposed strategies have successfully identified most of the simulated true loci. CONCLUSION: The data mining practice implies that IBD statistic and genomic context could be used as the informatics for locating the underlying genes for complex human diseases. Compared with the classical Haseman-Elston sib-pair regression method, the proposed strategies are more efficient for large-scale genomic mining.

Alcoholism↗

Estimating the number of clusters in DNA microarray data.

OBJECTIVES: The main objective of the research is an application of the clustering and cluster validity methods to estimate the number of clusters in cancer tumor datasets. A weighed voting technique is going to be used to improve the prediction of the number of clusters based on different data mining techniques. These tools may be used for the identification of new tumour classes using DNA microarray datasets. This estimation approach may perform a useful tool to support biological and biomedical knowledge discovery. METHODS: Three clustering and two validation algorithms were applied to two cancer tumor datasets. Recent studies confirm that there is no universal pattern recognition and clustering model to predict molecular profiles across different datasets. Thus, it is useful not to rely on one single clustering or validation method, but to apply a variety of approaches. Therefore, combination of these methods may be successfully used for the estimation of the number of clusters. RESULTS: The methods implemented in this research may contribute to the validation of clustering results and the estimation of the number of clusters. The results show that this estimation approach may represent an effective tool to support biomedical knowledge discovery and healthcare applications. CONCLUSION: The methods implemented in this research may be successfully used for the estimation of the number of clusters. The methods implemented in this research may contribute to the validation of clustering results and the estimation of the number of clusters. These tools may be used for the identification of new tumour classes using gene expression profiles.

Central Nervous System Neoplasms↗

Predicting enzyme class from protein structure using Bayesian classification.

Predicting enzyme class from protein structure parameters is a challenging problem in protein analysis. We developed a method to predict enzyme class that combines the strengths of statistical and data-mining methods. This method has a strong mathematical foundation and is simple to implement, achieving an accuracy of 45%. A comparison with the methods found in the literature designed to predict enzyme class showed that our method outperforms the existing methods.

Algorithms↗

[Global gene response to GSM 1800 MHz radiofrequency electromagnetic field in MCF-7 cells].

OBJECTIVE: To investigate whether GSM 1800 MHz radiofrequency electromagnetic field (RF EMF) can change the gene expression profile in MCF-7 cells and to screen RF EMF responsive genes. METHODS: Subcultured MCF-7 cells were intermittently (5-minute fields on/10-minute fields off) exposed or sham-exposed to GSM 1800 MHz RF EMF, which was modulated by 217 Hz EMF, for 24 hours at an average specific absorption rate (SAR) of 2.0 W/kg or 3.5 W/kg. Immediately after RF EMF exposure or sham-exposure, total RNA was isolated from MCF-7 cells and then purified. Affymetrix Human Genome U133A Genechip was applied to examine the change of gene expression profile according to the manufacturer's instruction. Data was analyzed by Affymetrix Microarray Suite 5.0 (MAS 5.0) and Affymetrix Data Mining Tool 3.0 (DMT 3.0). Quantitative reverse transcription polymerase chain reaction (RT-PCR) was used to validate the differentially expressed genes identified by Genechip analysis. RESULTS: A small number of differential expression genes were found in each comparison after RF EMF exposure. Through reproducible and consistent analysis, no gene or five up-regulated genes were screened out after exposure to RF EMF at SAR of 2.0 W/kg or 3.5 W/kg, respectively. However, these five genes could not be further confirmed by RT-PCR. CONCLUSION: The present study did not provide clear evidence that RF EMF exposure might distinctly change the gene expression profile in MCF-7 cells under current experimental conditions, implying that the exposure might not affect the MCF-7 cell physiology, or this cell line might be less sensitive to the RF EMF exposure.

Cell Line, Tumor↗

Patterns of p73 N-terminal isoform expression and p53 status have prognostic value in gynecological cancers.

The goal of this study was to determine whether patterns of expression profiles of p73 isoforms and of p53 mutational status are useful combinatorial biomarkers for predicting outcome in a gynecological cancer cohort. This is the first such study using matched tumor/normal tissue pairs from each patient. The median follow-up was over two years. The expression of all 5 N-terminal isoforms (TAp73, DeltaNp73, DeltaN'p73, Ex2p73 and Ex2/3p73) was measured by real-time RT-PCR and p53 status was analyzed by immunohistochemistry. TAp73, DeltaNp73 and DeltaN'p73 were significantly upregulated in tumors. Surprisingly, their range of overexpression was age-dependent, with the highest differences delta (tumor-normal) in the youngest age group. Correction of this age effect was important in further survival correlations. We used all 6 variables (five p73 isoform levels plus p53 status) as input into a principal component analysis with Varimax rotation (VrPCA) to filter out noise from non-disease related individual variability of p73 levels. Rationally selected and individually weighted principal components from each patient were then used to train a support vector machine (SVM) algorithm to predict clinical outcome. This SVM algorithm was able to predict correct outcome in 30 of the 35 patients. We use here a mathematical tool for pattern recognition that has been commonly used in e.g. microarray data mining and apply it for the first time in a prognostic model. We find that PCA/SVM is able to test a clinical hypothesis with robust statistics and show that p73 expression profiles and p53 status are useful prognostic biomarkers that differentiate patients with good vs. poor prognosis with gynecological cancers.

Age Factors↗

Sensors, medical image and signal processing. Findings from the Section on Sensor, Signal and Imaging Informatics.

OBJECTIVES: To summarize current excellent research in the field of sensor, signal and imaging informatics. METHODS: Synopsis of the articles selected for the IMIA Yearbook of Medical Informatics 2006. RESULTS: The selection process for this yearbook's section 'Sensor, signal and imaging informatics' results in six excellent articles, representing research in five different nations. We selected a cross section of the wide range of application, ranging from model based image segmentation, image retrieval and data mining, image based diagnosis assistance, bio-impedance based skin cancer screening, brain computer interfaces to MRI based computational models for fluid-structure-interactions. CONCLUSIONS: The selected articles indicate a small but meaningful extract from the research field of sensors, signal and image processing, which has a wide range of applications in medical informatics. The articles present excellent research with a possibility of having high relevance for the future in patient care.

Awards and Prizes↗

Unsupervised learning with independent component analysis can identify patterns of glaucomatous visual field defects.

PURPOSE: We previously reported the use of clustering by unsupervised learning with machine learning classifiers to segment clusters of patterns in standard automated perimetry (SAP) for glaucoma. In this study, the process of unsupervised learning by independent component analysis decomposed SAP field patterns into axes, and the information represented by these axes was evaluated. METHODS: SAP fields were obtained with the Humphrey Visual Field Analyzer on 189 normal eyes and 156 eyes with glaucomatous optic neuropathy (GON) determined by masked review with stereoscopic optic disc photos. The variational Bayesian independent component analysis mixture model (vB-ICA-mm) partitioned the SAP fields into the most informative number of clusters. Simultaneously, it learned an optimal number of maximally independent axes for each cluster. RESULTS: The most informative number of clusters was two. vB-ICA-mm placed 68.6% of the SAP fields from eyes with GON in a cluster labeled G and 98.4% of the fields from eyes with normal optic discs in a cluster labeled N. Cluster G optimally contained six axes. Post hoc analysis of patterns generated at -1 SD and +2 SD from the cluster G mean on the six axes revealed defects similar to those identified by experts as indicative of glaucoma. SAP fields associated with an axis showed increasing severity as they were located farther in the positive direction from the cluster G mean. CONCLUSIONS: vB-ICA-mm represented the SAP fields with patterns that were meaningful for glaucoma experts. This process also captured severity in the patterns uncovered. These findings should validate vB-ICA-mm as a data mining technique for new and unfamiliar complex tests.

Artificial Intelligence↗

Discovering biomedical relations utilizing the World-wide Web.

To crate a Semantic Web for Life Sciences discovering relations between biomedical entities is essential. Journals and conference proceedings represent the dominant mechanisms of reporting newly discovered biomedical interactions. The unstructured nature of such publications makes it difficult to utilize data mining or knowledge discovery techniques to automatically incorporate knowledge from these publications into the ontologies. On the other hand, since biomedical information is growing explosively, it is difficult to have human curators manually extract all the information from literature. In this paper we present techniques to automatically discover biomedical relations from the World-wide Web. For this purpose we retrieve relevant information from Web Search engines using various lexico-syntactic patterns as queries. Experiments are presented to show the usefulness of our techniques.

Classification↗