Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Evaluation of the DotMap algorithm for locating analytes of interest based on mass spectral similarity in data collected using comprehensive two-dimensional gas chromatography coupled with time-of-flight mass spectrometry.

Comprehensive two-dimensional gas chromatography coupled with time-of-flight mass spectrometry (GC x GC-TOF-MS) is a highly selective technique ideal for the analysis of complex mixtures. The instrument yields an abundance of data, with complete mass spectral scans at every time point in the GC x GC separation space. The development and application of appropriate tools for data mining is essential in making sense of the wealth of information available. An algorithm for locating analytes of interest based on mass spectral similarity in GC x GC-TOF-MS data, called DotMap, has been previously reported and is rigorously evaluated herein. A thorough investigation into the performance characteristics of DotMap, including the performance near the limit of detection and dynamic range of the algorithm as well as the capacity of the algorithm to deal with peak overlap, is investigated using jet fuel as a complex sample matrix. For instance, the algorithm can successfully identify a spiked compound at the single microg/ml level in a jet fuel sample with an overlapping interferent. The performance of the DotMap algorithm in situations with very limited mass spectral selectivity, specifically in the evaluation of spectra from isomer compounds, as well as the ability to tune DotMap results to provide the location of a specific analyte or of a class of compounds is demonstrated. The DotMap algorithm is demonstrated to be a sensitive tool that is useful in the analysis of complex mixtures and which possesses the capacity to be easily "tuned" to discern the location of analytes of interest.

Algorithms↗

A multi-step approach to time series analysis and gene expression clustering.

MOTIVATION: The huge growth in gene expression data calls for the implementation of automatic tools for data processing and interpretation. RESULTS: We present a new and comprehensive machine learning data mining framework consisting in a non-linear PCA neural network for feature extraction, and probabilistic principal surfaces combined with an agglomerative approach based on Negentropy aimed at clustering gene microarray data. The method, which provides a user-friendly visualization interface, can work on noisy data with missing points and represents an automatic procedure to get, with no a priori assumptions, the number of clusters present in the data. Cell-cycle dataset and a detailed analysis confirm the biological nature of the most significant clusters. AVAILABILITY: The software described here is a subpackage part of the ASTRONEURAL package and is available upon request from the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Artificial Intelligence↗

Photosynthetic acclimation is reflected in specific patterns of gene expression in drought-stressed loblolly pine.

Because the product of a single gene can influence many aspects of plant growth and development, it is necessary to understand how gene products act in concert and upon each other to effect adaptive changes to stressful conditions. We conducted experiments to improve our understanding of the responses of loblolly pine (Pinus taeda) to drought stress. Water was withheld from rooted plantlets of to a measured water potential of -1 MPa for mild stress and -1.5 MPa for severe stress. Net photosynthesis was measured for each level of stress. RNA was isolated from needles and used in hybridizations against a microarray consisting of 2173 cDNA clones from five pine expressed sequence tag libraries. Gene expression was estimated using a two-stage mixed linear model. Subsequently, data mining via inductive logic programming identified rules (relationships) among gene expression, treatments, and functional categories. Changes in RNA transcript profiles of loblolly pine due to drought stress were correlated with physiological data reflecting photosynthetic acclimation to mild stress or photosynthetic failure during severe stress. Analysis of transcript profiles indicated that there are distinct patterns of expression related to the two levels of stress. Genes encoding heat shock proteins, late embryogenic-abundant proteins, enzymes from the aromatic acid and flavonoid biosynthetic pathways, and from carbon metabolism showed distinctive responses associated with acclimation. Five genes shown to have different transcript levels in response to either mild or severe stress were chosen for further analysis using real-time polymerase chain reaction. The real-time polymerase chain reaction results were in good agreement with those obtained on microarrays.

Acclimatization↗

Multivariate sib-pair linkage analysis of longitudinal phenotypes by three step-wise analysis approaches.

BACKGROUND: Current statistical methods for sib-pair linkage analysis of complex diseases include linear models, generalized linear models, and novel data mining techniques. The purpose of this study was to further investigate the utility and properties of a novel pattern recognition technique (step-wise discriminant analysis) using the chromosome 10 linkage data from the Framingham Heart Study and by comparing it with step-wise logistic regression and linear regression. RESULTS: The three step-wise approaches were compared in terms of statistical significance and gene localization. Step-wise discriminant linkage analysis approach performed best; next was step-wise logistic regression; and step-wise linear regression was the least efficient because it ignored the categorical nature of disease phenotypes. Nevertheless, all three methods successfully identified the previously reported chromosomal region linked to human hypertension, marker GATA64A09. We also explored the possibility of using the discriminant analysis to detect gene x gene and gene x environment interactions. There was evidence to suggest the existence of gene x environment interactions between markers GATA64A09 or GATA115E01 and hypertension treatment and gene x gene interactions between markers GATA64A09 and GATA115E01. Finally, we answered the theoretical question "Is a trichotomous phenotype more efficient than a binary?" Unlike logistic regression, discriminant sib-pair linkage analysis might have more power to detect linkage to a binary phenotype than a trichotomous one. CONCLUSION: We confirmed our previous speculation that step-wise discriminant analysis is useful for genetic mapping of complex diseases. This analysis also supported the possibility of the pattern recognition technique for investigating gene x gene or gene x environment interactions.

Adult Children↗

Application of high-throughput computing in bioinformatics.

We describe the challenges faced when developing a Linux/PC-based cluster to apply bioinformatics algorithms to the rapidly increasing raw genomics data available. The calculations, which take around two months to complete, result in a powerful resource that can be used for data mining--most obviously for the human genome. Our current infrastructure consists of a 1314 node cluster with 1734 processors supporting both production and research. This paper highlights the problems in achieving high data throughput with such systems and shows that raw computer power is only one component of a complex problem.

Computational Biology↗

Public health, GIS, and the internet.

Internet access and use of georeferenced public health information for GIS application will be an important and exciting development for the nation's Department of Health and Human Services and other health agencies in this new millennium. Technological progress toward public health geospatial data integration, analysis, and visualization of space-time events using the Web portends eventual robust use of GIS by public health and other sectors of the economy. Increasing Web resources from distributed spatial data portals and global geospatial libraries, and a growing suite of Web integration tools, will provide new opportunities to advance disease surveillance, control, and prevention, and insure public access and community empowerment in public health decision making. Emerging supercomputing, data mining, compression, and transmission technologies will play increasingly critical roles in national emergency, catastrophic planning and response, and risk management. Web-enabled public health GIS will be guided by Federal Geographic Data Committee spatial metadata, OpenGIS Web interoperability, and GML/XML geospatial Web content standards. Public health will become a responsive and integral part of the National Spatial Data Infrastructure.

Geographic Information Systems↗

Increasing conclusiveness of metabonomic studies by chem-informatic preprocessing of capillary electrophoretic data on urinary nucleoside profiles.

Nowadays, bioinformatics offers advanced tools and procedures of data mining aimed at finding consistent patterns or systematic relationships between variables. Numerous metabolites concentrations can readily be determined in a given biological system by high-throughput analytical methods. However, such row analytical data comprise noninformative components due to many disturbances normally occurring in analysis of biological samples. To eliminate those unwanted original analytical data components advanced chemometric data preprocessing methods might be of help. Here, such methods are applied to electrophoretic nucleoside profiles in urine samples of cancer patients and healthy volunteers. The electrophoretic nucleoside profiles were obtained under following conditions: 100 mM borate, 72.5 mM phosphate, 160 mM SDS, pH 6.7; 25 kV voltage, 30 degrees C temperature; untreated fused silica capillary 70 cm effective length, 50 microm I.D. Different most advanced preprocessing tools were applied for baseline correction, denoising and alignment of electrophoretic data. That approach was compared to standard procedure of electrophoretic peak integration. The best results of preprocessing were obtained after application of the so-called correlation optimized warping (COW) to align the data. The principal component analysis (PCA) of preprocessed data provides a clearly better consistency of the nucleoside electrophoretic profiles with health status of subjects than PCA of peak areas of original data (without preprocessing).

Algorithms↗

Data management solutions for protein therapeutic research and development.

Protein therapeutics, including monoclonal antibodies, are a growing focus of drug discovery research organizations. High-throughput screening of large libraries of protein variants is therefore becoming increasingly important in R&D. As a result, there is a need to link large numbers of variant protein sequences with chemical and biological assay data. This integration will allow more efficient data mining and facilitate decision-making regarding hit identification, lead optimization and drug development. In this paper, we present an implementation in which a widely used small-molecule high-throughput screening data management system has been adapted to meet the unique needs of protein drug discovery and development.

Antibodies, Monoclonal↗

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity.

The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

Journal Article↗

Recent additions and improvements to the Onto-Tools.

The Onto-Tools suite is composed of an annotation database and six seamlessly integrated, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner and Pathway-Express. The Onto-Tools database has been expanded to include various types of data from 12 new databases. Our database now integrates different types of genomic data from 19 sequence, gene, protein and annotation databases. Additionally, our database is also expanded to include complete Gene Ontology (GO) annotations. Using the enhanced database and GO annotations, Onto-Express now allows functional profiling for 24 organisms and supports 17 different types of input IDs. Onto-Translate is also enhanced to fully utilize the capabilities of the new Onto-Tools database with an ultimate goal of providing the users with a non-redundant and complete mapping from any type of identification system to any other type. Currently, Onto-Translate allows arbitrary mappings between 29 types of IDs. Pathway-Express is a new tool that helps the users find the most interesting pathways for their input list of genes. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

Purdue ionomics information management system. An integrated functional genomics platform.

The advent of high-throughput phenotyping technologies has created a deluge of information that is difficult to deal with without the appropriate data management tools. These data management tools should integrate defined workflow controls for genomic-scale data acquisition and validation, data storage and retrieval, and data analysis, indexed around the genomic information of the organism of interest. To maximize the impact of these large datasets, it is critical that they are rapidly disseminated to the broader research community, allowing open access for data mining and discovery. We describe here a system that incorporates such functionalities developed around the Purdue University high-throughput ionomics phenotyping platform. The Purdue Ionomics Information Management System (PiiMS) provides integrated workflow control, data storage, and analysis to facilitate high-throughput data acquisition, along with integrated tools for data search, retrieval, and visualization for hypothesis development. PiiMS is deployed as a World Wide Web-enabled system, allowing for integration of distributed workflow processes and open access to raw data for analysis by numerous laboratories. PiiMS currently contains data on shoot concentrations of P, Ca, K, Mg, Cu, Fe, Zn, Mn, Co, Ni, B, Se, Mo, Na, As, and Cd in over 60,000 shoot tissue samples of Arabidopsis (Arabidopsis thaliana), including ethyl methanesulfonate, fast-neutron and defined T-DNA mutants, and natural accession and populations of recombinant inbred lines from over 800 separate experiments, representing over 1,000,000 fully quantitative elemental concentrations. PiiMS is accessible at www.purdue.edu/dp/ionomics.

Arabidopsis↗

arrayCGHbase: an analysis platform for comparative genomic hybridization microarrays.

BACKGROUND: The availability of the human genome sequence as well as the large number of physically accessible oligonucleotides, cDNA, and BAC clones across the entire genome has triggered and accelerated the use of several platforms for analysis of DNA copy number changes, amongst others microarray comparative genomic hybridization (arrayCGH). One of the challenges inherent to this new technology is the management and analysis of large numbers of data points generated in each individual experiment. RESULTS: We have developed arrayCGHbase, a comprehensive analysis platform for arrayCGH experiments consisting of a MIAME (Minimal Information About a Microarray Experiment) supportive database using MySQL underlying a data mining web tool, to store, analyze, interpret, compare, and visualize arrayCGH results in a uniform and user-friendly format. Following its flexible design, arrayCGHbase is compatible with all existing and forthcoming arrayCGH platforms. Data can be exported in a multitude of formats, including BED files to map copy number information on the genome using the Ensembl or UCSC genome browser. CONCLUSION: ArrayCGHbase is a web based and platform independent arrayCGH data analysis tool, that allows users to access the analysis suite through the internet or a local intranet after installation on a private server. ArrayCGHbase is available at http://medgen.ugent.be/arrayCGHbase/.

Base Sequence↗

Mass spectrometry-based proteomics and its application to studies of Porphyromonas gingivalis invasion and pathogenicity.

Porphyromonas gingivalis is a Gram-negative anaerobe that populates the subgingival crevice of the mouth. It is known to undergo a transition from its commensal status in healthy individuals to a highly invasive intracellular pathogen in human patients suffering from periodontal disease, where it is often the dominant species of pathogenic bacteria. The application of mass spectrometry-based proteomics to the study of P. gingivalis interactions with model host cell systems, invasion and pathogenicity is reviewed. These studies have evolved from qualitative identifications of small numbers of secreted proteins, using traditional gel-based methods, to quantitative whole cell proteomic studies using multiple dimension capillary HPLC coupled with linear ion trap mass spectrometry. It has become possible to generate a differential readout of protein expression change over the entire P. gingivalis proteome, in a manner analogous to whole genome mRNA arrays. Different strategies have been employed for generating protein level expression ratios from mass spectrometry data, including stable isotope metabolic labeling and most recently, spectral counting methods. A global view of changes in protein modification status remains elusive due to the limitations of existing computational tools for database searching and data mining. Such a view would be desirable for purposes of making global assessments of changes in gene regulation in response to host interactions during the course of adhesion, invasion and internalization. With a complete data matrix consisting of changes in transcription, protein abundance and protein modification during the course of invasion, the search for new protein drug targets would benefit from a more comprehensive understanding of these processes than what could be achieved prior to the advent of systems biology.

Bacterial Proteins↗

Analyzing Sub-Classifications of Glaucoma via SOM Based Clustering of Optic Nerve Images.

We present a data mining framework to cluster optic nerve images obtained by Confocal Scanning Laser Tomography (CSLT) in normal subjects and patients with glaucoma. We use self-organizing maps and expectation maximization methods to partition the data into clusters that provide insights into potential sub-classification of glaucoma based on morphological features. We conclude that our approach provides a first step towards a better understanding of morphological features in optic nerve images obtained from glaucoma patients and healthy controls.

Algorithms↗

Mathematical modeling of cancer: the future of prognosis and treatment.

BACKGROUND: Cancer research has undergone radical changes in the past few years. Producing information both at the basic and clinical levels is no longer the issue. Rather, how to handle this information has become the major obstacle to progress. Intuitive approaches are no longer feasible. The next big step will be to implement mathematical modeling approaches to interrogate the enormous amount of data being produced and extract useful answers (a "top-down" approach to biology and medicine). METHODS: Quantitative simulation of clinically relevant cancer situations-based on experimentally validated mathematical modeling-provides an opportunity for the researcher, and eventually the clinician, to address data and information in the context of well-formulated questions and "what if" scenarios. RESULTS AND CONCLUSIONS: At the Vanderbilt Integrative Cancer Biology Center (VICBC), we are integrating cancer researchers, oncologists, chemical and biological engineers, computational biologists, computer modelers, theoretical and applied mathematicians, and imaging scientists, in order to implement a vision for a combined web site and computational server that will be a home for our mathematical modeling of cancer invasion. The web site (www.vanderbilt.edu/VICBC/) will serve as a portal to our code, which simulates tumor growth by calculating the dynamics of individual cancer cells (an experimental "bottom-up" approach to complement the top-down model). Eventually, cancer researchers outside of Vanderbilt will be able to initiate a simulation based on providing individual cell data through a web page. We envision placing the web site and computer cluster directly in the hands of biological researchers involved in data mining and mathematical modeling. Furthermore, the web site will also contain teaching props for a new generation of biomedical researchers fluent in both mathematics and biology. This is unconventional bioinformatics: We will be incorporating biological data and functional information into a unified community-based mathematical framework. The result will be a tool for cancer modeling that will ultimately have basic research, therapeutic and educational value.

Animals↗

SPINS: standardized protein NMR storage. A data dictionary and object-oriented relational database for archiving protein NMR spectra.

Modern protein NMR spectroscopy laboratories have a rapidly growing need for an easily queried local archival system of raw experimental NMR datasets. SPINS (Standardized ProteIn Nmr Storage) is an object-oriented relational database that provides facilities for high-volume NMR data archival, organization of analyses, and dissemination of results to the public domain by automatic preparation of the header files required for submission of data to the BioMagResBank (BMRB). The current version of SPINS coordinates the process from data collection to BMRB deposition of raw NMR data by standardizing and integrating the storage and retrieval of these data in a local laboratory file system. Additional facilities include a data mining query tool, graphical database administration tools, and a NMRStar v2. 1.1 file generator. SPINS also includes a user-friendly internet-based graphical user interface, which is optionally integrated with Varian VNMR NMR data collection software. This paper provides an overview of the data model underlying the SPINS database system, a description of its implementation in Oracle, and an outline of future plans for the SPINS project.

Archives↗

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗

d-matrix - database exploration, visualization and analysis.

BACKGROUND: Motivated by a biomedical database set up by our group, we aimed to develop a generic database front-end with embedded knowledge discovery and analysis features. A major focus was the human-oriented representation of the data and the enabling of a closed circle of data query, exploration, visualization and analysis. RESULTS: We introduce a non-task-specific database front-end with a new visualization strategy and built-in analysis features, so called d-matrix. d-matrix is web-based and compatible with a broad range of database management systems. The graphical outcome consists of boxes whose colors show the quality of the underlying information and, as the name suggests, they are arranged in matrices. The granularity of the data display allows consequent drill-down. Furthermore, d-matrix offers context-sensitive categorization, hierarchical sorting and statistical analysis. CONCLUSIONS: d-matrix enables data mining, with a high level of interactivity between humans and computer as a primary factor. We believe that the presented strategy can be very effective in general and especially useful for the integration of distinct data types such as phenotypical and molecular data.

Cardiovascular Diseases↗