Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Statewide analysis of serum prostate specific antigen levels in Louisiana men without prostate cancer.

OBJECTIVES: To examine age, racial, and regional differences in serum PSA levels among men in Louisiana. METHODS: From January 1, 2001 through December 31, 2001, there were 10,012 serum PSA tests performed at Louisiana Health Care Services Division (HCSD) hospitals. Manual and electronic data mining were performed to select the earliest PSA value in those men who had multiple determinations. This PSA data file was then linked with those of the Louisiana Tumor Registry and from HCSD pathology laboratories, all matched cases were removed. Men younger than 40 years and older than 79 years were excluded from this study. The final data file contained 7,258 men, of whom 4,244 were African-Americans and 3,014 were Caucasians. Comparisons of median and geometric mean serum PSA level were made between and among races for each age-decade as well as among the hospitals to assess for racial and regional differences. RESULTS: Median PSA levels were statistically significantly higher in African-American men than in Caucasian men for each age group (p < or = 0.0002). The median PSA (ng/ml) for African-American men was 0.7, 0.9, 1.3, and 2.3 for age-decades 40-49, 50-59, 60-69, and 70-79, respectively, whereas for Caucasian men the median PSA levels were 0.8, 1.2, and 1.6 for age-decades 50-59, 60-69, and 70-79, respectively. Nonparametric analysis of variance did not demonstrate a regional pattern of PSA values among the hospitals. CONCLUSIONS: In a first statewide analysis of age and racial differences of serum PSA levels, African-American men without prostate cancer had significantly higher serum PSA levels than their age-matched Caucasian male counterparts. Additionally, there were no regional patterns of PSA values among the racial groups.

Adult↗

Design of a Multi Dimensional Database for the Archimed DataWarehouse.

The Archimed data warehouse project started in 1993 at the Geneva University Hospital. It has progressively integrated seven data marts (or domains of activity) archiving medical data such as Admission/Discharge/Transfer (ADT) data, laboratory results, radiology exams, diagnoses, and procedure codes. The objective of the Archimed data warehouse is to facilitate the access to an integrated and coherent view of patient medical in order to support analytical activities such as medical statistics, clinical studies, retrieval of similar cases and data mining processes. This paper discusses three principal design aspects relative to the conception of the database of the data warehouse: 1) the granularity of the database, which refers to the level of detail or summarization of data, 2) the database model and architecture, describing how data will be presented to end users and how new data is integrated, 3) the life cycle of the database, in order to ensure long term scalability of the environment. Both, the organization of patient medical data using a standardized elementary fact representation and the use of the multi dimensional model have proved to be powerful design tools to integrate data coming from the multiple heterogeneous database systems part of the transactional Hospital Information System (HIS). Concurrently, the building of the data warehouse in an incremental way has helped to control the evolution of the data content. These three design aspects bring clarity and performance regarding data access. They also provide long term scalability to the system and resilience to further changes that may occur in source systems feeding the data warehouse.

Databases, Factual↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

Substructure mining using elaborate chemical representation.

Substructure mining algorithms are important drug discovery tools since they can find substructures that affect physicochemical and biological properties. Current methods, however, only consider a part of all chemical information that is present within a data set of compounds. Therefore, the overall aim of our study was to enable more exhaustive data mining by designing methods that detect all substructures of any size, shape, and level of chemical detail. A means of chemical representation was developed that uses atomic hierarchies, thus enabling substructure mining to consider general and/or highly specific features. As a proof-of-concept, the efficient, multipurpose graph mining system Gaston learned substructures of any size and shape from a mutagenicity data set that was represented in this manner. From these substructures, we extracted a set of only six nonredundant, discriminative substructures that represent relevant biochemical knowledge. Our results demonstrate the individual and synergistic importance of elaborate chemical representation and mining for nonlinear substructures. We conclude that the combination of elaborate chemical representation and Gaston provides an excellent method for 2D substructure mining as this recipe systematically explores all substructures in different levels of chemical detail.

Databases as Topic↗

Bringing chemical data onto the Semantic Web.

Present chemical data storage methodologies place many restrictions on the use of the stored data. The absence of sufficient high-quality metadata prevents intelligent computer access to the data without human intervention. This creates barriers to the automation of data mining in activities such as quantitative structure-activity relationship modelling. The application of Semantic Web technologies to chemical data is shown to reduce these limitations. The use of unique identifiers and relationships (represented as uniform resource identifiers, URIs, and resource description framework, RDF) held in a triplestore provides for greater detail and flexibility in the sharing and storage of molecular structures and properties.

Journal Article↗

Computer-intensive methods in traffic safety research.

The analysis of traffic safety data archives has improved markedly with the development of procedures that are heavily dependent upon computers. Three such procedures are described here. The first procedure involves using computers to assist in the identification and correction of invalid data. The second procedure makes greater computational demands, and involves using computerized algorithms to fill in the "gaps" that typically occur in archival data when information regarding key variables is not available. The third and most computer-intensive procedure involves using data mining techniques to search archives for interesting and important relationships between variables. These procedures are illustrated using examples from data archives that describe the characteristics of traffic accidents in the USA and Australia.

Accidents, Traffic↗

SPECT electronic collimation resolution enhancement using chi-square minimization.

An electronic collimation technique is developed which utilizes the chi-square goodness-of-fit measure to filter scattered gammas incident upon a medical imaging detector. In this data mining technique, Compton kinematic expressions are used as the chi-square fitting templates for measured energy-deposition data involving multiple-interaction scatter sequences. Fit optimization is conducted using the Davidon variable metric minimization algorithm to simultaneously determine the best-fit gamma scatter angles and their associated uncertainties, with the uncertainty associated with the first scatter angle corresponding to the angular resolution precision for the source. The methodology requires no knowledge of materials and geometry. This pattern recognition application enhances the ability to select those gammas that will provide the best resolution for input to reconstruction software. Illustrative computational results are presented for a conceptual truncated-ellipsoid polystyrene position-sensitive fibre head-detector Monte Carlo model using a triple Compton scatter gamma sequence assessment for a 99mTc point source. A filtration rate of 94.3% is obtained, resulting in an estimated sensitivity approximately three orders of magnitude greater than a high-resolution mechanically collimated device. The technique improves the nominal single-scatter angular resolution by up to approximately 24 per cent as compared with the conventional analytic electronic collimation measure.

Algorithms↗

Incremental nonlinear dimensionality reduction by manifold learning.

Understanding the structure of multidimensional patterns, especially in unsupervised cases, is of fundamental importance in data mining, pattern recognition, and machine learning. Several algorithms have been proposed to analyze the structure of high-dimensional data based on the notion of manifold learning. These algorithms have been used to extract the intrinsic characteristics of different types of high-dimensional data by performing nonlinear dimensionality reduction. Most of these algorithms operate in a "batch" mode and cannot be efficiently applied when data are collected sequentially. In this paper, we describe an incremental version of ISOMAP, one of the key manifold learning algorithms. Our experiments on synthetic data as well as real world images demonstrate that our modified algorithm can maintain an accurate low-dimensional representation of the data in an efficient manner.

Algorithms↗

Mining frequent patterns in protein structures: a study of protease families.

MOTIVATION: Analysis of protein sequence and structure databases usually reveal frequent patterns (FP) associated with biological function. Data mining techniques generally consider the physicochemical and structural properties of amino acids and their microenvironment in the folded structures. Dynamics is not usually considered, although proteins are not static, and their function relates to conformational mobility in many cases. RESULTS: This work describes a novel unsupervised learning approach to discover FPs in the protein families, based on biochemical, geometric and dynamic features. Without any prior knowledge of functional motifs, the method discovers the FPs for each type of amino acid and identifies the conserved residues in three protease subfamilies; chymotrypsin and subtilisin subfamilies of serine proteases and papain subfamily of cysteine proteases. The catalytic triad residues are distinguished by their strong spatial coupling (high interconnectivity) to other conserved residues. Although the spatial arrangements of the catalytic residues in the two subfamilies of serine proteases are similar, their FPs are found to be quite different. The present approach appears to be a promising tool for detecting functional patterns in rapidly growing structure databases and providing insights in to the relationship among protein structure, dynamics and function. AVAILABILITY: Available upon request from the authors.

Algorithms↗

Technology Insight: tuning into the genetic orchestra using microarrays--limitations of DNA microarrays in clinical practice.

Scientific advances in the field of genetics and gene-expression profiling have revolutionized the concept of patient-tailored treatment. Analysis of differential gene-expression patterns across thousands of biological samples in a single experiment (as opposed to hundreds to thousands of experiments measuring the expression of one gene at a time), and extrapolation of these data to answer clinically pertinent questions such as those relating to tumor metastatic potential, can help define the best therapeutic regimens for particular patient subgroups. The use of microarrays provides a powerful technology, allowing in-depth analysis of gene-expression profiles. Currently, microarray technology is in a transition phase whereby scientific information is beginning to guide clinical practice decisions. Before microarrays qualify as a useful clinical tool, however, they must demonstrate reliability and reproducibility. The high-throughput nature of microarray experiments imposes numerous limitations, which apply to simple issues such as sample acquisition and data mining, to more controversial issues that relate to the methods of biostatistical analysis required to analyze the enormous quantities of data obtained. Methods for validating proposed gene-expression profiles and those for improving trial designs represent some of the recommendations that have been suggested. This Review focuses on the limitations of microarray analysis that are continuously being recognized, and discusses how these limitations are being addressed.

Clinical Trials as Topic↗

Applying the SOM model to text classification according to register and stylistic content.

We report on the application of the Self-Organizing Map (SOM) classification method to the task of categorizing texts according to their register and the style of their author. The SOM has been selected as its performance in various data-mining applications has been found to be highly successful. Here, the method is evaluated against the task of clustering textual data which are corpora of texts written in the Greek language; the parameters used depict linguistically important structural properties of the texts. The experiments reported indicate that the SOM results are equivalent to those generated by statistical methods.

Algorithms↗

Upregulation of the tumor suppressor gene menin in hepatocellular carcinomas and its significance in fibrogenesis.

The molecular mechanisms underlying the progression of cirrhosis toward hepatocellular carcinoma were investigated by a combination of DNA microarray analysis and literature data mining. By using a microarray screening of suppression subtractive hybridization cDNA libraries, we first analyzed genes differentially expressed in tumor and nontumor livers with cirrhosis from 15 patients with hepatocellular carcinomas. Seventy-four genes were similarly recovered in tumor (57.8% of differentially expressed genes) and adjacent nontumor tissues (64% of differentially expressed genes) compared with histologically normal livers. Gene ontology analyses revealed that downregulated genes (n = 35) were mostly associated with hepatic functions. Upregulated genes (n = 39) included both known genes associated with extracellular matrix remodeling, cell communication, metabolism, and post-transcriptional regulation gene (e.g., ZFP36L1), as well as the tumor suppressor gene menin (multiple endocrine neoplasia type 1; MEN1). MEN1 was further identified as an important node of a regulatory network graph that integrated array data with array-independent literature mining. Upregulation of MEN1 in tumor was confirmed in an independent set of samples and associated with tumor size (P = .016). In the underlying liver with cirrhosis, increased steady-state MEN1 mRNA levels were correlated with those of collagen alpha2(I) mRNA (P < .01). In addition, MEN1 expression was associated with hepatic stellate cell activation during fibrogenesis and involved in transforming growth factor beta (TGF-beta)-dependent collagen alpha2(I) regulation. In conclusion, menin is a key regulator of gene networks that are activated in fibrogenesis associated with hepatocellular carcinoma through the modulation of TGF-beta response.

Carcinoma, Hepatocellular↗

Annotating proteins by mining protein interaction networks.

MOTIVATION: In general, most accurate gene/protein annotations are provided by curators. Despite having lesser evidence strengths, it is inevitable to use computational methods for fast and a priori discovery of protein function annotations. This paper considers the problem of assigning Gene Ontology (GO) annotations to partially annotated or newly discovered proteins. RESULTS: We present a data mining technique that computes the probabilistic relationships between GO annotations of proteins on protein-protein interaction data, and assigns highly correlated GO terms of annotated proteins to non-annotated proteins in the target set. In comparison with other techniques, probabilistic suffix tree and correlation mining techniques produce the highest prediction accuracy of 81% precision with the recall at 45%. AVAILABILITY: Code is available upon request. Results and used materials are available online at http://kirac.case.edu/PROTAN.

Amino Acid Sequence↗

CDtool-an integrated software package for circular dichroism spectroscopic data processing, analysis, and archiving.

CDtool is a software package written to facilitate circular dichroism (CD) spectroscopic studies on both conventional lab-based instruments and synchrotron beamlines. It takes format-independent input data from any type of CD instrument, enables a wide range of standard and advanced processing methods, and, in a single user-friendly graphics-based package, takes raw data through the entire processing procedure and, importantly, uses data-mining techniques to retain in the final output all the information associated with the processing. It permits the facile comparison of data obtained from different instruments without the need for reformatting and displays it in graphical formats suitable for publication. It also includes the ability to automatically archive the processed data. This latter feature may be especially useful in light of recent funding institution directives with regard to data sharing and archiving and requirements for "good practice" and "traceability" within the pharmaceutical industry. In addition, CDtool includes a means of interfacing with protein data bank coordinate files and calculating secondary structures from them using alternate definitions and algorithms. This feature, along with a function that permits the facile production of new reference databases, enables the creation of specialized databases for secondary structural analyses of specific types of proteins. Thus the CDtool software not only enables rapid data processing and analyses but also includes many enhanced features not available in other CD data processing/analysis packages.

Archives↗

COPASAAR--a database for proteomic analysis of single amino acid repeats.

BACKGROUND: Single amino acid repeats make up a significant proportion in all of the proteomes that have currently been determined. They have been shown to be functionally and medically significant, and are associated with cancers and neuro-degenerative diseases such as Huntington's Chorea, where a poly-glutamine repeat is responsible for causing the disease. The COPASAAR database is a new tool to facilitate the rapid analysis of single amino acid repeats at a proteome level. The database aims to simplify the comparison of repeat distributions between proteomes in order to provide a better understanding of their function and evolution. RESULTS: A comparative analysis of all proteomes in the database (currently 244) shows that single amino acid repeats account for about 12-14% of the proteome of any given species. They are more common in eukaryotes (14%) than in either archaea or bacteria (both 13%). Individual analyses of proteomes show that long single amino acid repeats (6+ residues) are much more common in the Eukaryotes and that longer repeats are usually made up of hydrophilic amino acids such as glutamine, glutamic acid, asparagine, aspartic acid and serine. CONCLUSION: COPASAAR is a useful tool for comparative proteomics that provides rapid access to amino acid repeat data that can be readily data-mined. The COPASAAR database can be queried at the kingdom, proteome or individual protein level. As the amount of available proteome data increases this will be increasingly important in order to automate proteome comparison. The insights gained from these studies will give a better insight into the evolution of protein sequence and function.

Algorithms↗

The systematic functional characterisation of Xq28 genes prioritises candidate disease genes.

BACKGROUND: Well known for its gene density and the large number of mapped diseases, the human sub-chromosomal region Xq28 has long been a focus of genome research. Over 40 of approximately 300 X-linked diseases map to this region, and systematic mapping, transcript identification, and mutation analysis has led to the identification of causative genes for 26 of these diseases, leaving another 17 diseases mapped to Xq28, where the causative gene is still unknown. To expedite disease gene identification, we have initiated the functional characterisation of all known Xq28 genes. RESULTS: By using a systematic approach, we describe the Xq28 genes by RNA in situ hybridisation and Northern blotting of the mouse orthologs, as well as subcellular localisation and data mining of the human genes. We have developed a relational web-accessible database with comprehensive query options integrating all experimental data. Using this database, we matched gene expression patterns with affected tissues for 16 of the 17 remaining Xq28 linked diseases, where the causative gene is unknown. CONCLUSION: By using this systematic approach, we have prioritised genes in linkage regions of Xq28-mapped diseases to an amenable number for mutational screens. Our database can be queried by any researcher performing highly specified searches including diseases not listed in OMIM or diseases that might be linked to Xq28 in the future.

Animals↗

Creating experimental analogs with available clinical information: credible alternatives to "gold-standard" experiments?

Comparison of the implementation and findings of a "gold standard" evaluation of social work intervention and its experimental analog based on available clinical information illustrates the strengths and weaknesses of each. From a practice-research integration perspective, however, "clinical data-mining" may be a credible alternative to randomized controlled experiments.

Data Collection↗

Efficient streaming text clustering.

Clustering data streams has been a new research topic, recently emerged from many real data mining applications, and has attracted a lot of research attention. However, there is little work on clustering high-dimensional streaming text data. This paper combines an efficient online spherical k-means (OSKM) algorithm with an existing scalable clustering strategy to achieve fast and adaptive clustering of text streams. The OSKM algorithm modifies the spherical k-means (SPKM) algorithm, using online update (for cluster centroids) based on the well-known Winner-Take-All competitive learning. It has been shown to be as efficient as SPKM, but much superior in clustering quality. The scalable clustering strategy was previously developed to deal with very large databases that cannot fit into a limited memory and that are too expensive to read/scan multiple times. Using the strategy, one keeps only sufficient statistics for history data to retain (part of) the contribution of history data and to accommodate the limited memory. To make the proposed clustering algorithm adaptive to data streams, we introduce a forgetting factor that applies exponential decay to the importance of history data. The older a set of text documents, the less weight they carry. Our experimental results demonstrate the efficiency of the proposed algorithm and reveal an intuitive and an interesting fact for clustering text streams-one needs to forget to be adaptive.

Algorithms↗