Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Development of an open source laboratory information management system for 2-D gel electrophoresis-based proteomics workflow.

BACKGROUND: In the post-genome era, most research scientists working in the field of proteomics are confronted with difficulties in management of large volumes of data, which they are required to keep in formats suitable for subsequent data mining. Therefore, a well-developed open source laboratory information management system (LIMS) should be available for their proteomics research studies. RESULTS: We developed an open source LIMS appropriately customized for 2-D gel electrophoresis-based proteomics workflow. The main features of its design are compactness, flexibility and connectivity to public databases. It supports the handling of data imported from mass spectrometry software and 2-D gel image analysis software. The LIMS is equipped with the same input interface for 2-D gel information as a clickable map on public 2DPAGE databases. The LIMS allows researchers to follow their own experimental procedures by reviewing the illustrations of 2-D gel maps and well layouts on the digestion plates and MS sample plates. CONCLUSION: Our new open source LIMS is now available as a basic model for proteome informatics, and is accessible for further improvement. We hope that many research scientists working in the field of proteomics will evaluate our LIMS and suggest ways in which it can be improved.

Computational Biology↗

There's gold in them thar' databases.

Some health care organizations are using sophisticated data mining applications to unearth hidden truths buried in their online clinical and financial information. But the lack of a standard clinical vocabulary and standard work processes is an obstacle CIOs must blast through to reach their treasure.

Data Interpretation, Statistical↗

A comprehensive approach to the analysis of matrix-assisted laser desorption/ionization-time of flight proteomics spectra from serum samples.

For our analysis of the data from the First Annual Proteomics Data Mining Conference, we attempted to discriminate between 24 disease spectra (group A) and 17 normal spectra (group B). First, we processed the raw spectra by (i) correcting for additive sinusoidal noise (periodic on the time scale) affecting most spectra, (ii) correcting for the overall baseline level, (iii) normalizing, (iv) recombining fractions, and (v) using variable-width windows for data reduction. Also, we identified a set of polymeric peaks (at multiples of 180.6 Da) that is present in several normal spectra (B1-B8). After data processing, we found the intensities at the following mass to charge (m/z) values to be useful discriminators: 3077, 12 886 and 74 263. Using these values, we were able to achieve an overall classification accuracy of 38/41 (92.6%). Perfect classification could be achieved by adding two additional peaks, at 2476 and 6955. We identified these values by applying a genetic algorithm to a filtered list of m/z values using Mahalanobis distance between the group means as a fitness function.

Blood Proteins↗

International consensus on preliminary definitions of improvement in adult and juvenile myositis.

OBJECTIVE: To use a core set of outcome measures to develop preliminary definitions of improvement for adult and juvenile myositis as composite end points for therapeutic trials. METHODS: Twenty-nine experts in the assessment of myositis achieved consensus on 102 adult and 102 juvenile paper patient profiles as clinically improved or not improved. Two hundred twenty-seven candidate definitions of improvement were developed using the experts' consensus ratings as a gold standard and their judgment of clinically meaningful change in the core set of measures. Seventeen additional candidate definitions of improvement were developed from classification and regression tree analysis, a data-mining decision tree tool analysis. Six candidate definitions specifying percentage change or raw change in the core set of measures were developed using logistic regression analysis. Adult and pediatric working groups ranked the 13 top-performing candidate definitions for face validity, clinical sensibility, and ease of use, in which the sensitivity and specificity were >/=75% in adult, pediatric, and combined data sets. Nominal group technique was used to facilitate consensus formation. RESULTS: The definition of improvement (common to the adult and pediatric working groups) that ranked highest was 3 of any 6 of the core set measures improved by >/=20%, with no more than 2 worse by >/=25% (which could not include manual muscle testing to assess strength). Five and 4 additional preliminary definitions of improvement for adult and juvenile myositis, respectively, were also developed, with several definitions common to both groups. Participants also agreed to prospectively test 6 logistic regression definitions of improvement in clinical trials. CONCLUSION: Consensus preliminary definitions of improvement were developed for adult and juvenile myositis, and these incorporate clinically meaningful change in all myositis core set measures in a composite end point. These definitions require prospective validation, but they are now proposed for use as end points in all myositis trials.

Adult↗

[New knowledge derived from measurement of gene expression with the DNA microarray method].

BACKGROUND: The cDNA microarray method offers the first possibility of obtaining a global understanding of biological processes in living organisms, by simultaneous read-outs of tens of thousands of mRNAs. Initial experiments suggest that genes with similar function have similar expression patterns. MATERIAL AND METHODS: Understanding this level of biological complexity will, however, require completely new approaches to data analysis. Computer science methods, such as data mining and knowledge discovery, can synthesize interpretable if-then rules that model the relation between gene expressions and functions and use the rules to classify unknown genes. The huge body of existing biological and medical knowledge makes it necessary to develop methods for extracting knowledge from such repositories. RESULTS: Models of relations between gene expressions and gene functions in a data set from a publicly available source are synthesized semiautomatically and applied to classify unknown genes. Encouraging results have been achieved. The method is applied in the analysis of data from our microarray system which has recently become operational. INTERPRETATION: The principles are of general importance and will be used to evaluate a wide range of complex data sets like decision support in clinical medicine, for situations in which physicians need to handle a large volume of data for each patient.

Computational Biology↗

Causal relationship between educational attainment and the occurrence of venous thromboembolism.

BACKGROUND: The association between educational attainment (EA) and arterial thrombotic disease has been reported, but the causal relationship between EA and venous thromboembolism (VTE) is not clear. We aimed to assess the causal effect of EA on VTE using the two-sample mendelian randomization (MR) method. METHODS: Data mining was conducted on the genome wide association studies (GWAS), with exposure factor EA and outcome factor VTE. Two-sample Mendelian Randomization (TSMR) analysis was conducted, with the results obtained from the random effects inverse variance weighted method (IVW). Use the MR-Egger method for pleiotropy analysis and leave one method for sensitivity analysis to verify the reliability of the data. RESULTS: Genetically predicted decreased EA was associated with a decreased risk of VTE in both the FinnGen consortium and UK Biobank (FinnGen-VTE: OR = 0.848; 95% CI 0.776-0.927; P = 2.84 × 10-4; UKB-VTE OR = 0.996; 95% CI 0.994-0.999; P = 0.008) under a multiplicative random-effects IVW model. Results were consistent in all sensitivity analyses and no horizontal pleiotropy was detected. CONCLUSIONS: The MR technique instructed a potential inverse causative relationship between EA and occurrence of VTE. Therefore, patients with low EA should be more vigilant about the occurrence of VTE.

Venous Thromboembolism↗

Designing a framework of intelligent information processing for dentistry administration data.

OBJECTIVES: This study was designed to test a cumulative view of current data in the clinical database at the Faculty of Dentistry, Dalhousie University. We planned to examine associations among demographic factors and treatments. METHODS: Three tables were selected from the database of the faculty: patient, treatment and procedures. All fields and record numbers in each table were documented. Data was explored using SQL server and Visual Basic and then cleaned by removing incongruent fields. After transformation, a data warehouse was created. This was imported to SQL analysis services manager to create an OLAP (Online Analytic Process) cube. RESULTS: The multidimensional model used for access to data was created using a star schema. Treatment count was the measurement variable. Five dimensions--date, postal code, gender, age group and treatment categories--were used to detect associations. Another data warehouse of 8 tables (international tooth code # 1-8) was created and imported to SAS enterprise miner to complete data mining. Association nodes were used for each table to find sequential associations and minimum criteria were set to 2% of cases. Findings of this study confirmed most assumptions of treatment planning procedures. There were some small unexpected patterns of clinical interest. Further developments are recommended to create predictive models. CONCLUSIONS: Recent improvements in information technology offer numerous advantages for conversion of raw data from faculty databases to information and subsequently to knowledge. This knowledge can be used by decision makers, managers, and researchers to answer clinical questions, affect policy change and determine future research needs.

Adolescent↗

Identification of protein modifications using MS/MS de novo sequencing and the OpenSea alignment algorithm.

Algorithms that can robustly identify post-translational protein modifications from mass spectrometry data are needed for data-mining and furthering biological interpretations. In this study, we determined that a mass-based alignment algorithm (OpenSea) for de novo sequencing results could identify post-translationally modified peptides in a high-throughput environment. A complex digest of proteins from human cataractous lens, a tissue containing a high abundance of modified proteins, was analyzed using two-dimensional liquid chromatography, and data was collected on both high and low mass accuracy instruments. The data were analyzed using automated de novo sequencing followed by OpenSea mass-based sequence alignment. A total of 80 modifications were detected, 36 of which were previously unreported in the lens. This demonstrates the potential to identify large numbers of known and previously unknown protein modifications in a given tissue using automated data processing algorithms such as OpenSea.

Aged↗

Projective ART for clustering data sets in high dimensional spaces.

A new neural network architecture (PART) and the resulting algorithm are proposed to find projected clusters for data sets in high dimensional spaces. The architecture is based on the well known ART developed by Carpenter and Grossberg, and a major modification (selective output signaling) is provided in order to deal with the inherent sparsity in the full space of the data points from many data-mining applications. This selective output signaling mechanism allows the signal generated in a node in the input layer to be transmitted to a node in the clustering layer only when the signal is similar to the top-down weight between the two nodes and, hence, PART focuses on dimensions where information can be found. Illustrative examples are provided, simulations on high dimensional synthetic data and comparisons with Fuzzy ART module and PROCLUS are also reported.

Algorithms↗

MAGIIC-PRO: detecting functional signatures by efficient discovery of long patterns in protein sequences.

This paper presents a web service named MAGIIC-PRO, which aims to discover functional signatures of a query protein by sequential pattern mining. Automatic discovery of patterns from unaligned biological sequences is an important problem in molecular biology. MAGIIC-PRO is different from several previously established methods performing similar tasks in two major ways. The first remarkable feature of MAGIIC-PRO is its efficiency in delivering long patterns. With incorporating a new type of gap constraints and some of the state-of-the-art data mining techniques, MAGIIC-PRO usually identifies satisfied patterns within an acceptable response time. The efficiency of MAGIIC-PRO enables the users to quickly discover functional signatures of which the residues are not from only one region of the protein sequences or are only conserved in few members of a protein family. The second remarkable feature of MAGIIC-PRO is its effort in refining the mining results. Considering large flexible gaps improves the completeness of the derived functional signatures. The users can be directly guided to the patterns with as many blocks as that are conserved simultaneously. In this paper, we show by experiments that MAGIIC-PRO is efficient and effective in identifying ligand-binding sites and hot regions in protein-protein interactions directly from sequences. The web service is available at http://biominer.bime.ntu.edu.tw/magiicpro and a mirror site at http://biominer.cse.yzu.edu.tw/magiicpro.

Binding Sites↗

NEOBASE: databasing the neocortical microcircuit.

Mammals adapt to a rapidly changing world because of the sophisticated perceptual and cognitive function enabled by the neocortex. The neocortex, which has expanded to constitute nearly 80% of the human brain seems to have arisen from repeated duplication of a stereotypical template of neurons and synaptic circuits with subtle specializations in different brain regions and species. Determining the design and function of this microcircuitry is therefore of paramount importance to understanding normal and abnormal higher brain function. Recent advances in recording synaptically-coupled neurons has allowed rapid dissection of the neocortical microcircuitry thus yielding a massive amount of quantitative anatomical, electrical and gene expression data on the neurons and the synaptic circuits that connect the neurons. Due to the availability of the above mentioned data, it has now become imperative to database the neurons of the microcircuit and their synaptic connections. The NEOBASE project, aims to archive the neocortical microcircuit data in a manner that facilitates development of advanced data mining applications, statistical and bioinformatics analyses tools, custom microcircuit builders, and visualization and simulation applications. The database architecture is based on ROOT, a software environment that allows the construction of an object oriented database with numerous relational capabilities. The proposed architecture allows construction of a database that closely mimics the architecture of the real microcircuit, which facilitates the interface with virtually any application, allows for data format evolution, and aims for full interoperability with other databases. NEOBASE will provide an important resource and research tool for studying the microcircuit basis of normal and abnormal neocortical function. The database will be available to local as well as remote users using Grid based tools and technologies.

Animals↗

Mining time dependency patterns in clinical pathways.

Clinical pathways are widely adopted by many large hospitals around the world in order to provide high-quality patient treatment and reduce the length of hospital stay of each patient. The development of clinical pathways is a lengthy process, and may require the collaboration among physicians, nurses, and staffs in a hospital. However, the individual differences cause great variances in the execution of clinical pathways. It calls for a more dynamic and adaptive process to improve the performance of clinical pathways. This paper reports a data mining technique we have developed to discover the time dependency pattern of clinical pathways for managing brain stroke. The mining of time dependency pattern is to discover patterns of process execution sequences and to identify the dependent relation between activities in a majority of cases. By obtaining the time dependency patterns, we can predict the paths for new patients when he/she is admitted into a hospital; in turn, the health care procedure will be more effective and efficient.

Algorithms↗

A methodology to explain neural network classification.

Neural networks are still frustrating tools in the data mining arsenal. They exhibit excellent modelling performance, but do not give a clue about the structure of their models. We propose a methodology to explain the classification obtained by a multilayer perceptron. We introduce the concept of 'causal importance' and define a saliency measurement allowing the selection of relevant variables. Once the model is trained with the relevant variables only, we define a clustering of the data built from the hidden layer representation. Combining the saliency and the causal importance on a cluster by cluster basis allows an interpretation of the neural network classifier to be built. We illustrate the performances of this methodology on three benchmark datasets.

Classification↗

BioWarehouse: a bioinformatics database warehouse toolkit.

BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.

Computational Biology↗

T-SMmOTE: tweaked synthetic majority minority oversampling technique for data scarcity issue in multi omics studies.

MOTIVATION: Multiomics data offer a rich data mine for modeling complex as well as day-to-day diseases, but their practical deployment is constrained by the limited sample availability. To this end, generating synthetic samples is a viable remedy. Extant schemes operating along this line, however, are mostly limited to augmenting the minority class in imbalanced datasets and often produce synthetic samples that lack sufficient diversity and fail to faithfully capture the underlying data distribution. As a result, the full potential of synthetic augmentation in multi-omics learning remains underexplored. The aim is to address the data scarcity problem in multi-omics domain. We propose a synthetic oversampling framework, which is dedicated to addressing overall data scarcity in multi-omics datasets and the lack of diversity in synthetic samples. Contrary to conventional methods that restrict augmentation to minority classes and rely on interpolation of two neighbors, our method generates diverse yet distribution-aligned synthetic samples by interpolating three neighbors and extends this augmentation paradigm to the majority class. The framework first balances the dataset by generating synthetic minority samples, and subsequently augments the balanced dataset by oversampling both majority and minority classes. RESULTS: Empirical evaluation on multi-omics data obtained from three heterogeneous health scenarios-inflammatory bowel disease, multi-organ dysfunction syndrome, and colorectal cancer-substantiates the utility of the proposed scheme in improving the predictive performance. The models trained on T-SMmOTE-augmented data achieve higher Matthews correlation coefficient values, along with improvedscores for both majority and minority classes. Notably, oversampling of the majority class improves the cognition of the minority class as well. We also explore the consistency of the class distributions between the original and augmented class-specific datasets. These findings confirm the capability of our scheme to learn from small, high-dimensional multi-omics datasets and highlight its potential for non-invasive disease detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/payelu/TSMm.

Journal Article↗

Concepts and possibilities in forensic intelligence.

Forensic intelligence can be viewed as comprising two parts, one directly concerning intelligence delivery in forensic casework, the other considering performance aspects of forensic work, loosely termed here as business intelligence. Forensic casework can be viewed as processes that produce an intelligence product useful to police investigations. Traditionally, forensic intelligence production has been confined to discipline-specific activity. This paper examines the concepts, processes and intelligence products delivered in forensic casework, the information repositories available from forensic examinations, and ways to produce within- and across-discipline casework correlations by using information technology to capitalise on the information sets available. Such analysis presents opportunities to improve forensic intelligence services as well as challenges for technical solutions to deliver appropriate data-mining capabilities for available information sets, such as digital photographs. Business intelligence refers primarily to examination of efficiency and effectiveness of forensic service delivery. This paper discusses measures of forensic activity and their relationship to crime outcomes as a measure of forensic effectiveness.

Data Collection↗

GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.

GeneCards is an automatically mined database of human genes that strives to create, along with its auxiliary databases--GeneLoc, GeneNote and GeneAnnot--the most inclusive resource of gene-centered information of the human genome. GeneTide, the Gene Terra Incognita Discovery Endeavor (http://genecards.weizmann.ac.il/genetide/), the newest addition to this family, is a transcriptome-focused database which aims to enhance GeneCards with additional expressed sequence tag (EST)-based genes. This is achieved by comprehensively mapping >85% of the approximately 5.6 million human ESTs currently available at dbEST to known genes by means of data mining and integration of genomic resources including UniGene, DoTS, AceView and in-house resources. GeneTide thus creates comprehensive links between ESTs and GeneCards genes. Furthermore, groups of unassociated transcripts serve as a basis for defining novel EST-based GeneCards Candidates (EGCs). These EGCs, nearly 25,000 of which were defined in version 0.3 of GeneTide, are further annotated with various parameters, including splicing evidence and expression data extracted from the GeneNote database, to determine their validity as possible de novo genes.

Databases, Genetic↗

Using endophenotypes for pathway clusters to map complex disease genes.

Nature determines the complexity of disease etiology and the likelihood of revealing disease genes. While culprit genes for many monogenic diseases have been successfully unraveled, efforts to map major complex disease genes have not been as productive as hoped. The conceptual framework currently adopted to deal with the heterogeneous nature of complex diseases focuses on using homogeneous internal features of the disease phenotype for mapping. However, phenotypic homogeneity does not equal genotypic homogeneity. In this report, we advocate working with well-measured phenotypes portrayed by amounts of transcripts and activities of gene products or their metabolites, which are pertinent to relatively small pathway clusters. Reliable and controlled measures for oligogenic traits resulting from proper dissection efforts may enhance statistical power. The large amounts of information obtained on gene and protein expression from technological advances can add to the power of gene finding, particularly for diseases with unclear etiology. Data-mining tools for dimension reduction can assist biologists to reveal novel molecular endophenotypes. However, there are still hurdles to overcome, including high cost, relatively poor reproducibility and comparability among platforms, the cross-sectional nature of the information, and the accessibility of human tissues. Concerted efforts are required to carry out large-scale prospective studies that are integrated at the levels of phenotype characterization, high throughput experimental techniques, data analyses, and beyond.

Chromosome Mapping↗