Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

DEPD: a novel database for differentially expressed proteins.

SUMMARY: The Differentially Expressed Protein Database was designed to store the output of comparative proteomics studies and provides a publicly available query and analysis platform for data mining. The database contains information about more than 3000 differentially expressed proteins (DEPs) manually extracted from the published literature, including relevant biological, experimental and methodological elements. Tools for visualization and functional analysis of DEPs are provided via a user-friendly webinterface. AVAILABILITY: http://protchem.hunnu.edu.cn/depd/.

Computational Biology↗

A geometric invariant-based framework for the analysis of protein conformational space.

MOTIVATION: Characterization of the restricted nature of the protein local conformational space has remained a challenge, thereby necessitating a computationally expensive conformational search in protein modeling. Moreover, owing to the lack of unilateral structural descriptors, conventional data mining techniques, such as clustering and classification, have not been applied in protein structure analysis. RESULTS: We first map the local conformations in a fixed dimensional space by using a carefully selected suite of geometric invariants (GIs) and then reduce the number of dimensions via principal component analysis (PCA). Distribution of the conformations in the space spanned by the first four PCs is visualized as a set of conditional bivariate probability distribution plots, where the peaks correspond to the preferred conformations. The locations of the different canonical structures in the PC-space have been interpreted in the context of the weights of the GIs to the first four PCs. Clustering of the available conformations reveals that the number of preferred local conformations is several orders of magnitude smaller than that suggested previously. SUPPLEMENTARY INFORMATION: www.it.iitb.ac.in/~ashish/bioinfo2005/.

Algorithms↗

Athena: a resource for rapid visualization and systematic analysis of Arabidopsis promoter sequences.

SUMMARY: To better understand the regulatory networks that control plant gene expression, tools are needed to systematically analyze and visualize promoter regulatory sequences in Arabidopsis thaliana. We have developed the Athena database, which contains 30,067 predicted Arabidopsis promoter sequences and consensus sequences for 105 previously characterized transcription factor (TF) binding sites. Athena provides four novel tools to facilitate the analysis of promoter sequences: a promoter visualization tool to enable the rapid inspection of key regulatory sequences in multiple promoters; a TF binding site enrichment tool to identify statistically over-represented TF sites occurring in a user-selected subset of promoters; a data-mining tool to rapidly select promoter sequences containing the specified combination of TF binding sites; and a tool to display the distribution of TF binding site positions in a selected set of promoter sequences.

Arabidopsis↗

Detection of gene copy number changes in CGH microarrays using a spatially correlated mixture model.

MOTIVATION: Comparative genomic hybridization array experiments that investigate gene copy number changes present new challenges for statistical analysis and call for methods that incorporate spatial dependence between sequences along the chromosome. For this purpose, we propose a novel method called CGHmix. It is based on a spatially structured mixture model with three states corresponding to genomic sequences that are either unmodified, deleted or amplified. Inference is performed in a Bayesian framework. From the output, posterior probabilities of belonging to each of the three states are estimated for each genomic sequence and used to classify them. RESULTS: Using simulated data, CGHmix is validated and compared with both a conventional unstructured mixture model and with a recently proposed data mining method. We demonstrate the good performance of CGHmix for classifying copy number changes. In addition, the method provides a good estimate of the false discovery rate. We also present the analysis of a cancer related dataset. SUPPLEMENTARY INFORMATION: http://www.bgx.org.uk/papers.html

Algorithms↗

Global dynamics of biological systems from time-resolved omics experiments.

The emergent properties of biological systems, organized around complex networks of irregularly connected elements, limit the applications of the direct scientific method to their study. The current lack of knowledge opens new perspectives to the inverse scientific paradigm where observations are accumulated and analysed by advanced data-mining techniques to enable a better understanding and the formulation of testable hypotheses about the structure and functioning of these systems. The current technology allows for the wide application of omics analytical methods in the determination of time-resolved molecular profiles of biological samples. Here it is proposed that the theory of dynamical systems could be the natural framework for the proper analysis and interpretation of such experiments. A new method is described, based on the techniques of non-linear time series analysis, which is providing a global view on the dynamics of biological systems probed with time-resolved omics experiments.

Algorithms↗

caGrid: design and implementation of the core architecture of the cancer biomedical informatics grid.

MOTIVATION: The complexity of cancer is prompting researchers to find new ways to synthesize information from diverse data sources and to carry out coordinated research efforts that span multiple institutions. There is a need for standard applications, common data models, and software infrastructure to enable more efficient access to and sharing of distributed computational resources in cancer research. To address this need the National Cancer Institute (NCI) has initiated a national-scale effort, called the cancer Biomedical Informatics Grid (caBIGtrade mark), to develop a federation of interoperable research information systems. RESULTS: At the heart of the caBIG approach to federated interoperability effort is a Grid middleware infrastructure, called caGrid. In this paper we describe the caGrid framework and its current implementation, caGrid version 0.5. caGrid is a model-driven and service-oriented architecture that synthesizes and extends a number of technologies to provide a standardized framework for the advertising, discovery, and invocation of data and analytical resources. We expect caGrid to greatly facilitate the launch and ongoing management of coordinated cancer research studies involving multiple institutions, to provide the ability to manage and securely share information and analytic resources, and to spur a new generation of research applications that empower researchers to take a more integrative, trans-domain approach to data mining and analysis. AVAILABILITY: The caGrid version 0.5 release can be downloaded from https://cabig.nci.nih.gov/workspaces/Architecture/caGrid/. The operational test bed Grid can be accessed through the client included in the release, or through the caGrid-browser web application http://cagrid-browser.nci.nih.gov.

Biomarkers, Tumor↗

A scalable method for integration and functional analysis of multiple microarray datasets.

MOTIVATION: The diverse microarray datasets that have become available over the past several years represent a rich opportunity and challenge for biological data mining. Many supervised and unsupervised methods have been developed for the analysis of individual microarray datasets. However, integrated analysis of multiple datasets can provide a broader insight into genetic regulation of specific biological pathways under a variety of conditions. RESULTS: To aid in the analysis of such large compendia of microarray experiments, we present Microarray Experiment Functional Integration Technology (MEFIT), a scalable Bayesian framework for predicting functional relationships from integrated microarray datasets. Furthermore, MEFIT predicts these functional relationships within the context of specific biological processes. All results are provided in the context of one or more specific biological functions, which can be provided by a biologist or drawn automatically from catalogs such as the Gene Ontology (GO). Using MEFIT, we integrated 40 Saccharomyces cerevisiae microarray datasets spanning 712 unique conditions. In tests based on 110 biological functions drawn from the GO biological process ontology, MEFIT provided a 5% or greater performance increase for 54 functions, with a 5% or more decrease in performance in only two functions.

Algorithms↗

Neural network prediction of peptide separation in strong anion exchange chromatography.

MOTIVATION: The still emerging combination of technologies that enable description and characterization of all expressed proteins in a biological system is known as proteomics. Although many separation and analysis technologies have been employed in proteomics, it remains a challenge to predict peptide behavior during separation processes. New informatics tools are needed to model the experimental analysis method that will allow scientists to predict peptide separation and assist with required data mining steps, such as protein identification. RESULTS: We developed a software package to predict the separation of peptides in strong anion exchange (SAX) chromatography using artificial neural network based pattern classification techniques. A multi-layer perceptron is used as a pattern classifier and it is designed with feature vectors extracted from the peptides so that the classification error is minimized. A genetic algorithm is employed to train the neural network. The developed system was tested using 14 protein digests, and the sensitivity analysis was carried out to investigate the significance of each feature. AVAILABILITY: The software and testing results can be downloaded from ftp://ftp.bbc.purdue.edu.

Algorithms↗

Designer microarrays: from soup to nuts.

The recognition that multigene mechanisms control the pathways determining the aging process renders gene screening a necessary skill for biogerontologists. In the past few years, this task has become much more accessible, with the advent of DNA chip technology. Most commercially available microarrays are designed with prefixed templates of genes of general interest, allowing investigators little freedom of choice in attempting to focus gene screening on a particular thematic pathway of interest. This report describes our "designer microarray" approach as a next generation of DNA chips, allowing individual investigators to engage in gene screening with a user friendly, do-it-yourself approach, from designing the probe templates to data mining. The end result is the ability to use microarrays as a platform for versatile gene discovery.

Gene Expression Profiling↗

Lay person-based screening for early detection of Alzheimer's disease: development and validation of an instrument.

Symptoms of cognitive impairment reported to telephone interviewers by caregivers of 272 patients were analyzed with respect to research diagnoses of dementia. All patients received neuropsychological evaluation for establishing the research diagnoses. A data mining program that used machine learning algorithms produced an optimized binary decision tree for differentiating patient groups according to all available information. The results of this analysis were used to help four dementia experts create a dementia screening instrument amenable to application and scoring by nonclinical personnel. The validity of the resulting instrument was then evaluated in an independent sample of 103 patients administered neuropsychological testing within the previous 60 days. The psychometric properties of the empirically derived scale and its performance for discriminating control from probable or possible Alzheimer's patients indicate strong potential for use as a dementia screener for the general population.

Adult↗

Time-course metabolic profiling in Arabidopsis thaliana cell cultures after salt stress treatment.

Salt stress is one of the most important factors limiting plant cultivation. Many investigations of plant response to high salinity have been performed using conventional transcriptomics and/or proteomics approaches. However, transcriptomics and proteomics techniques are not all-encompassing methods that can achieve exclusive insights into the metabolite networks contributing to biochemical reactions. Hence, the functions of the complex stress response pathways are yet to be determined, especially at the metabolic level. A time-course metabolic profiling with Arabidopsis thaliana cell cultures after the imposition of salt stress is reported in this study. Analyses of primary metabolites, especially small polar metabolites such as amino acids, sugars, sugar alcohols, organic acids, and amines, was performed by GC/MS and LC/MS at 0.5, 1, 2, 4, 12, 24, 48, and 72 h after a salt-stress treatment with 100 mM NaCl being the final concentration. The mass chromatographic data were converted into matrix data sets, which were subjected to data mining processes, including principal component analysis (PCA) and batch-learning self-organizing mapping analysis (BL-SOM). The mining results suggest that the methylation cycle for the supply of methyl groups, the phenylpropanoid pathway for lignin production, and glycinebetaine biosynthesis are synergetically induced as a short-term response against salt-stress treatment. The results also suggest the the co-induction of glycolysis and sucrose metabolism as well as co-reduction of the methylation cycle as long-term responses to salt stress.

Amines↗

The Comprehensive Microbial Resource.

One challenge presented by large-scale genome sequencing efforts is effective display of uniform information to the scientific community. The Comprehensive Microbial Resource (CMR) contains robust annotation of all complete microbial genomes and allows for a wide variety of data retrievals. The bacterial information has been placed on the Web at http://www.tigr.org/CMR for retrieval using standard web browsing technology. Retrievals can be based on protein properties such as molecular weight or hydrophobicity, GC-content, functional role assignments and taxonomy. The CMR also has special web-based tools to allow data mining using pre-run homology searches, whole genome dot-plots, batch downloading and traversal across genomes using a variety of datatypes.

Bacteria↗

TissueInfo: high-throughput identification of tissue expression profiles and specificity.

We describe TissueInfo, a knowledge-based method for the high-throughput identification of tissue expression profiles and tissue specificity. TissueInfo defines a set of tissue information calculations that can be computed for large numbers of genes, expressed sequence tags (ESTs) or proteins. Tissue information records that result from the TissueInfo calculations are used to generate tables suitable for data mining and for the selection of genes according to a given expression profile or specificity. When benchmarked against a test set of 116 proteins and literature information, TissueInfo was found to be accurate for 69% of identified tissue specificities and for 80% of expression profiles. The accuracy of the identifications can be increased if query sequences for which little information is available from dbEST are ignored. Thus, with 80% coverage, TissueInfo achieves an accuracy of 76% for specificity and 89% for expression. For the same set of proteins, the curated tissue specificity offered in SWISS-PROT was accurate in 78% of cases. TissueInfo can be useful for the selection of clones for custom microarrays, selection of training sets for ab initio identification of tissue information, gene discovery and genome-wide predictions. Further information about the program can be found at http://icb.mssm.edu/tissueinfo.

Computational Biology↗

The Protein Information Resource: an integrated public resource of functional annotation of proteins.

The Protein Information Resource (PIR) serves as an integrated public resource of functional annotation of protein data to support genomic/proteomic research and scientific discovery. The PIR, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International Protein Information Database (JIPID), produces the PIR-International Protein Sequence Database (PSD), the major annotated protein sequence database in the public domain, containing about 250 000 proteins. To improve protein annotation and the coverage of experimentally validated data, a bibliography submission system is developed for scientists to submit, categorize and retrieve literature information. Comprehensive protein information is available from iProClass, which includes family classification at the superfamily, domain and motif levels, structural and functional features of proteins, as well as cross-references to over 40 biological databases. To provide timely and comprehensive protein data with source attribution, we have introduced a non-redundant reference protein database, PIR-NREF. The database consists of about 800 000 proteins collected from PIR-PSD, SWISS-PROT, TrEMBL, GenPept, RefSeq and PDB, with composite protein names and literature data. To promote database interoperability, we provide XML data distribution and open database schema, and adopt common ontologies. The PIR web site (http://pir.georgetown.edu/) features data mining and sequence analysis tools for information retrieval and functional identification of proteins based on both sequence and annotation information. The PIR databases and other files are also available by FTP (ftp://nbrfa.georgetown.edu/pir_databases).

Amino Acid Sequence↗

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗

GEPAS: A web-based resource for microarray gene expression data analysis.

We present a web-based pipeline for microarray gene expression profile analysis, GEPAS, which stands for Gene Expression Profile Analysis Suite (http://gepas.bioinfo.cnio.es). GEPAS is composed of different interconnected modules which include tools for data pre-processing, two-conditions comparison, unsupervised and supervised clustering (which include some of the most popular methods as well as home made algorithms) and several tests for differential gene expression among different classes, continuous variables or survival analysis. A multiple purpose tool for data mining, based on Gene Ontology, is also linked to the tools, which constitutes a very convenient way of analysing clustering results. On-line tutorials are available from our main web server (http://bioinfo.cnio.es).

Cluster Analysis↗

Onto-Tools: an ensemble of web-accessible, ontology-based tools for the functional design and interpretation of high-throughput gene expression experiments.

The Onto-Tools suite is composed of an annotation database and five seamlessly integrated web-accessible data mining tools: Onto-Express (OE), Onto-Compare (OC), Onto-Design (OD), Onto-Translate (OT) and Onto-Miner (OM). OM is a new tool that provides a unified access point and an application programming interface for most annotations available. Our database has been enhanced with more than 120 new commercial microarrays and annotations for Rattus norvegicus, Drosophila melanogaster and Carnorhabditis elegans. The Onto-Tools have been redesigned to provide better biological insight, improved performance and user convenience. The new features implemented in OE include support for gene names, LocusLink IDs and Gene Ontology (GO) IDs, ability to specify fold changes for the input genes, links to the KEGG pathway database and detailed output files. OC allows comparisons of the functional bias of more than 170 commercial microarrays. The latest version of OD allows the user to specify keywords if the exact GO term is not known as well as providing more details than the previous version. OE, OC and OD now have an integrated GO browser that allows the user to customize the level of abstraction for each GO category. The Onto-Tools are available online at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

PFD: a database for the investigation of protein folding kinetics and stability.

We have developed a new database that collects all protein folding data into a single, easily accessible public resource. The Protein Folding Database (PFD) contains annotated structural, methodological, kinetic and thermodynamic data for more than 50 proteins, from 39 families. A user-friendly web interface has been developed that allows powerful searching, browsing and information retrieval, whilst providing links to other protein databases. The database structure allows visualization of folding data in a useful and novel way, with a long-term aim of facilitating data mining and bioinformatics approaches. PFD can be accessed freely at http://pfd.med.monash.edu.au.

Databases, Protein↗