Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

EST mining and functional expression assays identify extracellular effector proteins from the plant pathogen Phytophthora.

Plant pathogenic microbes have the remarkable ability to manipulate biochemical, physiological, and morphological processes in their host plants. These manipulations are achieved through a diverse array of effector molecules that can either promote infection or trigger defense responses. We describe a general functional genomics approach aimed at identifying extracellular effector proteins from plant pathogenic microorganisms by combining data mining of expressed sequence tags (ESTs) with virus-based high-throughput functional expression assays in plants. PexFinder, an algorithm for automated identification of extracellular proteins from EST data sets, was developed and applied to 2147 ESTs from the oomycete plant pathogen Phytophthora infestans. The program identified 261 ESTs (12.2%) corresponding to a set of 142 nonredundant Pex (Phytophthora extracellular protein) cDNAs. Of these, 78 (55%) Pex cDNAs were novel with no significant matches in public databases. Validation of PexFinder was performed using proteomic analysis of secreted protein of P. infestans. To identify which of the Pex cDNAs encode effector proteins that manipulate plant processes, high-throughput functional expression assays in plants were performed on 63 of the identified cDNAs using an Agrobacterium tumefaciens binary vector carrying the potato virus X (PVX) genome. This led to the discovery of two novel necrosis-inducing cDNAs, crn1 and crn2, encoding extracellular proteins that belong to a large and complex protein family in Phytophthora. Further characterization of the crn genes indicated that they are both expressed in P. infestans during colonization of the host plant tomato and that crn2 induced defense-response genes in tomato. Our results indicate that combining data mining using PexFinder with PVX-based functional assays can facilitate the discovery of novel pathogen effector proteins. In principle, this strategy can be applied to a variety of eukaryotic plant pathogens, including oomycetes, fungi, and nematodes.

Algal Proteins↗

Metabolomics--the link between genotypes and phenotypes.

Metabolites are the end products of cellular regulatory processes, and their levels can be regarded as the ultimate response of biological systems to genetic or environmental changes. In parallel to the terms 'transcriptome' and proteome', the set of metabolites synthesized by a biological system constitute its 'metabolome'. Yet, unlike other functional genomics approaches, the unbiased simultaneous identification and quantification of plant metabolomes has been largely neglected. Until recently, most analyses were restricted to profiling selected classes of compounds, or to fingerprinting metabolic changes without sufficient analytical resolution to determine metabolite levels and identities individually. As a prerequisite for metabolomic analysis, careful consideration of the methods employed for tissue extraction, sample preparation, data acquisition, and data mining must be taken. In this review, the differences among metabolite target analysis, metabolite profiling, and metabolic fingerprinting are clarified, and terms are defined. Current approaches are examined, and potential applications are summarized with a special emphasis on data mining and mathematical modelling of metabolism.

Arabidopsis↗

Mining knowledge for HEp-2 cell image classification.

HEp-2 cells are used for the identification of antinuclear autoantibodies (ANAs). They allow for recognition of over 30 different nuclear and cytoplasmic patterns, which are given by upwards of 100 different autoantibodies. The identification of the patterns has recently been done manually by a human inspecting the slides with a microscope. In this paper, we present results on the analysis and classification of cells using image analysis and data mining techniques. Starting from a knowledge acquisition process with a human operator, we developed an image analysis and feature extraction algorithm. The collection of the dataset was done based on an expert's image reading and based on the automatic extracted features. A dataset containing 132 features for each entry was set up and given to a data mining algorithm to find out the relevant features among this large feature set and to construct the classification knowledge. The classifier was evaluated by cross validation. The results gave the expert new insights into the necessary features and the classification knowledge and show the feasibility of an automated inspection system.

Algorithms↗

Systems biology for cancer.

PURPOSE OF REVIEW: Significant insight can be gained into complex biologic mechanisms of cancer via a combined computational and experimental systems biology approach. This review highlights some of the major systems biology efforts that were applied to cancer in the past year. RECENT FINDINGS: Two main approaches to computational systems biology are discussed: mechanistic dynamical simulations and inferential data mining. Significant developments have occurred in both areas. For example, mechanistic simulations of the EGFR pathway are promoting understanding of cancer, and Bayesian inference approaches allow for the reconstruction of regulatory networks. In addition, the article reports on advancements in experimental systems biology for determining protein-protein interactions and quantifying protein expression to generate the necessary data for computational modeling and inferential data mining. Emerging approaches will further improve the ability to bridge the gap between in vitro systems and in vivo human biology. Technologies paving the way include in vitro models that better reflect in vivo tumors, microfabricated devices of human physiology, and improved animal models. SUMMARY: An important challenge facing the field is how better to translate in vitro discoveries to the clinic. Computational systems biology approaches that use omic data to predict biology along with novel experimental systems that better represent human in vivo biology will prove useful in bridging this gap. Although still early, the potential application of systems biology and the future evolution of the field will significantly affect understanding of cancer disease mechanisms and the ability to devise effective therapeutics.

Computational Biology↗

CHMIS-C: a comprehensive herbal medicine information system for cancer.

A comprehensive herbal medicine information system for cancer (CHMIS-C) has been developed. The current version of the database integrates information on more than 200 anticancer herbal recipes that have been used for the treatment of different types of cancer in clinic, 900 individual ingredients, and 8500 small organic molecules isolated from herbal medicines. Furthermore, subsidiary databases of literature references and molecular targets have been constructed. A number of web-based searching tools have been developed and integrated into the information system for efficient data mining. The compounds in the database have been linked to the corresponding entries in the National Cancer Institute's database, and to a database of drugs approved by the U.S. Food and Drug Administration. This paper provides a description of the individual subsidiary databases, integration of the entire database, and data mining tools. We demonstrate that this comprehensive information system may be used as an effective informatics tool for anticancer drug discovery.

Antineoplastic Agents, Phytogenic↗

Pharmacovigilance in the 21st century: new systematic tools for an old problem.

The large number of adverse-event reports generated by marketed drugs and devices argues for the application of validated computerized algorithms to supplement traditional methods of detecting adverse-event signals. Difficulties in accurately estimating patient exposure and background rates for a given event in a specific population hinder risk estimation in spontaneous adverse-event databases. The United States Food and Drug Administration (FDA) is evaluating a Bayesian data mining system called Multi-item Gamma Poisson Shrinker (MGPS) to enhance the FDA's ability to monitor the safety of drugs, biologics, and vaccines after they have been approved for use. The MGPS computes adjusted higher-than-expected reporting relationships between drugs and adverse events across 35 years of data relative to internal background rates. The MGPS can also adjust for random noise by using a model derived from the data, and corrects for temporal trends and confounding related to age, sex, and other variables by stratifying over 900 categories. Signals can then be compared with or used in conjunction with other sources (e.g. clinical trials, general practice databases) to further study the adverse-event risk. The example of pancreatitis risk with atypical antipsychotics, valproic acid, and valproate is used to discuss the strengths and limitations of MGPS versus traditional methods. Validated data mining techniques offer great promise to enhance pharmacovigilance practices.

Adverse Drug Reaction Reporting Systems↗

Search for a shared segment on chromosome 10q26 in patients with bipolar affective disorder or schizophrenia from the Faroe Islands.

Previous linkage studies have suggested a new locus for bipolar affective disorder and possibly also for schizophrenia on chromosome 10q26. We searched for allelic association and chromosome segment and haplotype sharing on chromosome 10q26 among distantly related patients with bipolar affective disorder or schizophrenia and controls from the relatively isolated population of the Faroe Islands by investigating 22 microsatellite markers from a 35 cM region. We used a combined approach with both assumption free tests and tests based on genealogical relationships. The 6.5 cM region between D10S1230 and D10S2322, which has been implied in previous linkage analyses, received some support. A search for segment sharing yielded empirical P-values around 0.02 among patients with bipolar affective disorder and around 0.03 for patients with schizophrenia. For both disorders combined allelic association yielded empirical P-values around 0.003 at marker D10S1723. A haplotype data mining approach supported haplotype sharing in this region. In another, more distal, 11.5 cM region between markers D10S214 and D10S505, which has received support in previous linkage studies, increased haplotype sharing in patients with bipolar affective disorder was supported by Fisher's exact test, tests based on genealogy and by haplotype data mining. Our findings yield some support for a risk gene for bipolar affective disorder and possibly also for schizophrenia.

Alleles↗

SIFamide is a highly conserved neuropeptide: a comparative study in different insect species.

Neb-LFamide or AYRKPPFNGSLFamide was originally purified from the grey flesh fly Neobellieria bullata as a myotropic neuropeptide. We studied the occurrence of this peptide and its isoforms in the central nervous system of different insect species by means of whole mount fluorescence immunohistochemistry, mass spectrometry, and data mining. We found that both sequence and immunoreactive distribution pattern are very conserved in the studied insects. In all species and stages we counted two pairs of immunoreactive cells in the pars intercerebralis. These cells projected axons throughout the ventral nerve cord. In the adult CNSs they formed a large number of immunoreactive varicosities as well. Mass spectrometry and data mining revealed that SIFamide exists in two isoforms: [G1]-SIFamide and [A1]-SIFamide. In addition, the SIFamide joining peptide is relatively well conserved throughout arthropod species. The conserved presence of two cysteine residues, separated by six amino acid residues, allows the formation of disulphide bridges.

Amino Acid Sequence↗

Blast2GO goes grid: developing a grid-enabled prototype for functional genomics analysis.

The vast amount in complexity of data generated in Genomic Research implies that new dedicated and powerful computational tools need to be developed to meet their analysis requirements. Blast2GO (B2G) is a bioinformatics tool for Gene Ontology-based DNA or protein sequence annotation and function-based data mining. The application has been developed with the aim of affering an easy-to-use tool for functional genomics research. Typical B2G users are middle size genomics labs carrying out sequencing, ETS and microarray projects, handling datasets up to several thousand sequences. In the current version of B2G. The power and analytical potential of both annotation and function data-mining is somehow restricted to the computational power behind each particular installation. In order to be able to offer the possibility of an enhanced computational capacity within this bioinformatics application, a Grid component is being developed. A prototype has been conceived for the particular problem of speeding up the Blast searches to obtain fast results for large datasets. Many efforts have been done in the literature concerning the speeding up of Blast searches, but few of them deal with the use of large heterogeneous production Grid Infrastructures. These are the infrastructures that could reach the largest number of resources and the best load balancing for data access. The Grid Service under development will analyse requests based on the number of sequences, splitting them accordingly to the available resources. Lower-level computation will be performed through MPIBLAST. The software architecture is based on the WSRF standard.

Computational Biology↗

Searching QTL by gene expression: analysis of diabesity.

BACKGROUND: Recent developments in sequence databases provide the opportunity to relate the expression pattern of genes to their genomic position, thus creating a transcriptome map. Quantitative trait loci (QTL) are phenotypically-defined chromosomal regions that contribute to allelically variant biological traits, and by overlaying QTL on the transcriptome, the search for candidate genes becomes extremely focused. RESULTS: We used our novel data mining tool, ExQuest, to select genes within known diabesity QTL showing enriched expression in primary diabesity affected tissues. We then quantified transcripts in adipose, pancreas, and liver tissue from Tally Ho mice, a multigenic model for Type II diabetes (T2D), and from diabesity-resistant C57BL/6J controls. Analysis of the resulting quantitative PCR data using the Global Pattern Recognition analytical algorithm identified a number of genes whose expression is altered, and thus are novel candidates for diabesity QTL and/or pathways associated with diabesity. CONCLUSION: Transcription-based data mining of genes in QTL-limited intervals followed by efficient quantitative PCR methods is an effective strategy for identifying genes that may contribute to complex pathophysiological processes.

Algorithms↗

A data review and re-assessment of ovarian cancer serum proteomic profiling.

BACKGROUND: The early detection of ovarian cancer has the potential to dramatically reduce mortality. Recently, the use of mass spectrometry to develop profiles of patient serum proteins, combined with advanced data mining algorithms has been reported as a promising method to achieve this goal. In this report, we analyze the Ovarian Dataset 8-7-02 downloaded from the Clinical Proteomics Program Databank website, using nonparametric statistics and stepwise discriminant analysis to develop rules to diagnose patients, as well as to understand general patterns in the data that may guide future research. RESULTS: The mass spectrometry serum profiles derived from cancer and controls exhibited numerous statistical differences. For example, use of the Wilcoxon test in comparing the intensity at each of the 15,154 mass to charge (M/Z) values between the cancer and controls, resulted in the detection of 3,591 M/Z values whose intensities differed by a p-value of 10-6 or less. The region containing the M/Z values of greatest statistical difference between cancer and controls occurred at M/Z values less than 500. For example the M/Z values of 2.7921478 and 245.53704 could be used to significantly separate the cancer from control groups. Three other sets of M/Z values were developed using a training set that could distinguish between cancer and control subjects in a test set with 100% sensitivity and specificity. CONCLUSION: The ability to discriminate between cancer and control subjects based on the M/Z values of 2.7921478 and 245.53704 reveals the existence of a significant non-biologic experimental bias between these two groups. This bias may invalidate attempts to use this dataset to find patterns of reproducible diagnostic value. To minimize false discovery, results using mass spectrometry and data mining algorithms should be carefully reviewed and benchmarked with routine statistical methods.

Artificial Intelligence↗

An aphasia database on the internet: a model for computer-assisted analysis in aphasiology.

A web-based software model was developed as an example for data mining in aphasiology. It is used for educating medical and engineering students. It is based upon a database of 254 aphasic patients which contains the diagnosis of the aphasia type, profiles of an aphasia test battery (Aachen Aphasia Test), and some further clinical information. In addition, the cerebral lesion profiles of 147 of these cases were standardized by transferring the coordinates of the lesions to a 3D reference brain based upon the ACPC coordinate system. Two artificial neural networks were used to perform a classification of the aphasia type. First, a coarse classification was achieved by using an assessment of spontaneous speech of the patient which produced correct results in 87% of the test cases. Data analysis tools were used to select four features of the 30 available test features to yield a more accurate diagnosis. This classifier produced correct results in 92% of the test cases. The neural network approach is similar to grouping performed in group studies, while the nearest-neighbor method shows a design more similar to case studies. It finds the neurolinguistic and the lesion data of patients whose AAT profiles are most similar to the user's input. This way lesion profiles can be compared to each other interindividually. The Aphasia Diagnoser is available on the Web address http://fuzzy.iau.dtu.dk/aphasia.nsf and thus should facilitate a discussion about the reliability and possibilities of data-mining techniques in aphasiology.

Aphasia↗

Discovery of association rules in medical data.

Data mining is a technique for discovering useful information from large databases. This technique is currently being profitably used by a number of industries. A common approach for information discovery is to identify association rules which reveal relationships among different items. In this paper, we use this approach to analyse a large database containing medical-record data. Our aim is to obtain association rules indicating relationships between procedures performed on a patient and the reported diagnoses. Random sampling was used to obtain these association rules. After reviewing the basic concepts associated with data mining, we discuss our approach for identifying association rules and report on the rules generated.

Algorithms↗

Internet-enabled high-resolution brain mapping and virtual microscopy.

Virtual microscopy involves the conversion of histological sections mounted on glass microscope slides to high-resolution digital images. Virtual microscopy offers several advantages over traditional microscopy, including remote viewing and data sharing, annotation, and various forms of data mining. We describe a method utilizing virtual microscopy for generation of internet-enabled, high-resolution brain maps and atlases. Virtual microscopy-based digital brain atlases have resolutions approaching 100,000 dpi, which exceeds by three or more orders of magnitude resolutions obtainable in conventional print atlases, MRI, and flat-bed scanning. Virtual microscopy-based digital brain atlases are superior to conventional print atlases in five respects: (1) resolution, (2) annotation, (3) interaction, (4) data integration, and (5) data mining. Implementation of virtual microscopy-based digital brain atlases is located at BrainMaps.org, which is based on more than 10 million megapixels (35 terabytes) of scanned images of serial sections of primate and non-primate brains with a resolution of 0.46 microm/pixel (55,000 dpi). The method can be replicated by labs seeking to increase accessibility and sharing of neuroanatomical data. Online tools offer the possibility of visualizing and exploring completely digitized sections of brains at a sub-neuronal level and can facilitate large-scale connectional tracing, histochemical, and stereological analyses.

Animals↗

Modular transcriptional activity characterizes the initiation and progression of autoimmune encephalomyelitis.

Murine experimental autoimmune encephalomyelitis is a well-established model that recapitulates many clinical and physiopathological aspects of multiple sclerosis (MS). An important conceptual development in the understanding of both experimental autoimmune encephalomyelitis and MS pathogenesis has been the compartmentalization of the mechanistic process into two distinct but overlapping and connected phases, inflammatory and neurodegenerative. However, the dynamics of CNS transcriptional changes that underlie the development and regression of the phenotype are not well understood. Our report presents the first high frequency longitudinal study looking at the earliest transcriptional changes in the CNS of NOD mice immunized with myelin oligodendrocyte glycoprotein 35-55 in CFA. Microarray-based gene expression profiling and histopathological analysis were performed from spinal cord samples obtained at 13 time points around the first clinical symptom (every other day until day 11 and every day onward until day 19 postimmunization). Advanced statistics and data-mining algorithms were used to identify expression signatures that correlated with disease stage and histological profiles. Discrete phases of neuroinflammation were accompanied by distinctive expression signatures, in which altered immune to neural gene expression ratios were observed. By using high frequency gene expression analysis we captured expression profiles that were characteristic of the transition from innate to adaptive immune response in this experimental paradigm between days 11 and 12 postimmunization. Our study demonstrates the utility of large-scale transcriptional studies and advanced data mining to decipher complex biological processes such as those involved in MS and other neurodegenerative disorders.

Adjuvants, Immunologic↗

A technique for identifying three diagnostic findings using association analysis.

In diagnosing diseases in clinical practice, a combination of three clinical findings is often used to represent each disease. This is largely because it is often difficult or impractical to assess for all possible combinations of symptoms and abnormal exam findings that occur in any particular disease. For most diseases, diagnostic triads are based on empirical observations. In this study, we determined diagnostic triads for chronic diseases using data mining procedures. We also verified the combinations' validity as well as our procedure for determining them. We used symptoms and examination findings from 477 patients with chronic diseases, collected as part of a 35-year longitudinal study begun in 1968. For each patient there were 295 items from examinations in internal medicine, dermatology, ophthalmology, dentistry and blood tests. We judged each item to be either normal or abnormal, and restricted the analysis to the abnormal findings. To analyze such an exhaustive assortment, we used the data mining technique of association analysis. The analysis generated three clinical findings for each disease. Diseases were defined based on blood tests. Searching through all 295 items to find the three most useful clinical findings would be impractical on a commodity PC. However, by excluding normal items, we were able to sufficiently reduce the total number of combinations so as to make combinatorial analysis on a PC feasible. In addition to more accurate diagnoses, we believe our technique can identify those diagnostic data that are more cost effective in terms of time and other resources required for their collection.

Data Collection↗

Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program.

The UMLS Metathesaurus, the largest thesaurus in the biomedical domain, provides a representation of biomedical knowledge consisting of concepts classified by semantic type and both hierarchical and non-hierarchical relationships among the concepts. This knowledge has proved useful for many applications including decision support systems, management of patient records, information retrieval (IR) and data mining. Gaining effective access to the knowledge is critical to the success of these applications. This paper describes MetaMap, a program developed at the National Library of Medicine (NLM) to map biomedical text to the Metathesaurus or, equivalently, to discover Metathesaurus concepts referred to in text. MetaMap uses a knowledge intensive approach based on symbolic, natural language processing (NLP) and computational linguistic techniques. Besides being applied for both IR and data mining applications, MetaMap is one of the foundations of NLM's Indexing Initiative System which is being applied to both semi-automatic and fully automatic indexing of the biomedical literature at the library.

Abstracting and Indexing↗

Metabolic profiling allows comprehensive phenotyping of genetically or environmentally modified plant systems.

Metabolic profiling using gas chromatography-mass spectrometry technologies is a technique whose potential in the field of functional genomics is largely untapped. To demonstrate the general usefulness of this technique, we applied to diverse plant genotypes a recently developed profiling protocol that allows detection of a wide range of hydrophilic metabolites within a single chromatographic run. For this purpose, we chose four independent potato genotypes characterized by modifications in sucrose metabolism. Using data-mining tools, including hierarchical cluster analysis and principle component analysis, we were able to assign clusters to the individual plant systems and to determine relative distances between these clusters. Extraction analysis allowed identification of the most important components of these clusters. Furthermore, correlation analysis revealed close linkages between a broad spectrum of metabolites. In a second, complementary approach, we subjected wild-type potato tissue to environmental manipulations. The metabolic profiles from these experiments were compared with the data sets obtained for the transgenic systems, thus illustrating the potential of metabolic profiling in assessing how a genetic modification can be phenocopied by environmental conditions. In summary, these data demonstrate the use of metabolic profiling in conjunction with data-mining tools as a technique for the comprehensive characterization of a plant genotype.

Cluster Analysis↗