Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Conceptual biology, hypothesis discovery, and text mining: Swanson's legacy.

Innovative biomedical librarians and information specialists who want to expand their roles as expert searchers need to know about profound changes in biology and parallel trends in text mining. In recent years, conceptual biology has emerged as a complement to empirical biology. This is partly in response to the availability of massive digital resources such as the network of databases for molecular biologists at the National Center for Biotechnology Information. Developments in text mining and hypothesis discovery systems based on the early work of Swanson, a mathematician and information scientist, are coincident with the emergence of conceptual biology. Very little has been written to introduce biomedical digital librarians to these new trends. In this paper, background for data and text mining, as well as for knowledge discovery in databases (KDD) and in text (KDT) is presented, then a brief review of Swanson's ideas, followed by a discussion of recent approaches to hypothesis discovery and testing. 'Testing' in the context of text mining involves partially automated methods for finding evidence in the literature to support hypothetical relationships. Concluding remarks follow regarding (a) the limits of current strategies for evaluation of hypothesis discovery systems and (b) the role of literature-based discovery in concert with empirical research. Report of an informatics-driven literature review for biomarkers of systemic lupus erythematosus is mentioned. Swanson's vision of the hidden value in the literature of science and, by extension, in biomedical digital databases, is still remarkably generative for information scientists, biologists, and physicians.

Journal Article↗

Integration of high-resolution array comparative genomic hybridization analysis of chromosome 16q with expression array data refines common regions of loss at 16q23-qter and identifies underlying candidate tumor suppressor genes in prostate cancer.

We have constructed a high-resolution genomic microarray of human chromosome 16q, and used it for comparative genomic hybridization analysis of 16 prostate tumors. We demarcated 10 regions of genomic loss between 16q23.1 and 16qter that occurred in five or more samples. Mining expression array data from four independent studies allowed us to identify 11 genes that were frequently underexpressed in prostate cancer and that co-localized with a region of genomic loss. Quantitative expression analyses of these genes in matched tumor and benign tissue from 13 patients showed that six of these 11 (WWOX, WFDC1, MAF, FOXF1, MVD and the predicted novel transcript Q9H0B8 (NM_031476)) had significant and consistent downregulation in the tumors relative to normal prostate tissue expression making them candidate tumor suppressor genes.

Chromosomes, Human, Pair 16↗

VisANT: an online visualization and analysis tool for biological interaction data.

BACKGROUND: New techniques for determining relationships between biomolecules of all types--genes, proteins, noncoding DNA, metabolites and small molecules--are now making a substantial contribution to the widely discussed explosion of facts about the cell. The data generated by these techniques promote a picture of the cell as an interconnected information network, with molecular components linked with one another in topologies that can encode and represent many features of cellular function. This networked view of biology brings the potential for systematic understanding of living molecular systems. RESULTS: We present VisANT, an application for integrating biomolecular interaction data into a cohesive, graphical interface. This software features a multi-tiered architecture for data flexibility, separating back-end modules for data retrieval from a front-end visualization and analysis package. VisANT is a freely available, open-source tool for researchers, and offers an online interface for a large range of published data sets on biomolecular interactions, including those entered by users. This system is integrated with standard databases for organized annotation, including GenBank, KEGG and SwissProt. VisANT is a Java-based, platform-independent tool suitable for a wide range of biological applications, including studies of pathways, gene regulation and systems biology. CONCLUSION: VisANT has been developed to provide interactive visual mining of biological interaction data sets. The new software provides a general tool for mining and visualizing such data in the context of sequence, pathway, structure, and associated annotations. Interaction and predicted association data can be combined, overlaid, manipulated and analyzed using a variety of built-in functions. VisANT is available at http://visant.bu.edu.

Animals↗

Virus infection and the interferon response: a global view through functional genomics.

The primary focus of this chapter is on providing an overview of how we use the tools of functional genomics to study virus infection, the interferon response, and the mechanisms by which viruses attenuate or evade this response to ensure successful replication. We provide examples of the types of analyses we perform, experimental design considerations, and the tools and techniques we use for data processing and mining. We have not attempted to provide detailed protocols for performing microarray or proteomics experiments because even individual components of such analyses, for example, techniques for the isolation and amplification of ribonucleic acid, could easily be the subject of an entire chapter. Rather, our goal is to show how data obtained from global gene expression and protein profiling can be used to gain new insights into virus-host interactions, with particular emphasis on the interferon response and its modulation by virus infection.

Animals↗

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals↗

MiCoViTo: a tool for gene-centric comparison and visualization of yeast transcriptome states.

BACKGROUND: Information obtained by DNA microarray technology gives a rough snapshot of the transcriptome state, i.e., the expression level of all the genes expressed in a cell population at any given time. One of the challenging questions raised by the tremendous amount of microarray data is to identify groups of co-regulated genes and to understand their role in cell functions. RESULTS: MiCoViTo (Microarray Comparison Visualization Tool) is a set of biologists' tools for exploring, comparing and visualizing changes in the yeast transcriptome by a gene-centric approach. A relational database includes data linked to genome expression and graphical output makes it easy to visualize clusters of co-expressed genes in the context of available biological information. To this aim, upload of personal data is possible and microarray data from fifty publications dedicated to S. cerevisiae are provided on-line. A web interface guides the biologist during the usage of this tool and is freely accessible at http://www.transcriptome.ens.fr/micovito/. CONCLUSIONS: MiCoViTo offers an easy-to-read picture of local transcriptional changes connected to current biological knowledge. This should help biologists to mine yeast microarray data and better understand the underlying biology. We plan to add functional annotations from other organisms. That would allow inter-species comparison of transcriptomes via orthology tables.

Cluster Analysis↗

New tools for cancer chemotherapy: computational assistance for tailoring treatments.

Computational models of cancer chemotherapy have the potential to streamline clinical trial design, contribute to the design of rational, tailored treatments, and facilitate our understanding of experimental results. Mechanistic models based on functional data from tumor biopsies will enable physicians to predict response to treatment for a specific patient, in contrast to statistical models in which the probability of response for a given patient may differ substantially from the population average. While microarray analyses of gene expression also show promise for guiding individualized treatments, it may be difficult to link statistical mining of microarray data with mechanistic, tailored treatments. Furthermore, gene expression does not identify how drugs should be scheduled. This review summarizes mechanistic mathematical models developed to improve the design of chemotherapy regimens. Mechanistic models that incorporate both genetic resistance and cell cycle-mediated resistance during treatment with multiple drugs will be most useful in designing treatment regimens tailored for individuals. Because there are already a number of papers that address the applications of microarray technology, we will limit our discussion to the contrasts between mechanistic computational models and microarray technology, and how these two approaches may complement one another.

Drug Resistance, Neoplasm↗

The impact of computing technology on pharmaceutical and biotech research.

Pharmaceutical and biotechnology companies continue to make investments in research techniques including those using high-performance computing technology. Much of this research is fueled by the need to analyze the ever-growing store of genomic and proteomic data. However, many research and development investments, including techniques to mine newly produced genomic data, have not provided the anticipated improvements in productivity. This raises the questions of whether advancements expected from next-generation computing technology can translate into tangible benefits for pharmaceutical and biotechnology researchers, and whether researchers can capitalize on an increased understanding of biological systems using tools made more readily available in high-performance computing. The author provides some background on the source of the advancements in the computing industry and offers examples from other scientific applications that point to potential benefits in life science research.

Biotechnology↗

[The efficacy of speleotherapy in salt mines in children with bronchial asthma based on the data from immediate and late observations].

Speleotherapy was conducted in 216 children with bronchial asthma treated in conditions of salt mines situated near the town of Nakhichevan. The assessment of clinical, immunological and functional parameters showed that the best results had been achieved in atopic asthma running a light or moderate course. Speleotherapy courses noticeably diminished broncho-obstructive syndrome, improved pulmonary ventilation. The improvement proved stable in the majority of the patients. It is recommended to include speleotherapy in salt mines into combined rehabilitation treatment of pediatric asthmatics.

Adolescent↗

Mining literature for systems biology.

Currently, literature is integrated in systems biology studies in three ways. Hand-curated pathways have been sufficient for assembling models in numerous studies. Second, literature is frequently accessed in a derived form, such as the concepts represented by the Medical Subject Headings (MeSH) and Gene Ontologies (GO), or functional relationships captured in protein-protein interaction (PPI) databases; both of these are convenient, consistent reductions of more complex concepts expressed as free text in the literature. Moreover, their contents are easily integrated into computational processes required for dealing with large data sets. Last, mining text directly for specific types of information is on the rise as text analytics methods become more accurate and accessible. These uses of literature, specifically manual curation, derived concepts captured in ontologies and databases, and indirect and direct application of text mining, will be discussed as they pertain to systems biology.

Animals↗

Mortality of lead smelter workers.

To examine patterns of death in lead smelter workers, a retrospective analysis of mortality was conducted in a cohort of 1,987 males employed between 1940 and 1965 at a primary lead smelter in Idaho. Overall mortality was similar to that of the United States white male population (standardized mortality ratio (SMR) = 98). Excess mortality, however, was found from chronic renal disease (SMR = 192; confidence interval (CI) = 88-364), and the risk of death from renal disease increased with increasing duration of employment, such that after 20 years employment, the standardized mortality ratio reached 392 (CI = 107-1,004). Excess mortality was also noted for nonmalignant respiratory disease (SMR = 187, CI = 128-264). Eight of 32 deaths in this category were caused by silicosis; at least five workers who died of silicosis had been miners for a part of their lives. An additional 11 deaths resulted from tuberculosis (SMR = 139; CI = 69-249); in six of these cases, silicosis was a contributory cause of death. Cancer mortality was not increased overall (SMR = 95; CI = 78-114). An increase, however, was noted for deaths from kidney cancer (six cases; SMR = 204; CI = 75-444). Finally, excess mortality was noted for injuries (SMR = 138; CI = 104-179); 13 (23%) of the 56 deaths in this category were caused by mining injuries. The data from this study are consistent with previous reports of increased mortality from chronic renal disease in persons exposed occupationally to lead.

Actuarial Analysis↗

Identifying potential receptors and routes of contaminant exposure in the traditional territory of the Ouje-Bougoumou Cree: land use and a geographical information system.

Great concern has been raised with respect to the 13 traplines that constitute the traditional territory of the Ouje-Bougoumou Cree located in the James Bay region of northern Quebec, Canada, with respect to mine wastes originating from three local mines. As a result, an "Integrative Risk Assessment" was initiated consisting of three interrelated components: a comprehensive human health study, an assessment of the existing ecological/environmental database, and a land use/potential sites of concern study. In this paper, we document past and present land use in the traditional territory of the Ouje-Bougoumou Cree for 72 heads of households, including 13 tallymen, and use a Geographic Information System (GIS) to layer harvest/hunting and gathering/collecting data over known mining areas and potential sites of concern. In this way, potential receptors of contamination and routes of human exposure were identified. Areas of overlap with respect to land use activity and mining operations were relatively extensive for certain harvesting activities (e.g., beaver, Castor canadensis and various species of game birds), less so for fish harvesting (all species) and water collection, and relatively restrictive for large mammal harvesting and collection of fire wood (and other collection activities). Potential receptors of contaminants associated with mining activity (e.g., fish and small mammals) and potential routes of exposure (e.g., ingestion of contaminated game and drinking of contaminated water) were identified.

Conservation of Natural Resources↗

Bioaccessible arsenic in the home environment in southwest England.

Samples of household dust and garden soil were collected from twenty households in the vicinity of an ex-mining site in southwest England and from nine households in a control village. All samples were analysed by ICP-MS for pseudo-total arsenic (As) concentrations and the results show clearly elevated levels, with maximum As concentrations of 486 microg g(-1) in housedusts and 471 microg g(-1) in garden soils (and mean concentrations of 149 microg g(-1) and 262 microg g(-1), respectively). Arsenic concentrations in all samples from the mining area exceeded the UK Soil Guideline Value (SGV) of 20 microg g(-1). No significant correlation was observed between garden soil and housedust As concentrations. Bioaccessible As concentrations were determined in a small subset of samples using the Physiologically Based Extraction Test (PBET). For the stomach phase of the PBET, bioaccessibility percentages of 10-20% were generally recorded. Higher percentages (generally 30-45%) were recorded in the intestine phases with a maximum value (for one of the housedusts) of 59%. Data from the mining area were used, together with default values for soil ingestion rates and infant body weights from the Contaminated Land Exposure Assessment (CLEA) model, to derive estimates of As intake for infants and small children (0-6 years old). Dose estimates of up to 3.53 microg kg(-1) bw day(-1) for housedusts and 2.43 microg kg(-1) bw day(-1) for garden soils were calculated, compared to the index dose used for the derivation of the SGV of 0.3 microg kg(-1) bw day(-1) (based on health risk assessments). The index dose was exceeded by 75% (18 out of 24) of the estimated As doses that were calculated for children aged 0-6 years, a group which is particularly at risk from exposure via soil and dust ingestion. The results of the present study support the concerns expressed by previous authors about the significant As contamination in southwest England and the potential implications for human health.

Air Pollution, Indoor↗

A Fourier transformation based method to mine peptide space for antimicrobial activity.

BACKGROUND: Naturally occurring antimicrobial peptides are currently being explored as potential candidate peptide drugs. Since antimicrobial peptides are part of the innate immune system of every living organism, it is possible to discover new candidate peptides using the available genomic and proteomic data. High throughput computational techniques could also be used to virtually scan the entire peptide space for discovering out new candidate antimicrobial peptides. RESULT: We have identified a unique indexing method based on biologically distinct characteristic features of known antimicrobial peptides. Analysis of the entries in the antimicrobial peptide databases, based on our indexing method, using Fourier transformation technique revealed a distinct peak in their power spectrum. We have developed a method to mine the genomic and proteomic data, for the presence of peptides with potential antimicrobial activity, by looking for this distinct peak. We also used the Euclidean metric to rank the potential antimicrobial peptides activity. We have parallelized our method so that virtually any given protein space could be data mined, in search of antimicrobial peptides. CONCLUSION: The results show that the Fourier transform based method with the property based coding strategy could be used to scan the peptide space for discovering new potential antimicrobial peptides.

Amino Acid Sequence↗

Effective mine risk education in war-zone areas--a shared responsibility.

The focus of this paper is effective health education and promotion in the field of mine awareness, or what has more recently been re-titled 'mine risk education'. According to the United Nations, mine risk education comprises educational activities that aim to reduce the risk of injury from landmine/unexploded ordnance (UXO) through raising awareness and promoting behavioural change and includes public information dissemination, education and training, and community mine action liaison. Specifically, this paper is an empirical study of mine risk education practices using data collected during the implementation of a mine risk education programme that commenced in Lao PDR in 1996 and is ongoing. In particular, it considers lessons learned from the programme's monitoring and evaluation process. The authors argue that in a country such as Lao PDR, where communities have lived with UXO infestation for over 25 years, more mine risk education is not necessarily needed. This paper concludes that common programmes of mine risk education using top-down educational methods, based on the assumption that ignorance of landmine/UXO risk is the key factor in mine accidents, are inadequate. Evidence from the literature on health promotion and the experience of the programme indicate that there is a need to supplement or replace existing common mine risk education practices with techniques that incorporate an understanding of the economic, social and political circumstances faced by communities at risk.

Blast Injuries↗

A hierarchical mixture of Markov models for finding biologically active metabolic paths using gene expression and protein classes.

With the recent development of experimental high-throughput techniques, the type and volume of accumulating biological data have extremely increased these few years. Mining from different types of data might lead us to find new biological insights. We present a new methodology for systematically combining three different datasets to find biologically active metabolic paths/patterns. This method consists of two steps: First it synthesizes metabolic paths from a given set of chemical reactions, which are already known and whose enzymes are co-expressed, in an efficient manner. It then represents the obtained metabolic paths in a more comprehensible way through estimating parameters of a probabilistic model by using these synthesized paths. This model is built upon an assumption that an entire set of chemical reactions corresponds to a Markov state transition diagram. Furthermore, this model is a hierarchical latent variable model, containing a set of protein classes as a latent variable, for clustering input paths in terms of existing knowledge of protein classes. We tested the performance of our method using a main pathway of glycolysis, and found that our method achieved higher predictive performance for the issue of classifying gene expressions than those obtained by other unsupervised methods. We further analyzed the estimated parameters of our probabilistic models, and found that biologically active paths were clustered into only two or three patterns for each expression experiment type, and each pattern suggested some new long-range relations in the glycolysis pathway.

Computer Simulation↗

Exposures in the alumina and primary aluminium industry: an historical review.

We reviewed specific chemical exposures and exposure assessment methods relating to published and unpublished epidemiological studies in the alumina and primary aluminium industry. Our focus was to review limitations in the current literature and make recommendations for future research. Although some of the exposures in the smelting of aluminium have been well characterised, particularly in potrooms, little has been published regarding the exposures in bauxite mining and alumina refining. Past epidemiological studies in the industry have concentrated on the smelting of aluminium, with many limitations in the methodology used in their exposure assessment. We found that in aluminium smelting, exposures to fluorides, coal tar pitch volatiles (CTPV) and sulfur dioxide (SO2) have tended to decrease in recent years, but insufficient information exists for the other known exposures. Although excess cancers have been found among workers in the smelting of aluminium, the exposure assessment methods in future studies need to be improved to better characterise possible causative agents. The small number of cohort studies has been a factor in the failure to identify clear exposure-response relationships for respiratory diseases. A dose-response relationship has been recently described for fluoride exposure and bronchial hyper-responsiveness, but whether fluorides are the causative agent, co-agent or simply markers for the causative agent(s) for potroom asthma, remains to be determined. Published epidemiological studies and quantitative exposure data for bauxite mining and alumina refining are virtually non-existent. Determination of possible exposure-response relationships for this part of the industry through improved exposure assessment methods should be the focus of future studies.

Aluminum↗