Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

RNAiDB and PhenoBlast: web tools for genome-wide phenotypic mapping projects.

RNA interference (RNAi) is being used in large-scale genomic studies as a rapid way to obtain in vivo functional information associated with specific genes. How best to archive and mine the complex data derived from these studies provides a series of challenges associated with both the methods used to elicit the RNAi response and the functional data gathered. RNAiDB (RNAi Database; http://www. rnai.org) has been created for the archival, distribution and analysis of phenotypic data from large-scale RNAi analyses in Caenorhabditis elegans. The database contains a compendium of publicly available data and provides information on experimental methods and phenotypic results, including raw data in the form of images and streaming time-lapse movies. Phenotypic summaries together with graphical displays of RNAi to gene mappings allow quick intuitive comparison of results from different RNAi assays and visualization of the gene product(s) potentially inhibited by each RNAi experiment based on multiple sequence analysis methods. RNAiDB can be searched using combinatorial queries and using the novel tool PhenoBlast, which ranks genes according to their overall phenotypic similarity. RNAiDB could serve as a model database for distributing and navigating in vivo functional information from large-scale systematic phenotypic analyses in different organisms.

Animals↗

STRING: known and predicted protein-protein associations, integrated and transferred across organisms.

A full description of a protein's function requires knowledge of all partner proteins with which it specifically associates. From a functional perspective, 'association' can mean direct physical binding, but can also mean indirect interaction such as participation in the same metabolic pathway or cellular process. Currently, information about protein association is scattered over a wide variety of resources and model organisms. STRING aims to simplify access to this information by providing a comprehensive, yet quality-controlled collection of protein-protein associations for a large number of organisms. The associations are derived from high-throughput experimental data, from the mining of databases and literature, and from predictions based on genomic context analysis. STRING integrates and ranks these associations by benchmarking them against a common reference set, and presents evidence in a consistent and intuitive web interface. Importantly, the associations are extended beyond the organism in which they were originally described, by automatic transfer to orthologous protein pairs in other organisms, where applicable. STRING currently holds 730,000 proteins in 180 fully sequenced organisms, and is available at http://string.embl.de/.

Databases, Protein↗

Fungal biology and agriculture: revisiting the field.

Plant pathology has made significant progress over the years, a process that involved overcoming a variety of conceptual and technological hurdles. Descriptive mycology and the advent of chemical plant-disease management have been followed by biochemical and physiological studies of fungi and their hosts. The later establishment of biochemical genetics along with the introduction of DNA-mediated transformation have set the stage for dissection of gene function and advances in our understanding of fungal cell biology and plant-fungus interactions. Currently, with the advent of high-throughput technologies, we have the capacity to acquire vast data sets that have direct relevance to the numerous subdisciplines within fungal biology and pathology. These data provide unique opportunities for basic research and for engineering solutions to important agricultural problems. However, we also are faced with the challenge of data organization and mining to analyze the relationships between fungal and plant genomes and to elucidate the physiological function of pertinent DNA sequences. We present our perspective of fungal biology and agriculture, including administrative and political challenges to plant protection research.

Agriculture↗

Prevalence survey of respiratory abnormalities in New Mexico uranium miners.

To obtain additional data concerning uranium mining and nonmalignant respiratory diseases, we conducted a prevalence survey of 192 long-term New Mexico uranium miners. Survey procedures included spirometry, completion of a respiratory symptoms questionnaire, physical examination and interpretation of available chest x rays. Total duration of underground uranium mining was used as the exposure index. Of the major respiratory symptoms, only the prevalence of dyspnea increased significantly with duration of uranium mining. With linear multiple-regression analysis, small but statistically significant effects of mining were found for two spirometric parameters, the forced expiratory volume in one sec and the maximal midexpiratory flow. By the 1980 International Labor Organization (ILO) U/C classification, 12 of 143 participants with x rays available for interpretation had at least category 1/0 pneumoconiosis. The opacities were predominantly nodular and compatible with silicosis.

Adult↗

Estimates of lifetime lung cancer risks resulting from Rn progeny exposure.

Data on five mining populations exposed to Rn progeny have been used to estimate the lifetime risk of lung cancer resulting from occupational and environmental exposure under current standards. Slopes of dose-response relations for lung cancer show a tendency to decrease with increasing dose. Our best estimate of curvilinearity is given by raising dose to the power 0.92 +/- 0.07, but the improvement in fit beyond simple linearity is not significant. On the other hand, the addition of a cell-killing term significantly improves the fit of the linear model. In any event, linear extrapolation is unlikely to underestimate the excess risk at low doses by more than a factor of 1.5. However, these inferences about curvilinearity are highly subject to error from the choice of reference populations, dosimetry, and latency. Under the linear-cell-killing model, our best estimate of excess relative risk is 2.28 +/- 0.35 per 100 working level month (WLM) (a doubling dose of 44 WLM). Attributable risks in these five studies range from 3.4-17.8 per 10(6) person-yr WLM-1. Risks from Rn progeny appear to interact with age and smoking in a form intermediate between additive and multiplicative. The "relative risk" model is therefore preferable for projecting lifetime risks, but life-table projections are described for a wide variety of assumptions. Our best estimate of the effect of a 50-yr occupational exposure to 4 WLM yr-1 is 130 excess lung cancer deaths per 1000 persons (0.65 per 1000 person-WLM), with a range from 60-250 per 1000. Similar calculations for lifetime exposure to an additional 0.02 working level (WL) beyond normal background produces an estimate of 20 excess lung cancers per 1000 persons.

Adolescent↗

Application of genome-wide gene expression profiling by high-density DNA arrays to the treatment and study of inflammatory bowel disease.

Identification of factors involved in the initiation, amplification, and perpetuation of the chronic immune response and the identification of markers for the characterization of patient subgroups remain critical objectives for ongoing research in inflammatory bowel disease (IBD). The Human Genome Project and the development of the expressed sequence tag (EST) clone collection and database have made possible a new revolution in gene expression analysis. Instead of measuring one or a few genes, parallel DNA microarrays are capable of simultaneously measuring expression of thousands of genes, providing a glimpse into the logic and functional grouping of gene programs encoded by our genome. Applied to clinical specimens from affected and normal individuals, this methodology has the potential to provide a new level of information about disease pathogenesis not previously possible. Two dominant platforms for the construction of high-density microarrays have emerged: cDNA arrays and GeneChips. The first involves robotic spotting of DNA molecules, often derived from EST clone collections, onto a suitable solid phase matrix such as a glass slide. The second involves direct in situ synthesis of sets of gene-specific oligonucleotides on a silicon wafer by an eloquent derivative of the photolithography process. Both cDNA and oligonucleotide arrays are interrogated by hybridization with a fluorescent-labeled cDNA or cRNA representation of the original tissue mRNA. This enables measurement of the expression levels for thousands of mucosal genes in a single experiment. These technologies have recently become less expensive and more widely accessible to all researchers. This review details the principles and methods behind DNA array technology, data analysis and mining, and potential application to research and treatment of IBD.

Gene Expression Profiling↗

A new set of Arabidopsis expressed sequence tags from developing seeds. The metabolic pathway from carbohydrates to seed oil.

Large-scale single-pass sequencing of cDNAs from different plants has provided an extensive reservoir for the cloning of genes, the evaluation of tissue-specific gene expression, markers for map-based cloning, and the annotation of genomic sequences. Although as of January 2000 GenBank contained over 220,000 entries of expressed sequence tags (ESTs) from plants, most publicly available plant ESTs are derived from vegetative tissues and relatively few ESTs are specifically derived from developing seeds. However, important morphogenetic processes are exclusively associated with seed and embryo development and the metabolism of seeds is tailored toward the accumulation of economically valuable storage compounds such as oil. Here we describe a new set of ESTs from Arabidopsis, which has been derived from 5- to 13-d-old immature seeds. Close to 28,000 cDNAs have been screened by DNA/DNA hybridization and approximately 10,500 new Arabidopsis ESTs have been generated and analyzed using different bioinformatics tools. Approximately 40% of the ESTs currently have no match in dbEST, suggesting many represent mRNAs derived from genes that are specifically expressed in seeds. Although these data can be mined with many different biological questions in mind, this study emphasizes the import of photosynthate into developing embryos, its conversion into seed oil, and the regulation of this pathway.

Arabidopsis↗

Analysing the developing brain transcriptome with the GenePaint platform.

We discuss technical means by which the complexity of gene and protein signalling cascades can be projected onto the complex structure of the mammalian brain. We argue that this requires both robotics and novel computational tools to register images of gene expression, annotate expression patterns and quantify gene expression. When sufficiently enriched and detailed, such gene expression/neuroanatomical atlases are hypothesis-generating tools and contain in themselves much of the information needed to investigate function in normal and genetically or otherwise modified brains. To be successful and useful, data-rich and comprehensive gene expression/neuroanatomical atlases have to be web accessible and structured in a way that allows the application of data exploration and mining tools.

Animals↗

A novel isoform of tensin-1 promotes actin filament assembly for efficient erythroblast enucleation.

Mammalian red blood cells are generated via a terminal erythroid differentiation pathway culminating in cell polarization and enucleation. Actin filament (F-actin) polymerization is critical for enucleation, but the underlying molecular regulatory mechanisms remain poorly understood. We used publicly available RNA sequencing and proteomic data sets to mine for actin-regulatory factors differentially expressed during human erythroid differentiation and discovered that a focal adhesion (FA) protein, tensin-1 (TNS1), dramatically increases in expression late in differentiation. Remarkably, we found that differentiating human CD34+ cells express a novel truncated form of TNS1 (erythroid TNS1 [eTNS1]; Mr ∼125 kDa) missing the N-terminal half of the protein containing the actin-binding domain, due to an internal messenger RNA translation start site resulting in a unique exon 1E. The region upstream of eTNS1 has features of an active erythroid promoter, demonstrating increasing chromatin accessibility during terminal differentiation, paralleling increasing gene expression. Sequence comparisons across species indicate that eTNS1 is expressed in humans and nonhuman primates, but not in zebrafish, mice, or other rodents. Confocal microscopy showed that eTNS1 localized to the cytoplasm during terminal erythroid differentiation but, surprisingly, did not appear to form focal adhesions nor to colocalize with F-actin. Knockout of eTNS1 did not affect terminal differentiation or assembly of the spectrin membrane skeleton but led to reduced F-actin assembly and abnormal organization in polarized and enucleating erythroblasts, resulting in impaired enucleation efficiency. We conclude that eTNS1 is a novel regulator of F-actin during human erythroid terminal differentiation that is required for efficient enucleation.

Humans↗

Performance of a genetic algorithm for mass spectrometry proteomics.

BACKGROUND: Recently, mass spectrometry data have been mined using a genetic algorithm to produce discriminatory models that distinguish healthy individuals from those with cancer. This algorithm is the basis for claims of 100% sensitivity and specificity in two related publicly available datasets. To date, no detailed attempts have been made to explore the properties of this genetic algorithm within proteomic applications. Here the algorithm's performance on these datasets is evaluated relative to other methods. RESULTS: In reproducing the method, some modifications of the algorithm as it is described are necessary to get good performance. After modification, a cross-validation approach to model selection is used. The overall classification accuracy is comparable though not superior to other approaches considered. Also, some aspects of the process rely upon random sampling and thus for a fixed dataset the algorithm can produce many different models. This raises questions about how to choose among competing models. How this choice is made is important for interpreting sensitivity and specificity results as merely choosing the model with lowest test set error rate leads to overestimates of model performance. CONCLUSIONS: The algorithm needs to be modified to reduce variability and care must be taken in how to choose among competing models. Results derived from this algorithm must be accompanied by a full description of model selection procedures to give confidence that the reported accuracy is not overstated.

Algorithms↗

QPath: a method for querying pathways in a protein-protein interaction network.

BACKGROUND: Sequence comparison is one of the most prominent tools in biological research, and is instrumental in studying gene function and evolution. The rapid development of high-throughput technologies for measuring protein interactions calls for extending this fundamental operation to the level of pathways in protein networks. RESULTS: We present a comprehensive framework for protein network searches using pathway queries. Given a linear query pathway and a network of interest, our algorithm, QPath, efficiently searches the network for homologous pathways, allowing both insertions and deletions of proteins in the identified pathways. Matched pathways are automatically scored according to their variation from the query pathway in terms of the protein insertions and deletions they employ, the sequence similarity of their constituent proteins to the query proteins, and the reliability of their constituent interactions. We applied QPath to systematically infer protein pathways in fly using an extensive collection of 271 putative pathways from yeast. QPath identified 69 conserved pathways whose members were both functionally enriched and coherently expressed. The resulting pathways tended to preserve the function of the original query pathways, allowing us to derive a first annotated map of conserved protein pathways in fly. CONCLUSION: Pathway homology searches using QPath provide a powerful approach for identifying biologically significant pathways and inferring their function. The growing amounts of protein interactions in public databases underscore the importance of our network querying framework for mining protein network data.

Algorithms↗

Sequence- and structure-based protein function prediction from genomic information.

Existing functional annotation transfer is fraught with inaccuracies that may hinder forward interpretation and mining of genomic data. Hand-curation of the annotation placed into databases is not practical. In lieu of experimental evidence, computational biological approaches offer high-throughput tools to predict function accurately; however, these methods are still notably deficient in defining and describing the complexity of protein function. Enriching genomic sequences obtained from sequencing efforts and expression array methods with protein function information and classification will be an efficient first step for incorporating genomic data into drug discovery programs.

Computational Biology↗

Sequence similarity as a predictor of the transmembrane topology of membrane-intrinsic subunits of bacterial respiratory chain enzymes.

Integral membrane proteins usually have a predominantly alpha-helical secondary structure in which transmembrane segments are connected by membrane-extrinsic loops. Although a number of membrane protein structures have been reported in recent years, in most cases transmembrane topologies are initially predicted using a variety of theoretical techniques, including hydropathy analyses and the "positive inside" rule. We have explored the use of plots of the distribution of sequence similarity within families of membrane proteins comprising homeomorphic domains as a new method for the prediction/verification of the orientation of transmembrane topology models within certain families of multimeric respiratory chain enzymes. Within such proteins, analyses of sequence similarity can: i) identify heme and/or quinol binding sites; ii) identify potential electron-transfer conduits to/from prosthetic groups; and iii) locate regions defining potential subunit-subunit interactions. We mined emerging bioinformatic data for sequences of 11 families of membrane-intrinsic proteins that are part of multimeric respiratory chain complexes that also have membrane-extrinsic subunits. The sequences of each family were then aligned and the resultant alignments converted into a graphical format recording an empirical measure of the sequence similarity plotted versus residue position. In each case, this plot was compared to the predicted transmembrane topology. With one exception, there is a strong correlation between the existence

Amino Acid Sequence↗

Discovering the interaction propensities of amino acids and nucleotides from protein-RNA complexes.

With the availability of many genome sequences, the mining of biological data is attracting much attention, most of it limited to the sequences of macromolecules. Sequence data are easy to analyze as they can be treated as strings of characters, whereas the structure of a macromolecule is much more complex. We developed a set of algorithms to analyze the structures of protein-RNA complexes at the atomic level and used them to analyze protein-RNA interactions using structural data on 51 protein-RNA complexes. The analysis revealed, among other things, that: (1) polar and charged amino acids have a strong tendency to interact with nucleotides, (2) arginine and asparagine tend to hydrogen bond with uracil, and (3) histidine favors uracil in water-mediated bonding with RNA. We analyzed a large set of structural data of protein-RNA complexes involving water-mediated hydrogen bonds as well as direct hydrogen bonds. The interaction patterns discovered from the analysis provide useful information for predicting the structure of RNA that binds proteins, and of proteins that bind RNA.

Algorithms↗

[The abnormity and clinical value of tyrosine kinase signaling pathway in lung adenocarcinoma].

OBJECTIVE: To study the abnormity and clinical value of tyrosine kinase signaling pathway in lung adenocarcinoma. METHODS: Total RNA was extracted form 10 lung adenocarcinoma samples and matched normal tissues. cDNA was labeled fluorescence during reverse transcription and hybridized on 20 slides of microarray with 13 824 genes. After washed and scanned, the cy3:cy5 ratios were obtained by computer analysis. Each experiment was repeated two times and the mean value was chosen. Cluster analysis and discriminant analysis were used to mine the pool data. The results were combined with clinical prognostic factors for further analysis. RESULTS: According to the abnormity of tyrosine kinase signaling pathway in lung adenocarcinoma, sample cluster divided the 10 samples into 3 groups and gene cluster divided all genes into 5 groups. Different prognosis samples had different characteristic genes. The patients with the best prognosis overexpressed some genes related to immunity and inhibition proliferation, and those with the worst prognosis overexpressed some genes related to proliferation while the genes related to immunity and inhibition proliferation were expressed at a low level. CONCLUSION: There are 3 type abnormity of tyrosine kinase signaling pathway in lung adenocarcinoma that are related to different biological behavior and different prognosis.

Adenocarcinoma↗

Landscape of current toxicity databases and database standards.

Having readily available historical information for modeling toxicity has become important throughout the various stages of research and development. The high cost of late-phase attrition and recent international regulatory legislations have made even more acute the need to be able to mine the fragmented data and information available across diverse databases. In addition, the general trend to accelerate regulatory processes globally makes the effective use of existing data an imperative. To achieve efficient screening, develop profiles and gain the ability to cross reference, databases must be interoperated to allow data exchange and integration. Several database standards and controlled vocabulary initiatives have been used in the development of toxicity data models to transform the current landscape. This review describes the major databases of toxicological information now available, and provides a simple example of standardization that demonstrates the benefits of a toxicity database containing such qualified data.

Animals↗

INFOBIOMED: European Network of Excellence on Biomedical Informatics to support individualised healthcare.

INFOBIOMED is an European Network of Excellence (NoE) funded by the Information Society Directorate-General of the European Commission (EC). A consortium of European organizations from ten different countries is involved within the network. Four pilots, all related to linking clinical and genomic information, are being carried out. From an informatics perspective, various challenges, related to data integration and mining, are included.

Biomedical Research↗

Respirable dust exposures in U.S. surface coal mines (1982-1986).

Exposure of miners to respirable coal mine dust and to respirable quartz silica at surface coal mines in the United States during 1982-1986 were evaluated by job category using data collected by coal mine operators and Mine Safety and Health Administration (MSHA) inspectors. Average coal mine dust concentrations were usually well below the MSHA Permissible Exposure Limit (PEL) for all job categories, but at least 10% of the samples obtained from some coal preparation plant job areas and most drilling job areas had concentrations that exceeded the 2.0 mg/m3 limit. In contrast, a very high proportion of samples from surface mine driller areas exceeded the quartz PEL. Of all samples collected for highwall drill operators and helpers, 78% and 77%, respectively, were greater than the 0.1 mg/m3 quartz exposure limit (average concentrations were .32 and .36 mg/m3, respectively). Although MSHA compliance data may not be entirely adequate for assessing chronic exposure to quartz, these data and the results of other NIOSH studies nonetheless indicate excessive exposure to silica in a group of surface coal miners.

Air Pollutants, Occupational↗