Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

A genomics approach to the early stages of triterpene saponin biosynthesis in Medicago truncatula.

The saponins of the model legume Medicago truncatula are glycosides of at least five different triterpene aglycones: soyasapogenol B, soyasapogenol E, medicagenic acid, hederagenin and bayogenin. These aglycones are most likely derived from beta-amyrin, a product of the cyclization of 2,3-oxidosqualene. Mining M. truncatula EST data sets led to the identification of sequences putatively encoding three early enzymes of triterpene aglycone formation: squalene synthase (SS), squalene epoxidase (SE), and beta-amyrin synthase (beta-AS). SS was functionally characterized by expression in Escherichia coli, two forms of SE by complementation of the yeast erg1 mutant, and beta-AS by expression in yeast. Beta-amyrin was the sole product of the cyclization of squalene epoxide by the recombinant M. truncatulabeta-AS, as judged by GC-MS and NMR. Transcripts encoding beta-AS, SS and one form of SE were strongly and co-ordinately induced, associated with accumulation of triterpenes, upon exposure of M. truncatula cell suspension cultures to methyl jasmonate. Sterol composition remained unaffected by jasmonate treatment. Molecular verification of induction of the triterpene pathway in a cell culture system provides a new tool for saponin pathway gene discovery by DNA array-based approaches.

Acetates↗

Chronic cocaine-mediated changes in non-human primate nucleus accumbens gene expression.

Chronic cocaine use elicits changes in the pattern of gene expression within reinforcement-related, dopaminergic regions. cDNA hybridization arrays were used to illuminate cocaine-regulated genes in the nucleus accumbens (NAcc) of non-human primates (Macaca fascicularis; cynomolgus macaque), treated daily with escalating doses of cocaine over one year. Changes seen in mRNA levels by hybridization array analysis were confirmed at the level of protein (via specific immunoblots). Significantly up-regulated genes included: protein kinase A alpha catalytic subunit (PKA(calpha)); cell adhesion tyrosine kinase beta (PYK2); mitogen activated protein kinase kinase 1 (MEK1); and beta-catenin. While some of these changes exist in previously described cocaine-responsive models, others are novel to any model of cocaine use. All of these adaptive responses coexist within a signaling scheme that could account for known inductions of genes(e.g. fos and jun proteins, and cyclic AMP response element binding protein) previously shown to be relevant to cocaine's behavioral actions. The complete data set from this experiment has been posted to the newly created Drug and Alcohol Abuse Array Data Consortium (http://www.arraydata.org) for mining by the general research community.

Animals↗

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis↗

The derivation of estimated dust exposures for U.S. coal miners working before 1970.

A number of reports on the prevalence of coal workers' pneumoconiosis in U.S. coal miners have been published, yet very little is known about the relationship between dust exposure and pneumoconiosis levels in the U.S. This report describes the derivation of cumulative dust exposure estimates by back-extrapolation of data processed by the Mine Safety and Health Administration after 1970 by using a ratio of dust concentrations based on information collected during environmental surveys at certain U.S. mines by the Bureau of Mines between 1968 and 1969. Cumulative personal dust exposure estimates were calculated by using occupational histories obtained from the miners and job-specific estimates of dust concentration. In other reports, the resulting estimated exposures have been shown to correlate well with various measures of respiratory morbidity.

Adult↗

Contributions of domesticated plant studies to our understanding of plant evolution.

BACKGROUND: Plant evolutionary theory has been greatly enriched by studies on crop species. Over the last century, important information has been generated on many aspects of population biology, speciation and polyploid genetics. SCOPE: Searches for quantitative trait loci (QTL) in crop species have uncovered numerous blocks of genes that have dramatic effects on adaptation, particularly during the domestication process. Many of these QTL have epistatic and pleiotropic effects making rapid evolutionary change possible. Most of the pioneering work on the molecular basis of self-incompatibility has been conducted on crop species, along with the sequencing of the phytopathogenic resistance genes (R genes) responsible for the 'gene-to-gene' relations of coevolution observed in host-pathogen relationships. Some of the better examples of co-adaptation and early acting inbreeding depression have also been elucidated in crops. Crop-wild progenitor interactions have provided rich opportunities to study the evolution of novel adaptations subsequent to hybridization. Most crop/wild F1 hybrids have reduced fitness, but in some instances the crop relatives have acquired genes that make them more efficient weeds through crop mimicry. Studies on autopolyploid alfalfa and potato have uncovered the means by which polyploid gametes are formed and have led to hypotheses about how multiallelic interactions are associated with fitness and self-fertility. Research on the cole crops and wheat has discovered that newly formed polyploids can undergo dramatic genome rearrangements that could lead to rapid evolutionary change. CONCLUSIONS: Many more important evolutionary discoveries are on the horizon, now that the whole genome sequence is available of the two major subspecies of rice Oryza sativa ssp. japonica and O. sativa ssp. indica. The rice sequence data can be used to study the origin of genes and gene families, track rates of sequence divergence over time, and provide hints about how genes evolve and generate products with novel biological properties. The rice sequence data has already been mined to show that transposable elements often carry fragments of cellular genes. This type of genome shuffling could play a role in creating novel, reorganized genes with new adaptive properties.

Biological Evolution↗

Farm animal genomics and informatics: an update.

Farm animal genomics is of interest to a wide audience of researchers because of the utility derived from understanding how genomics and proteomics function in various organisms. Applications such as xenotransplantation, increased livestock productivity, bioengineering new materials, products and even fabrics are several reasons for thriving farm animal genome activity. Currently mined in rapidly growing data warehouses, completed genomes of chicken, fish and cows are available but are largely stored in decentralized data repositories. In this paper, we provide an informatics primer on farm animal bioinformatics and genome project resources which drive attention to the most recent advances in the field. We hope to provide individuals in biotechnology and in the farming industry with information on resources and updates concerning farm animal genome projects.

Animals↗

Indoor radon progeny exposure in the Florida phosphate mining region: a review.

This paper reviews the data generated in studies of land radioactivity and indoor airborne radon progeny associated with mined and reclaimed phosphate lands in Florida. Highest indoor radon progeny levels are associated with the slab-on-grade type of construction. Concentrations exceeding 0.03 WL are associated with overburden soils, deposits and fill, while concentrations up to about 0.03 WL are associated with tailings. The lower limit for distinguishing increases above non-enhanced natural concentrations is on the order of 0.01-0.02 WL. Results of this study show that about 25% of the land produced by present methods of mining and reclamation practices would require restrictions on the type of construction or would require special construction methods. It is projected that with modification of mining and tailings disposal practices, virtually all of the land produced by mining and reclamation would be satisfactory for unrestricted construction use.

Background Radiation↗

AffyMiner: mining differentially expressed genes and biological knowledge in GeneChip microarray data.

BACKGROUND: DNA microarrays are a powerful tool for monitoring the expression of tens of thousands of genes simultaneously. With the advance of microarray technology, the challenge issue becomes how to analyze a large amount of microarray data and make biological sense of them. Affymetrix GeneChips are widely used microarrays, where a variety of statistical algorithms have been explored and used for detecting significant genes in the experiment. These methods rely solely on the quantitative data, i.e., signal intensity; however, qualitative data are also important parameters in detecting differentially expressed genes. RESULTS: AffyMiner is a tool developed for detecting differentially expressed genes in Affymetrix GeneChip microarray data and for associating gene annotation and gene ontology information with the genes detected. AffyMiner consists of the functional modules, GeneFinder for detecting significant genes in a treatment versus control experiment and GOTree for mapping genes of interest onto the Gene Ontology (GO) space; and interfaces to run Cluster, a program for clustering analysis, and GenMAPP, a program for pathway analysis. AffyMiner has been used for analyzing the GeneChip data and the results were presented in several publications. CONCLUSION: AffyMiner fills an important gap in finding differentially expressed genes in Affymetrix GeneChip microarray data. AffyMiner effectively deals with multiple replicates in the experiment and takes into account both quantitative and qualitative data in identifying significant genes. AffyMiner reduces the time and effort needed to compare data from multiple arrays and to interpret the possible biological implications associated with significant changes in a gene's expression.

Algorithms↗

An expressed sequence tag (EST) library from developing fruits of an Hawaiian endemic mint (Stenogyne rugosa, Lamiaceae): characterization and microsatellite markers.

BACKGROUND: The endemic Hawaiian mints represent a major island radiation that likely originated from hybridization between two North American polyploid lineages. In contrast with the extensive morphological and ecological diversity among taxa, ribosomal DNA sequence variation has been found to be remarkably low. In the past few years, expressed sequence tag (EST) projects on plant species have generated a vast amount of publicly available sequence data that can be mined for simple sequence repeats (SSRs). However, these EST projects have largely focused on crop or otherwise economically important plants, and so far only few studies have been published on the use of intragenic SSRs in natural plant populations. We constructed an EST library from developing fleshy nutlets of Stenogyne rugosa principally to identify genetic markers for the Hawaiian endemic mints. RESULTS: The Stenogyne fruit EST library consisted of 628 unique transcripts derived from 942 high quality ESTs, with 68% of unigenes matching Arabidopsis genes. Relative frequencies of Gene Ontology functional categories were broadly representative of the Arabidopsis proteome. Many unigenes were identified as putative homologs of genes that are active during plant reproductive development. A comparison between unigenes from Stenogyne and tomato (both asterid angiosperms) revealed many homologs that may be relevant for fruit development. Among the 628 unigenes, a total of 44 potentially useful microsatellite loci were predicted. Several of these were successfully tested for cross-transferability to other Hawaiian mint species, and at least five of these demonstrated interesting patterns of polymorphism across a large sample of Hawaiian mints as well as close North American relatives in the genus Stachys. CONCLUSION: Analysis of this relatively small EST library illustrated a broad GO functional representation. Many unigenes could be annotated to involvement in reproductive development. Furthermore, first tests of microsatellite primer pairs have proven promising for the use of Stenogyne rugosa EST SSRs for evolutionary and phylogeographic studies of the Hawaiian endemic mints and their close relatives. Given that allelic repeat length variation in developmental genes of other organisms has been linked with morphological evolution, these SSRs may also prove useful for analyses of phenotypic differences among Hawaiian mints.

5' Untranslated Regions↗

Genome-wide review of transcriptional complexity in mouse protein kinases and phosphatases.

BACKGROUND: Alternative transcripts of protein kinases and protein phosphatases are known to encode peptides with altered substrate affinities, subcellular localizations, and activities. We undertook a systematic study to catalog the variant transcripts of every protein kinase-like and phosphatase-like locus of mouse http://variant.imb.uq.edu.au. RESULTS: By reviewing all available transcript evidence, we found that at least 75% of kinase and phosphatase loci in mouse generate alternative splice forms, and that 44% of these loci have well supported alternative 5' exons. In a further analysis of full-length cDNAs, we identified 69% of loci as generating more than one peptide isoform. The 1,469 peptide isoforms generated from these loci correspond to 1,080 unique Interpro domain combinations, many of which lack catalytic or interaction domains. We also report on the existence of likely dominant negative forms for many of the receptor kinases and phosphatases, including some 26 secreted decoys (seven known and 19 novel: Alk, Csf1r, Egfr, Epha1, 3, 5,7 and 10, Ephb1, Flt1, Flt3, Insr, Insrr, Kdr, Met, Ptk7, Ptprc, Ptprd, Ptprg, Ptprl, Ptprn, Ptprn2, Ptpro, Ptprr, Ptprs, and Ptprz1) and 13 transmembrane forms (four known and nine novel: Axl, Bmpr1a, Csf1r, Epha4, 5, 6 and 7, Ntrk2, Ntrk3, Pdgfra, Ptprk, Ptprm, Ptpru). Finally, by mining public gene expression data (MPSS and microarrays), we confirmed tissue-specific expression of ten of the novel isoforms. CONCLUSION: These findings suggest that alternative transcripts of protein kinases and phosphatases are produced that encode different domain structures, and that these variants are likely to play important roles in phosphorylation-dependent signaling pathways.

Alternative Splicing↗

BoCaTFBS: a boosted cascade learner to refine the binding sites suggested by ChIP-chip experiments.

Comprehensive mapping of transcription factor binding sites is essential in postgenomic biology. For this, we propose a mining approach combining noisy data from ChIP (chromatin immunoprecipitation)-chip experiments with known binding site patterns. Our method (BoCaTFBS) uses boosted cascades of classifiers for optimum efficiency, in which components are alternating decision trees; it exploits interpositional correlations; and it explicitly integrates massive negative information from ChIP-chip experiments. We applied BoCaTFBS within the ENCODE project and showed that it outperforms many traditional binding site identification methods (for instance, profiles).

Algorithms↗

Signs of latent handedness in families.

Luria argued that several behaviors suggested a "latent" handedness. He suggested that such things as the way in which one crossed one's arms or clasped one's hands might reflect a latent preference for the left hand. Arm-folding refers to the preferential tendency for individuals to fold one forearm over the other, whereas hand-clasping refers to the preferential tendency for individuals to clasp the hands together. We investigated hand-clasping and arm-folding in 292 families (mother, father, and offspring). In this study about 55% of the population are left-hand-claspers, 44% are right-hand-claspers, and the remaining 1% report that they have no preference or are indifferent. About 54% of the population are left-arm-folders, 42% are right-arm-folders, and the remaining 4% report that they have no preference or are indifferent. Familial data suggest that hand-clasping and arm-folding may be under genetic control: although the data do not fit any straightforward recessive or dominant Mendelian model, they are compatible with the type of model invoking fluctuating asymmetry which has been used to explain the inheritance of handedness. It is possible that hand-clasping and arm-folding as well as leg-crossing may be idiosyncrasies due to or influenced by physical bilateral differences in the hands or arms. All family data including others and mine together (arithmetical sum) suggest a genetic contribution, although environmental influences are also evident.

Adult↗

Croatian experience with the landmines: deaths and injuries from the landmines in the area of Sisak during five-years period (1995-2000).

After the war in Croatia, thousands of landmines were left on the fields and because of this today we have many casualties. Area around Sisak is a region where a large number of landmines were laid and accidents are frequent. The authors of this article present results of retrospective analyses of accidents (death, injuries and type of injuries) which occurred after the military operation "Oluja" since August 1995 until March 2000. Data is collected from local hospital and police department in Sisak and compared with data obtained from Croatian Mines Action Center.

Blast Injuries↗

Antisense oligonucleotides for target validation and gene function determination.

Antisense technology is attracting attention from the biotechnology and pharmaceutical industries because it provides a high-throughput and systematic approach to drug target validation and gene function discovery. Antisense represents a logical approach to gene function analysis and discovery as it is specific, broadly applicable, and can be designed with minimal information (ie, expressed sequence tags). This technology in combination with other emerging technologies (eg, microarray technology), will enable efficient 'mining' of the sequence data generated by the human genome project. This review addresses recent advances in the antisense field and discusses the potential use of antisense technology for functional genomics approaches.

Journal Article↗

In silico approaches to mechanistic and predictive toxicology: an introduction to bioinformatics for toxicologists.

Bioinformatics, or in silico biology, is a rapidly growing field that encompasses the theory and application of computational approaches to model, predict, and explain biological function at the molecular level. This information rich field requires new skills and new understanding of genome-scale studies in order to take advantage of the rapidly increasing amount of sequence, expression, and structure information in public and private databases. Toxicologists are poised to take advantage of the large public databases in an effort to decipher the molecular basis of toxicity. With the advent of high-throughput sequencing and computational methodologies, expressed sequences can be rapidly detected and quantitated in target tissues by database searching. Novel genes can also be isolated in silico, while their function can be predicted and characterized by virtue of sequence homology to other known proteins. Genomic DNA sequence data can be exploited to predict target genes and their modes of regulation, as well as identify susceptible genotypes based on single nucleotide polymorphism data. In addition, highly parallel gene expression profiling technologies will allow toxicologists to mine large databases of gene expression data to discover molecular biomarkers and other diagnostic and prognostic genes or expression profiles. This review serves to introduce to toxicologists the concepts of in silico biology most relevant to mechanistic and predictive toxicology, while highlighting the applicability of in silico methods using select examples.

Cluster Analysis↗

Constructing biological networks through combined literature mining and microarray analysis: a LMMA approach.

MOTIVATION: Network reconstruction of biological entities is very important for understanding biological processes and the organizational principles of biological systems. This work focuses on integrating both the literatures and microarray gene-expression data, and a combined literature mining and microarray analysis (LMMA) approach is developed to construct gene networks of a specific biological system. RESULTS: In the LMMA approach, a global network is first constructed using the literature-based co-occurrence method. It is then refined using microarray data through a multivariate selection procedure. An application of LMMA to the angiogenesis is presented. Our result shows that the LMMA-based network is more reliable than the co-occurrence-based network in dealing with multiple levels of KEGG gene, KEGG Orthology and pathway. AVAILABILITY: The LMMA program is available upon request.

Abstracting and Indexing↗

Comparison of mass concentrations determined with personal respirable coal mine dust samplers operating at 1.2 liters per minute and the Casella 113A gravimetric sampler (MRE).

Measuring respirable dust concentrations in coal mine environments is currently done using approved personal respirable dust sampling equipment operating at a flow rate of 2.0 liters per minute (Lpm). Measurements made with approved coal mine dust sampling equipment are converted to equivalent concentrations that would be obtained with a Mining Research Establishment (MRE) instrument known as the Casella 113A gravimetric sampler, using a conversion factor of 1.38. NIOSH has recently recommended that coal mine dust samplers (CMDS) used to measure respirable dust levels in mine environments be operated at 1.2 Lpm and measured concentrations be multiplied by 0.91 to obtain an equivalent MRE concentration. The purpose of this recommendation is to reduce systematic error caused by the variation in mine dust distributions. This paper presents and discusses data collected in the laboratory and in underground coal mines to evaluate the recommended 1.2 Lpm flow rate and 0.91 conversion factor. Comparative measurements were obtained in the laboratory with the CMDS operating at 2.0 and 1.2 liters per minute and the MRE, using aerosols of coal and limestone dust of varying particle size distribution. Similar comparative measurements were made in a number of underground coal mines at locations with environments having particle size distributions representative of different underground mining operations. It was concluded from this study that there is no significant change in the variability associated with the constant factor used to convert respirable dust measurements, obtained with approved respirable CMDS, to equivalent MRE measurements when the flow rate of the CMDS is reduced from 2.0 to 1.2 Lpm.

Coal↗

EXPANDER--an integrative program suite for microarray data analysis.

BACKGROUND: Gene expression microarrays are a prominent experimental tool in functional genomics which has opened the opportunity for gaining global, systems-level understanding of transcriptional networks. Experiments that apply this technology typically generate overwhelming volumes of data, unprecedented in biological research. Therefore the task of mining meaningful biological knowledge out of the raw data is a major challenge in bioinformatics. Of special need are integrative packages that provide biologist users with advanced but yet easy to use, set of algorithms, together covering the whole range of steps in microarray data analysis. RESULTS: Here we present the EXPANDER 2.0 (EXPression ANalyzer and DisplayER) software package. EXPANDER 2.0 is an integrative package for the analysis of gene expression data, designed as a 'one-stop shop' tool that implements various data analysis algorithms ranging from the initial steps of normalization and filtering, through clustering and biclustering, to high-level functional enrichment analysis that points to biological processes that are active in the examined conditions, and to promoter cis-regulatory elements analysis that elucidates transcription factors that control the observed transcriptional response. EXPANDER is available with pre-compiled functional Gene Ontology (GO) and promoter sequence-derived data files for yeast, worm, fly, rat, mouse and human, supporting high-level analysis applied to data obtained from these six organisms. CONCLUSION: EXPANDER integrated capabilities and its built-in support of multiple organisms make it a very powerful tool for analysis of microarray data. The package is freely available for academic users at http://www.cs.tau.ac.il/~rshamir/expander.

Algorithms↗