Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

PEDF and the serpins: phylogeny, sequence conservation, and functional domains.

Pigment epithelium derived factor (PEDF) is non-inhibitory serpin with neurotrophic and antiangiogenic functions. In this study, we have assembled PEDF sequences for 9 additional species by data base mining and performed cross-species alignment for 14 PEDF sequences to identify conserved structural domains. We found evolutionary conservation of a leader sequence, a single C-terminal glycosylation site, collagen-binding residues, and four specific conserved PEDF peptides. The C-terminus, 384--415 and an N-terminal region 78--95, show close homology with many other serpins, and there is strong conservation of 39 of 51 consensus key residues involved in serpin structure and function. Two peptide regions, 40--67 and 277--301, are unique to PEDF but conserved in all species. Conserved residues at the N-terminus, helix d (hD), and helix A (hA) of PEDF form a structure similar to the heparin-binding groove of other serpins. We identified a motif in PEDF that is homologous to the nuclear localization signals of other proteins. A bitopographical localization of PEDF was confirmed by immunocytochemistry and Western blots. Our results suggest that secretion is required for PEDF's activity, that PEDF can migrate to the nucleus, and that PEDF has structural and functional features more common with inhibitory serpins.

Amino Acid Motifs↗

Regulation of murine TGFbeta2 by Pax3 during early embryonic development.

Previously our laboratory identified TGFbeta2 as a potential downstream target of Pax3 by utilizing microarray analysis and promoter data base mining (Mayanil, C. S. K., George, D., Freilich, L., Miljan, E. J., Mania-Farnell, B. J., McLone, D. G., and Bremer, E. G. (2001) J. Biol. Chem. 276, 49299-49309). Here we report that Pax3 directly regulates TGFbeta2 transcription by binding to cis-regulatory elements within its promoter. Chromatin immunoprecipitation revealed that Pax3 bound to the cis-regulatory elements on the TGFbeta2 promoter (GenBanktrade mark accession number AF118263). Both TGFbeta2 promoter-luciferase activity measurements in transient cotransfection experiments and electromobility shift assays supported the idea that Pax3 regulates TGFbeta2 by directly binding to its cis-regulatory regions. Additionally, by using a combination of co-immunoprecipitation and chromatin immunoprecipitation, we show that the TGFbeta2 cis-regulatory elements between bp 741-940 and bp 1012-1212 bind acetylated Pax3 and are associated with p300/CBP and histone deacetylases. The cis-regulatory elements between bp 741 and 940 in addition to associating with acetylated Pax3 and HDAC1 also associated with SIRT1. Whole mount in situ hybridization and quantitative real time reverse transcription-PCR showed diminished levels of TGFbeta2 transcripts in Pax3(-/-) mouse embryos (whose phenotype is characterized by neural tube defects) as compared with Pax3(+/+) littermates (embryonic day 10.0; 30 somite stage), suggesting that Pax3 regulation of TGFbeta2 may play a pivotal role during early embryonic development.

Animals↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

An approach to the characterization of silica exposure in U.S. industry.

Quantitative evaluation of worker exposure to silica in nine Standard Industrial Classification (SIC) codes was conducted, using data derived from OSHA compliance inspections, in order to assess the silica exposure problem in the U.S. The nine SICs studied were those in which OSHA inspections were concentrated. They include: construction; chemical manufacture; stone, glass, and clay manufacturing; primary metal industries; metal fabrication; machinery; transportation; and miscellaneous manufacturing industries. High exposures to silica were documented in each industry, with the number of test samples over the permissible exposure limit ranging from 14% (aluminum foundries) to 73% (pottery). An estimation is made that 24,889 workers employed in ferrous and nonferrous foundries are at risk of silica-related pulmonary effects. The data developed in this analysis also indicate the need to investigate certain industries that had high exposures but few inspections. The limitations of the data base for estimating the scope of the silica problem, including lack of data on mining and milling, are discussed. We conclude that exposure to silica represents a continuing and significant problem in a number of U.S. industries.

Air Pollutants, Occupational↗

Nonsynonymous SNPs: validation characteristics, derived allele frequency patterns, and suggestive evidence for natural selection.

We experimentally investigated more than 1,200 entries in dbSNP that would change amino-acids (nsSNPs), using various subsets of DNA samples drawn from 18 global populations (approximately 1,000 subjects in total). First, we mined the data for any SNP features that correlated with a high validation rate. Useful predictors of valid SNPs included multiple submissions to dbSNP, having a dbSNP validation statement, and being present in a low number of ESTs. Together, these features improved validation rates by almost 10-fold. Higher-abundance SNPs (e.g., T/C variants) also validated more frequently. Second, we considered derived alleles and noted a considerably (approximately 10%) increased average derived allele frequency (DAF) in Europeans vs. Africans, plus a further increase in some other populations. This was not primarily due to an SNP ascertainment bias, nor to the effects of natural selection. Instead, it can be explained as a drift-based, progressive increase in DAF that occurs over many generations and becomes exaggerated during population bottlenecks. This observation could be used as the basis for novel DAF-based tests for comparing demographic histories. Finally, we considered individual marker patterns and identified 37 SNPs with allele frequency variance or FST values consistent with the effects of population-specific natural selection. Four particularly striking clusters of these markers were apparent, and three of these coincide with genes/regions from among only several dozen such domains previously suggested by others to carry signatures of selection.

Alleles↗

Diversity and selection in sorghum: simultaneous analyses using simple sequence repeats.

Although molecular markers and DNA sequence data are now available for many crop species, our ability to identify genetic variation associated with functional or adaptive diversity is still limited. In this study, our aim was to quantify and characterize diversity in a panel of cultivated and wild sorghums (Sorghum bicolor), establish genetic relationships, and, simultaneously, identify selection signals that might be associated with sorghum domestication. We assayed 98 simple sequence repeat (SSR) loci distributed throughout the genome in a panel of 104 accessions comprising 73 landraces (i.e., cultivated lines) and 31 wild sorghums. Evaluation of SSR polymorphisms indicated that landraces retained 86% of the diversity observed in the wild sorghums. The landraces and wilds were moderately differentiated (F st=0.13), but there was little evidence of population differentiation among racial groups of cultivated sorghums (F st=0.06). Neighbor-joining analysis showed that wild sorghums generally formed a distinct group, and about half the landraces tended to cluster by race. Overall, bootstrap support was low, indicating a history of gene flow among the various cultivated types or recent common ancestry. Statistical methods (Ewens-Watterson test for allele excess, lnRH, and F st) for identifying genomic regions with patterns of variation consistent with selection gave significant results for 11 loci (approx. 15% of the SSRs used in the final analysis). Interestingly, seven of these loci mapped in or near genomic regions associated with domestication-related QTLs (i.e., shattering, seed weight, and rhizomatousness). We anticipate that such population genetics-based statistical approaches will be useful for re-evaluating extant SSR data for mining interesting genomic regions from germplasm collections.

Cluster Analysis↗

Survey of long terminal repeat retrotransposons of domesticated silkworm (Bombyx mori).

Long terminal retrotransposons are major components of eukaryotic transposable elements. We have surveyed the long terminal repeats (LTR) retrotransposons of domesticated silkworm (Bombyx mori) by mining the data produced by Bombyx mori Genome Sequencing Project. At least 29 separate families of LTR retrotransposons are identified in this survey, comprising of 11.8% of the complete sequence. Families of domesticated silkworm LTR retrotransposons can be mainly classified into three groups: gypsy-like, copia-like, Pao-Bel. Fourteen families identified consist of gypsy-like elements, four families consist of copia-like elements and seven families consist of Pao-Bel elements. In addition to the three groups of LTR retrotransposons, two families of unusual non-coding elements are identified in the genome of this species. Further phylogenetic analysis of RT domain indicates that the elements of B.mori show high diversity and can form different clades in each group. An analysis of sequence variation from different families reveals distinct patterns of variation for the elements belonging to three groups. The analysis of the domesticated silkworm LTR retrotransposons should assist in our understanding of the roles of retroelement in lepidopteron insect genome evolution.

Animals↗

Using persuasive messages to encourage voluntary hearing protection among coal miners.

INTRODUCTION: This longitudinal field study was designed to encourage Appalachian coal miners in West Virginia and Pennsylvania to engage in hearing-protection behaviors. METHOD: Participants were mailed postcards that featured either a positive, negative, or neutral message on the outside of the postcard and a message encouraging hearing protection behaviors on the inside. The first posttest measurement of the effectiveness of the persuasive messages was conducted about a week after the postcards were mailed. The delayed posttest measurement was conducted six weeks later. RESULTS: Responses from 307 coal miners revealed that the positive or neutral messages generated significantly more self-reported hearing protection behaviors than the negative message. Identical results were obtained in a delayed posttest assessment of miners' self-reported hearing protection behaviors. The positive message was also more effective than either the neutral or negative message in preventing defensive mechanisms from emerging over time. IMPACT ON INDUSTRY: Positive and neutral messages were convincingly more successful than negative messages in facilitating self-reported hearing protection behaviors among coal miners. Similarly, the positive messages kept defensive processes at bay.

Adult↗

Screening the receptorome to discover the molecular targets for plant-derived psychoactive compounds: a novel approach for CNS drug discovery.

Because psychoactive plants exert profound effects on human perception, emotion, and cognition, discovering the molecular mechanisms responsible for psychoactive plant actions will likely yield insights into the molecular underpinnings of human consciousness. Additionally, it is likely that elucidation of the molecular targets responsible for psychoactive drug actions will yield validated targets for CNS drug discovery. This review article focuses on an unbiased, discovery-based approach aimed at uncovering the molecular targets responsible for psychoactive drug actions wherein the main active ingredients of psychoactive plants are screened at the "receptorome" (that portion of the proteome encoding receptors). An overview of the receptorome is given and various in silico, public-domain resources are described. Newly developed tools for the in silico mining of data derived from the National Institute of Mental Health Psychoactive Drug Screening Program's (NIMH-PDSP) K(i) Database (K(i) DB) are described in detail. Additionally, three case studies aimed at discovering the molecular targets responsible for Hypericum perforatum, Salvia divinorum, and Ephedra sinica actions are presented. Finally, recommendations are made for future studies.

Animals↗

Modeling toxicity by using supervised kohonen neural networks.

Counterprogation neural network is shown to be a powerful and suitable tool for the investigation of toxicity. This study mined a data set of 568 chemicals. Two hundred eighty-two objects were used as the training set and 286 as the test set. The final model developed presents high performances on the data set R(2) = 0.83 (R(2) = 0.97 on the training set, R(2) = 0.59 on the test set). This technique distinguishes itself also for the ability to give to the expert two-dimensional maps suitable for the study of the distribution/clustering of the data and the identification of outliers.

Journal Article↗

Mining high-throughput screening data of combinatorial libraries: development of a filter to distinguish hits from nonhits.

Kohonen neural networks generate projections of large data sets defined in high-dimensional space. The resulting self-organizing maps can be used in many applications in the drug discovery process, such as to analyze combinatorial libraries for their similarity or diversity and to select descriptors for structure-activity relationships. The ability to investigate thousands of compounds in parallel also allows one to conduct a study based on single-dose experiments of high-throughput screening campaigns, which are known to have a greater uncertainty than IC50 or Ki values. This is demonstrated here for a data set of 5513 compounds from one combinatorial library. Furthermore, a method was developed that uses self-organizing maps not only as an indicator of structure-activity relationships, but as the basis of a classification system allowing predictive modeling of combinatorial libraries.

Journal Article↗

Expression profiling with oligonucleotide arrays: technologies and applications for neurobiology.

DNA microarrays have been used in applications ranging from the assignment of gene function to analytical uses in prognostics. However, the detection sensitivity, cross hybridization, and reproducibility of these arrays can affect experimental design and data interpretation. Moreover, several technologies are available for fabrication of oligonucleotide microarrays. We review these technologies and performance attributes and, with data sets generated from human brain RNA, present statistical tools and methods to analyze data quality and to mine and visualize the data. Our data show high reproducibility and should allow an investigator to discern biological and regional variability from differential expression. Although we have used brain RNA as a model system to illustrate some of these points, the oligonucleotide arrays and methods employed in this study can be used with cell lines, tissue sections, blood, and other fluids. To further demonstrate this point, we provide data generated from total RNA sample sizes of 200 ng.

Brain↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Microbial genomes.

Microbial genome sequencing is driven by the need to understand and control pathogens and to exploit extremophiles and their enzymes in bioremediation and industry. It is hard for the traditional bacteriologist to grasp the scale and pace of the venture. Around two dozen microbial genomes have now been completed and, within a decade, genomes from every significant species of bacterial pathogen of humans, animals and plants will have been sequenced. Indeed, we will often have more than one sequence from a species or genus--for example, we already have sequences from two strains of Helicobacter pylori, from two strains of Mycobacterium tuberculosis and from three species of Pyrococcus. However, genome sequencing risks becoming expensive molecular stamp-collecting without the tools to mine the data and fuel hypothesis-driven laboratory-based research. Bioinformatics, twinned with the new experimental approaches forming functional genomics', provides some of the needed tools. Nonetheless, there will be an increasing need for us to explore the detailed implications of genomic findings. Microbial genome sequencing thus represents not a threat, but an exciting opportunity for molecular microbiologists.

Computational Biology↗

Novel OCRL1 mutations in patients with the phenotype of Dent disease.

BACKGROUND: Dent disease is an X-linked tubulopathy frequently caused by mutations affecting the voltage-gated chloride channel and chloride/proton antiporter ClC-5. A recent study showed that defects in OCRL1, encoding a phosphatidylinositol 4,5-bisphosphate 5-phosphatase (Ocrl) and usually found mutated in patients with Lowe syndrome, also can provoke a Dent-like phenotype (Dent 2 disease). METHODS: We investigated 20 CLCN5-negative males from 17 families with a phenotype resembling Dent disease for defects in OCRL1. RESULTS: In our complete series of 35 families with a phenotype of Dent disease, a mutation in the OCRL1 gene was detected in 6 kindreds. All were novel frameshift (Q70RfsX88 and T121NfsX122, detected twice) or missense mutations (I257T and R476W). None of our patients had cognitive or behavioral impairment or cataracts, 2 classic hallmarks of Lowe syndrome. All patients had mild increases in lactate dehydrogenase and/or creatine kinase levels, which rarely is observed in CLCN5-positive patients, but frequently found in patients with Lowe syndrome. To explain the phenotypic heterogeneity caused by OCRL1 mutations, we performed extensive data-bank mining and extended reverse-transcriptase polymerase chain reaction analysis, which provided no evidence for yet unknown (tissue-specific) alternative OCRL1 transcripts. CONCLUSION: Mutations in the OCRL1 gene are found in approximately 23% of kindreds with a Dent phenotype. Defective protein sorting/targeting of Ocrl might be the reason for mildly elevated creatine kinase and lactate dehydrogenase serum concentrations in these patients and a clue to suspect Dent disease unrelated to CLCN5 mutations. It remains to be elucidated why the various OCRL1 mutations found in patients with Dent 2 disease do not cause cataracts.

Chloride Channels↗

Large-scale analysis of the human and mouse transcriptomes.

High-throughput gene expression profiling has become an important tool for investigating transcriptional activity in a variety of biological samples. To date, the vast majority of these experiments have focused on specific biological processes and perturbations. Here, we have generated and analyzed gene expression from a set of samples spanning a broad range of biological conditions. Specifically, we profiled gene expression from 91 human and mouse samples across a diverse array of tissues, organs, and cell lines. Because these samples predominantly come from the normal physiological state in the human and mouse, this dataset represents a preliminary, but substantial, description of the normal mammalian transcriptome. We have used this dataset to illustrate methods of mining these data, and to reveal insights into molecular and physiological gene function, mechanisms of transcriptional regulation, disease etiology, and comparative genomics. Finally, to allow the scientific community to use this resource, we have built a free and publicly accessible website (http://expression.gnf.org) that integrates data visualization and curation of current gene annotations.

Animals↗

Slug is a novel downstream target of MyoD. Temporal profiling in muscle regeneration.

Temporal expression profiling was utilized to define transcriptional regulatory pathways in vivo in a mouse muscle regeneration model. Potential downstream targets of MyoD were identified by temporal expression, promoter data base mining, and gel shift assays; Slug and calpain 6 were identified as novel MyoD targets. Slug, a member of the snail/slug family of zinc finger transcriptional repressors critical for mesoderm/ectoderm development, was further shown to be a downstream target by using promoter/reporter constructs and demonstration of defective muscle regeneration in Slug null mice.

Animals↗

Prediction of penicillin resistance in Staphylococcus aureus isolates from dairy cows with mastitis, based on prior test results.

AIM: To gauge how well prior laboratory test results predict in vitro penicillin resistance of Staphylococcus aureus isolates from dairy cows with mastitis. METHODS: Population-based data on the farm of origin (n=79), genotype based on pulsed-field gel electrophoresis (PFGE) results, and the penicillin-resistance status of Staph. aureus isolates (n=115) from milk samples collected from dairy cows with mastitis submitted to two diagnostic laboratories over a 6-month period were used. Data were mined stochastically using the all-possible-pairs method, binomial modelling and bootstrap simulation, to test whether prior test results enhance the accuracy of prediction of penicillin resistance on farms. RESULTS: Of all Staph. aureus isolates tested, 38% were penicillin resistant. A significant aggregation of penicillin-resistance status was evident within farms. The probability of random pairs of isolates from the same farm having the same penicillin-resistance status was 76%, compared with 53% for random pairings of samples across all farms. Thus, the resistance status of randomly selected isolates was 1.43 times more likely to correctly predict the status of other isolates from the same farm than the random population pairwise concordance probability (p=0.011). This effect was likely due to the clonal relationship of isolates within farms, as the predictive fraction attributable to prior test results was close to nil when the effect of within-farm clonal infections was withdrawn from the model. CONCLUSIONS: Knowledge of the penicillin-resistance status of a prior Staph. aureus isolate significantly enhanced the predictive capability of other isolates from the same farm. In the time and space frame of this study, clinicians using previous information from a farm would have more accurately predicted the penicillin-resistance status of an isolate than they would by chance alone on farms infected with clonal Staph. aureus isolates, but not on farms infected with highly genetically heterogeneous bacterial strains.

Animals↗