Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Analysis of remnant reticulocyte mRNA reveals new genes and antisense transcripts expressed in the human erythroid lineage.

BACKGROUND AND OBJECTIVES: We studied the gene expression profile of human purified reticulocytes to provide a transcriptional basis for the study of erythroid biology, differentiation and hematologic disorders. DESIGN AND METHODS: We screened highly purified blood reticulocytes from ten healthy adult volunteers. We chose a modified protocol of serial analysis of gene expression (SAGE), the serial analysis of downsized extracts (SADE). RESULTS: Data analysis revealed that 64% of gene signatures (tags) matched with known genes; mainly hemoglobin. In addition to the abundant globin mRNA, SAGE analysis identified previously described genes and new transcripts. In reticulocytes, which are poor in mRNA, we also identified 9% of EST and 27% of tags that did not match with any known genes. Mining our data, 70% of the unknown tags and 39% of tags identifying EST were found to be specific to the reticulocyte. We demonstrated the presence of a mRNA that matched with the reverse sequence of the hemoglobin b (HBB) transcript. INTERPRETATION AND CONCLUSIONS: This is the first description of an antisense transcript of the human HBB gene suggesting regulation by way of sense-antisense pairing. The well-characterized genes found in the SAGE library were genes specific to the blood cell lineage, housekeeping genes and, interestingly, genes not previously described in the reticulocyte. Furthermore the study provides markers of the erythroid lineage regulated during the differentiation process as observed in in vitro experiments.

Adult↗

Modtector: ultra-fast modification signal mining on mapped sequencing reads.

SUMMARY: Existing tools for RNA epitranscriptomic modification and structural signal analysis are often fragmented, inefficiency, and limited to single signal types. We developed Modtector, an unified tool for extracting mutation and reverse-transcription stop signals from aligned sequencing reads. By using a "count-then-correct" strategy, Modtector reduces computational complexity and enables efficient dual-signal analysis. It achieves multi-fold speedups on large-genome and high-coverage datasets, including completing HEK293 22G data analysis in 5 minutes, and show strong scalability on single-cell datasets with speedups exceeding 50-fold. AVAILABILITY: The source code is available at GitHub (https://github.com/TongZhou2017/modtector) and Crates.io (https://crates.io/crates/modtector). The archived source-code snapshot used in this study is available at Zenodo (DOI: 10.5281/zenodo.20967747), corresponding to GitHub commit 7c60e9d. Workflow examples, datasets, and analysis scripts are available at Zenodo (DOI: 10.5281/zenodo.17316476 and 10.5281/zenodo.18523297).

Humans↗

ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation.

MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).

Semantics↗

Preparing for the next influenza pandemic: lessons from multinational data.

BACKGROUND: In the past decade, avian influenza has made several incursions of increasing scope and virulence into humans. The likelihood of another pandemic is increasing with time. In work recently published, influenza was found to be the principal cause of the increase in mortality in the United States during the winter months. In a companion report, the U.S. national vaccination program was shown to have increased coverage of high risk groups 5-fold from 1980 to 1999, but excess mortality did not decline in any elderly age group. The Multinational Influenza Seasonal Mortality Study has assembled and has begun to mine mortality data from many countries. Early results indicate that the U.S. results extend to other economically developed countries and probably worldwide. RESULTS: The Multinational Influenza Seasonal Mortality Study data extend the observations of others that there were heralding events that provided advance warning for all of the pandemics of the 20th century. Moreover, in the first year of emergence of A(H3N2) viruses, the 1968-1969 pandemic produced little excess mortality outside of North America. It appears that there were at least 2 variants of the pandemic virus, differing at 1 or more internal gene loci, and that the more virulent form emerged as dominant in the second pandemic season. CONCLUSIONS: Integrating these findings, it seems clear that the influenza control strategy now used in about 50 countries is less than optimal. While it is likely that there will be more time to react in the pandemic season than previously imagined, an enhancement of the historical strategy is clearly indicated. Furthermore, the vaccine shortage that is presently inevitable suggests that a departure from the historical strategy if calamitous ineffectiveness is to be avoided.

Aged↗

Cloning of ovocalyxin-36, a novel chicken eggshell protein related to lipopolysaccharide-binding proteins, bactericidal permeability-increasing proteins, and plunc family proteins.

The avian eggshell is a composite biomaterial composed of noncalcifying eggshell membranes and the overlying calcified shell matrix. The shell is deposited in a uterine fluid where the concentration of different protein species varies at different stages of its formation. The role of avian eggshell proteins during shell formation remains poorly understood, and we have sought to identify and characterize the individual components in order to gain insight into their function during elaboration of the eggshell. In this study, we have used direct sequencing, immunochemistry, expression screening, and EST data base mining to clone and characterize a 1995-bp full-length cDNA sequence corresponding to a novel chicken eggshell protein that we have named Ovocalyxin-36 (OCX-36). Ovocalyxin-36 protein was only detected in the regions of the oviduct where egg-shell formation takes place; uterine OCX-36 message was strongly up-regulated during eggshell calcification. OCX-36 localized to the calcified eggshell predominantly in the inner part of the shell, and to the shell membranes. BlastN data base searching indicates that there is no mammalian version of OCX-36; however, the protein sequence is 20-25% homologous to proteins associated with the innate immune response as follows: lipopolysaccharide-binding proteins, bactericidal permeability-increasing proteins, and Plunc family proteins. Moreover, the genomic organization of these proteins and OCX-36 appears to be highly conserved. These observations suggest that OCX-36 is a novel and specific chicken eggshell protein related to the superfamily of lipopolysaccharide-binding proteins/bactericidal permeability-increasing proteins and Plunc proteins. OCX-36 may therefore participate in natural defense mechanisms that keep the egg free of pathogens.

Acute-Phase Proteins↗

Lung cancer in the Schneeberg mines: a reappraisal of the data reported by Harting and Hesse in 1879.

The first description of occupational lung cancer, by Harting and Hesse in 1879, unfortunately is not readily accessible. Its account of the vicissitudes of the Schneeberg miners merits study and is therefore presented in summary and set in a historical and geological context. The authors attempted to discover the cause of the disease and made recommendations for improving the health of miners. In the course of their programme of investigations, they developed methods for measuring airborne dust and inhaled dust by personal monitoring. It was left to subsequent discovery for radon and its daughter products to be identified as the causal agents. Later generations were to discover the impact of radioactive spoils from mines situated in the mountain range in which Schneeberg was located.

Germany↗

Transport injuries in small coal mines: an exploratory analysis.

Mine Safety and Health Administration (MSHA) surveillance data were analyzed to elucidate mine characteristics or injury characteristics that distinguished mines with high rates of transport-related injuries from mines with lower transport injury rates. The results showed that most high-rate mines are small, high-rate mines have a disproportionate number of injuries involving young and less experienced workers, and injuries in high-rate mines are proportionally more severe. Further analyses of the MSHA injury data showed that smaller mines have a greater share of fatal and permanently disabling injuries, whereas larger mines have a greater share of injuries involving no lost time. Based on these results, we explored two explanations for the small mine injury risk: (1) a suggestion that differences in injury reporting between large and small mines may contribute to an apparent small mine injury risk, and (2) identification of factors contributing to a true difference in transport-related injury risk between small and large mines. Whereas it was true that most high injury rate mines were small, most small mines were actually zero-rate, having reported employment but no injuries to MSHA. An analysis employing binomial probability theory showed that a substantial proportion of small mines reported zero injuries when it was statistically probable that injuries would have occurred. This indicated that small mines may underreport injuries relative to larger mines. The possibility that reporting bias affected the associations found in this study was explored by eliminating the least severe injuries from the data set and evaluating changes in associations. This "adjustment" for reporting bias did not change previously observed relationships. Finally, MSHA injury data were analyzed in concert with mining population data collected by the Bureau of Mines. With such denominator information, the results indicated a disproportionately high risk of injury among workers in their first year at a mine and indicated that higher injury risk in small mines might be explained by the fact that workers at small mines have substantially less experience than workers at large mines. An effect of age was not found in these analyses. These results suggest the potential importance of targeted training programs for newly hired miners. Results also point to the need to explore specific factors contributing to the small mine injury risk, and to the necessity for complete and accurate reporting of injury data.

Accidents, Occupational↗

Text-based knowledge discovery: search and mining of life-sciences documents.

Text literature is playing an increasingly important role in biomedical discovery. The challenge is to manage the increasing volume, complexity and specialization of knowledge expressed in this literature. Although information retrieval or text searching is useful, it is not sufficient to find specific facts and relations. Information extraction methods are evolving to extract automatically specific, fine-grained terms corresponding to the names of entities referred to in the text, and the relationships that connect these terms. Information extraction is, in turn, a means to an end, and knowledge discovery methods are evolving for the discovery of still more-complex structures and connections among facts. These methods provide an interpretive context for understanding the meaning of biological data.

Biological Science Disciplines↗

An atlas of forecasted molecular data. 1. Internuclear separations of main-group and transition-metal neutral gas-phase diatomic molecules in the ground state.

Needed spectroscopic data on diatomic molecules can often be found in the superb critical tables of Huber and Herzberg or in the literature published since 1979. Unfortunately, these sources apply to only a fraction of the diatomic species that can exist and so investigators have had to rely on interpolation, additivity, or ad hoc rules to estimate needed values, all of which require other information that is often lacking. This Atlas presents 1001 additional internuclear separations for use until critical tables are available to fill the needs more precisely. The Atlas was produced by mining the data from Huber and Herzberg for trends with least-squares analysis and with neural network software. There are 162 molecules about whose data Huber and Herzberg had no qualifications and whose data were employed for this work; 248 copies of data with low and high magnitudes were added to reduce the effects of frequency. Internuclear separations for 1001 species not found in Huber and Herzberg are presented, and least-squares predictions supplement some of them. The results, i.e., the Atlas, are presented as Table A, Supporting Information. The average error, based on the average of the absolute differences between the predicted values and tabulated values for the molecules having Huber and Herzberg data, is 0.074 A; if each error is expressed as a percent of the forecast to which it pertains, the average of these errors is 2.94%. There are 25 "questionable" data from Huber and Herzberg, not used in the preparation of the Atlas, for which predictions are included in the Atlas. Of these, 14 agree with the predicted internuclear separations to within twice the stated errors. Additional atlases for other properties of diatomic molecules are in preparation.

Journal Article↗

tbCPSF30 depletion by RNA interference disrupts polycistronic RNA processing in Trypanosoma brucei.

Gene expression in eukaryotes requires the post-transcriptional cleavage of mRNA precursors into mature mRNAs. In Trypanosoma brucei, mRNA processing is of particular importance, since most transcripts are derived from polycistronic transcription units. This organization dictates that regulated gene expression is promoter-independent and governed at the posttranscriptional level. We have identified tbCPSF30, a protein containing five CCCH zinc finger motifs, which is a homologue of the cleavage and polyadenylation specificity factor (CPSF) 30-kDa subunit, a component of the machinery required for 3'-end formation in yeast and mammals. Using gene silencing of tbCPSF30 by RNA interference, we demonstrate that this gene is essential in bloodstream and procyclic forms of T. brucei. Interestingly, tbCPSF30-specific RNA interference results in the accumulation of an aberrant tbCPSF30 mRNA species concomitant with depletion of tbCPSF30 protein. tbCPSF30 protein depletion is accompanied by the accumulation of unprocessed tubulin RNAs, implicating tbCPSF30 in polycistronic RNA processing. By genome data base mining, we also identify several other putative components of the T. brucei cleavage and polyadenylation machinery, indicating their conservation throughout eukaryotic evolution. This study is the first to identify and characterize a core component of the T. brucei CPSF and show its involvement in polycistronic RNA processing.

Amino Acid Sequence↗

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans↗

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software↗

Integrating gene and protein expression data: pattern analysis and profile mining.

Proteomics and functional genomics are emerging new research fields devoted to the study of the entire collection of proteins and mRNA transcripts (collectively known as gene products) that define a biological system. DNA microarrays are now a popular platform for measuring changes in messenger RNA transcript levels on a genome-wide scale, while gel-free shotgun profiling methods based on tandem mass spectrometry are increasingly being used to determine the identity, modification states, and relative abundance of large numbers of proteins. By defining the behavior of entire biological pathways and networks under various physiological states, these studies aim to extend traditional reductionist molecular genetic approaches regarding the biological roles of the vast array of uncharacterized gene products. A key goal is to determine how the information encoded by the myriad of expressed gene products is integrated at the molecular, cellular, and even whole organism level to create the dynamic biochemical processes and complex physiological controls that sustain life. While comparison of the complementary information contained in proteomic and mRNA data sets poses considerable analytical challenges, these efforts should provide added insight into the fundamental mechanisms underlying physiology, development, and the emergence of disease. Here, we outline several analytical approaches, methods, and tools that have proven to be helpful in the face of this important challenge.

Algorithms↗

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Acceptance of rules generated by machine learning among medical experts.

OBJECTIVES: The aim was to evaluate the potential for monotonicity constraints to bias machine learning systems to learn rules that were both accurate and meaningful. METHODS: Two data sets, taken from problems as diverse as screening for dementia and assessing the risk of mental retardation, were collected and a rule learning system, with and without monotonicity constraints, was run on each. The rules were shown to experts, who were asked how willing they would be to use such rules in practice. The accuracy of the rules was also evaluated. RESULTS: Rules learned with monotonicity constraints were at least as accurate as rules learned without such constraints. Experts were, on average, more willing to use the rules learned with the monotonicity constraints. CONCLUSIONS: The analysis of medical databases has the potential of improving patient outcomes and/or lowering the cost of health care delivery. Various techniques, from statistics, pattern recognition, machine learning, and neural networks, have been proposed to "mine" this data by uncovering patterns that may be used to guide decision making. This study suggests cognitive factors make learned models coherent and, therefore, credible to experts. One factor that influences the acceptance of learned models is consistency with existing medical knowledge.

Alzheimer Disease↗

Peptidomics of the larval Drosophila melanogaster central nervous system.

Neuropeptides regulate most, if not all, biological processes in the animal kingdom, but only seven have been isolated and sequenced from Drosophila melanogaster. In analogy with the proteomics technology, where all proteins expressed in a cell or tissue are analyzed, the peptidomics approach aims at the simultaneous identification of the whole peptidome of a cell or tissue, i.e. all expressed peptides with their posttranslational modifications. Using nanoscale liquid chromatography combined with tandem mass spectrometry and data base mining, we analyzed the peptidome of the larval Drosophila central nervous system at the amino acid sequence level. We were able to provide biochemical evidence for the presence of 28 neuropeptides using an extract of only 50 larval Drosophila central nervous systems. Eighteen of these peptides are encoded in previously cloned or annotated precursor genes, although not all of them were predicted correctly. Eleven of these peptides were never purified before. Eight other peptides are entirely novel and are encoded in five different, not yet annotated genes. This neuropeptide expression profiling study also opens perspectives for other eukaryotic model systems, for which genome projects are completed or in progress.

Amino Acid Sequence↗

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans↗

Identification of a novel human kinase supporter of Ras (hKSR-2) that functions as a negative regulator of Cot (Tpl2) signaling.

Kinase suppressor of Ras (KSR) is an integral and conserved component of the Ras signaling pathway. Although KSR is a positive regulator of the Ras/mitogen-activated protein (MAP) kinase pathway, the role of KSR in Cot-mediated MAPK activation has not been identified. The serine/threonine kinase Cot (also known as Tpl2) is a member of the MAP kinase kinase kinase (MAP3K) family that is known to regulate oncogenic and inflammatory pathways; however, the mechanism(s) of its regulation are not precisely known. In this report, we identify an 830-amino acid novel human KSR, designated hKSR-2, using predictions from genomic data base mining based on the structural profile of the KSR kinase domain. We show that, similar to the known human KSR, hKSR-2 co-immunoprecipitates with many signaling components of the Ras/MAPK pathway, including Ras, Raf, MEK-1, and ERK-1/2. In addition, we demonstrate that hKSR-2 co-immunoprecipitates with Cot and that co-expression of hKSR-2 with Cot significantly reduces Cot-mediated MAPK and NF-kappaB activation. This inhibition is specific to Cot, because Ras-induced ERK and IkappaB kinase-induced NF-kappaB activation are not significantly affected by hKSR-2 co-expression. Moreover, Cot-induced interleukin-8 production in HeLa cells is almost completely inhibited by the concurrent expression of hKSR-2, whereas transforming growth factor beta-activated kinase 1 (TAK1)/TAK1-binding protein 1 (TAB1)-induced interleukin-8 production is not affected by hKSR-2 co-expression. Taken together, these results indicate that hKSR-2, a new member of the KSR family, negatively regulates Cot-mediated MAP kinase and NF-kappaB pathway signaling.

Base Sequence↗