Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

An omnidirectional M-mode echocardiography system and its clinical application.

This paper introduces the omnidirectional M-mode echocardiography (OME), which can detect dynamic information from sequential echocardiography images. The method for detecting dynamic information is based on the rebuilding of their "gray (position)-time" function [Qiang L, Wenjing J, Li Z. A method for detecting dynamic information of sequential images--omnidirectional gray-time waveform and its applications in echocardiography images. In: Proceedings of CISST' 2001. 2001. p. 760-3; Qiang L, Wenjing J, Xiuzhi Y. A method for mining data of sequetial images-rebuilding of gray (position)-time function on arbitrary direction lines. In: Proceeding of CISST' 2002, vol. 3-6. 2002; Qiang L, Li Z, Wenjing J. The realization of omnidirectional gray-time waveform system and its application on echocardiography. J Electron Meas Instrum, 2002;16(1):70-5] on direction lines. The system can obtain motion and inherent dynamic information of a certain part of the cardiac structure at a certain moment. The system also shows a group of omnidirectional M-mode echocardiography images with synchronous ECG, which is rebuilt from 2D echocardiography images. The ECG supplies a standard time for the omnidirectional M-mode echocardiography images. The system has been applied in clinical application for 3 years and the results are good.

Algorithms↗

Record systems for beef stocker production enterprises.

The accelerated growth of individual animal identification systems is likely to generate significant amounts of data that need to be synchronized, filtered, analyzed, managed, and acted on in real time by data-mining software and animal health professionals who possess a dual understanding of beef systems production and technology associated with management information and record-keeping systems. Ultimately, the resulting information can be used seamlessly throughout a vertically coordinated production system to conduct management and animal health compliance audits, initiate timely animal and product recall measures, and reveal complex biologic and economic relations.

Animal Husbandry↗

A novel molecular pathway of lipid accumulation in human hepatocytes caused by PFOA and PFOS.

Exposed to ubiquitously perfluorooctanoic acid (PFOA) and perfluorooctane sulfonate (PFOS) has been associated with non-alcoholic fatty liver disease (NAFLD), yet the underlying molecular mechanism remains elusive. The extrapolation of empirical studies correlating per- and polyfluoroalkyl substance (PFAS) exposure with NAFLD occurrence to real-life exposure was hindered by the limited availability of mechanistic data at environmentally relevant concentrations. Herein, a novel pathway mediating hepatocyte lipid accumulation by PFOA and PFOS at human-relevant dose (<10&#xa0;&#x3bc;M) was identified by integrating CRISPR-Cas9 genome screening, concentration-dependent transcriptional assay in HepG2 cell and epidemiological data mining. 1) At genetic level, nudt7 showed the highest enriched potency among 569 NAFLD-related genes, and the transcription of nudt7 was significantly downregulated by PFOA and PFOS exposure (<7 &#x3bc;M). 2) At molecular pathway, upon exposure to&#xa0;&#x2264;10-4&#xa0;&#x3bc;M PFOA and PFOS, the downregulation of nudt7 transcriptional expression triggered the reduction of Ace-CoA hydrolase activity. 3) At cellular level, increased lipids were measured in HepG2 cells with PFOA and PFOS (<2&#xa0;&#x3bc;M). Overall, we identified a novel mechanism mediated by transcriptional downregulation of nudt7 gene in hepatocellular lipid increase treated with PFOA and PFOS, which could potentially explain the NAFLD occurrence associated with exposure to PFASs in humans.

Humans↗

Molecular profiling of temporal lobe epilepsy: comparison of data from human tissue samples and animal models.

The advent of gene chip technology and the era of functional genomics have initially been accompanied by huge anticipations to quickly unravel the molecular pathogenesis of multifactorial diseases. Expectations have, today, given way to some concerns about this non-hypothesis driven approach. However, the careful and controlled application of expression microarrays in concert with refined bioinformatic tools may provide novel insights in major disorders particularly of highly complex organs such as the central nervous system (CNS). Epilepsies are among the most frequent CNS disorders affecting approximately 1.5% of the population worldwide. In temporal lobe epilepsy (TLE), the seizure origin typically involves the hippocampal formation, a structure located in the mesial temporal lobe. Many TLE patients develop pharmacoresistance, i.e. seizures can no more be controlled by antiepileptic drugs. In order to achieve seizure control, surgical removal of the epileptogenic focus has been established as successful therapeutic strategy. Hippocampal biopsy tissue of pharmacoresistant TLE patients represents an excellent substrate to analyze molecular mechanisms related to structural and cellular reorganization in epilepsy. The complexity of alterations in TLE hippocampi suggests numerous genes and signaling cascades to be involved in the pathogenesis. By microarrays, genome wide expression profiles can be constituted from TLE tissues. However, hippocampi of pharmacoresistant TLE patients represent an advanced stage of the disease. Early stages of epilepsy development are not available for functional genome analysis in humans. Animal models of TLE appear particularly helpful to study molecular mechanisms of highly dynamic processes such as the development of hyperexcitability and pharmacoresistance. In this review, we summarize recent data of gene expression profiles in human and experimental TLE and discuss the relevance of novel tools for bioinformatic analysis and data mining.

Animals↗

The A-loop, a novel conserved aromatic acid subdomain upstream of the Walker A motif in ABC transporters, is critical for ATP binding.

ATP-binding cassette (ABC) transporters represent one of the largest families of proteins, and transport a variety of substrates ranging from ions to amphipathic anticancer drugs. The functional unit of an ABC transporter is comprised of two transmembrane domains and two cytoplasmic ABC ATPase domains. The energy of the binding and hydrolysis of ATP is used to transport the substrates across membranes. An ABC domain consists of conserved regions, the Walker A and B motifs, the signature (or C) region and the D, H and Q loops. We recently described the A-loop (Aromatic residue interacting with the Adenine ring of ATP), a highly conserved aromatic residue approximately 25 amino acids upstream of the Walker A motif that is essential for ATP-binding. Here, we review the mutational analysis of this subdomain in human P-glycoprotein as well as homology modeling, structural and data mining studies that provide evidence for a functional role of the A-loop in ATP-binding in most members of the superfamily of ABC transporters.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Combining different standards and different approaches for health information retrieval in a quality-controlled gateway.

Internet as source of information is increasing in preeminence in numerous fields, including health. We describe in this paper the CISMeF project (acronym of Catalogue and Index of French-speaking Medical Sites) which has been designed to help the health information consumers and health professionals to find what they are looking for among the numerous health documents available online. The catalogue is founded on two standards: a set of metadata and a terminology based on the MeSH thesaurus which has the same structure and use as an ontology of the medical domain. The structure of the catalogue allows us to place the project at an overlap between the present Web, which is informal, and the forthcoming Semantic Web. Many features of information retrieval and navigation through the catalogue were developed. These features take into account the kind of the end-user (health professional, medical student, patient). The CISMeF-patients catalogue is a sub-catalogue of CISMeF and is dedicated to the patients and the general public. It shares the same model as CISMeF whereas MEDLINE and MedlinePlus do not. We also propose to couple two approaches (morphological processing and data mining) to help the users by correcting and refining their queries.

France↗

Multidimensional protein identification technology (MudPIT): technical overview of a profiling method optimized for the comprehensive proteomic investigation of normal and diseased heart tissue.

An optimized analytical expression profiling strategy based on gel-free multidimensional protein identification technology (MudPIT) is reported for the systematic investigation of biochemical (mal)-adaptations associated with healthy and diseased heart tissue. Enhanced shotgun proteomic detection coverage and improved biological inference is achieved by pre-fractionation of excised mouse cardiac muscle into subcellular components, with each organellar fraction investigated exhaustively using multiple repeat MudPIT analyses. Functional-enrichment, high-confidence identification, and relative quantification of hundreds of organelle- and tissue-specific proteins are achieved readily, including detection of low abundance transcriptional regulators, signaling factors, and proteins linked to cardiac disease. Important technical issues relating to data validation, including minimization of artifacts stemming from biased under-sampling and spurious false discovery, together with suggestions for further fine-tuning of sample preparation, are discussed. A framework for follow-up bioinformatic examination, pattern recognition, and data mining is also presented in the context of a stringent application of MudPIT for probing fundamental aspects of heart muscle physiology as well as the discovery of perturbations associated with heart failure.

Animals↗

Towards the development of a conceptual distance metric for the UMLS.

The objective of this work is to investigate the feasibility of conceptual similarity metrics in the framework of the Unified Medical Language System (UMLS). We have investigated an approach based on the minimum number of parent links between concepts, and evaluated its performance relative to human expert estimates on three sets of concepts for three terminologies within the UMLS (i.e., MeSH, ICD9CM, and SNOMED). The resulting quantitative metric enables computer-based applications that use decision thresholds and approximate matching criteria. The proposed conceptual matching supports problem solving and inferencing (using high-level, generic concepts) based on readily available data (typically represented as low-level, specific concepts). Through the identification of semantically similar concepts, conceptual matching also enables reasoning in the absence of exact, or even approximate, lexical matching. Finally, conceptual matching is relevant for terminology development and maintenance, machine learning research, decision support system development, and data mining research in biomedical informatics and other fields.

Algorithms↗

Mapping high-dimensional data onto a relative distance plane--an exact method for visualizing and characterizing high-dimensional patterns.

We introduce a distance (similarity)-based mapping for the visualization of high-dimensional patterns and their relative relationships. The mapping preserves exactly the original distances between points with respect to any two reference patterns in a special two-dimensional coordinate system, the relative distance plane (RDP). As only a single calculation of a distance matrix is required, this method is computationally efficient, an essential requirement for any exploratory data analysis. The data visualization afforded by this representation permits a rapid assessment of class pattern distributions. In particular, we can determine with a simple statistical test whether both training and validation sets of a 2-class, high-dimensional dataset derive from the same class distributions. We can explore any dataset in detail by identifying the subset of reference pairs whose members belong to different classes, cycling through this subset, and for each pair, mapping the remaining patterns. These multiple viewpoints facilitate the identification and confirmation of outliers. We demonstrate the effectiveness of this method on several complex biomedical datasets. Because of its efficiency, effectiveness, and versatility, one may use the RDP representation as an initial, data mining exploration that precedes classification by some classifier. Once final enhancements to the RDP mapping software are completed, we plan to make it freely available to researchers.

Algorithms↗

A bioinformatics framework for genotype-phenotype correlation in humans with Marfan syndrome caused by FBN1 gene mutations.

Mutations in the human FBN1 gene are known to be associated with the Marfan syndrome, an autosomal dominant inherited multi-systemic connective tissue disorder. However, in the absence of solid genotype-phenotype correlations, the identification of an FBN1 mutation has only little prognostic value. We propose a bioinformatics framework for the mutated FBN1 gene which comprises the collection, management, and analysis of mutation data identified by molecular genetic analysis (DHPLC) and data of the clinical phenotype. To query our database at different levels of information, a relational data model, describing mutational events at the cDNA and protein levels, and the disease's phenotypic expression from two alternative views, was implemented. For database similarity requests, a query model which uses a distance measure based on log-likelihood weights for each clinical manifestation, was introduced. A data mining strategy for discovering diagnostic markers, classification and clustering of phenotypic expressions was provided which enabled us to confirm some known and to identify some new genotype-phenotype correlations.

Computational Biology↗

Spectral analysis and fingerprinting for biomedia characterisation.

Classical culture media, as well as domestic and/or industrial wastewater treated by biological processes, have a complex composition. The on-line and/or in situ determination of some substances is possible, but expensive, as sample collection and pre-treatment are often necessary with strict rules of sterility. More global methods can be used to detect rapidly "accidents" such as the appearance of an undesirable by-product in a fermentation broth or of a toxic substance in wastewater. These methods combine a "hard" part, for sensing, and a "soft" part, for data treatment. Among potential "hard" candidates, spectroscopy can be the basis for non-invasive and non-destructive measuring systems. Some of them have been already tested in situ: ultra-violet-visible, infra-red (mid or near), fluorescence (mono-dimensional, two-dimensional or synchronous), dielectric, while others, more sophisticated, such as mass spectrometry, coupled or not to pyrolysis, nuclear magnetic resonance and Raman spectroscopy, have been proposed. All these methods provide spectra, i.e. large sets of data, from which meaningful information should be rapidly extracted, either for analysis or fingerprinting. The recourse to data-mining techniques (the "soft" part) such as principal components analysis, projection on latent structures or artificial neural networks, is a necessary step for that task. A review of techniques, mostly based on spectroscopy, with examples taken in the bioengineering field in general is proposed.

Biotechnology↗

Enhanced detectability in proteome studies.

The discovery of candidate biomarkers from biological materials coupled with the development of detection methods holds both incredible clinical potential as well as significant challenges. However, the proteomic techniques still provide the low dynamic range of protein detection at lower abundances. This review describes the current development of potential methods to enhance the detection and quantification in proteome studies. It also includes the bioinformatics tools that are helpfully used for data mining of protein ontology. Therefore, we believe that this review provided many proteomic approaches, which would be very potent and useful for proteome studies and for further diagnostic and therapeutic applications.

Biomarkers↗

Elevated Triggering Receptor Expressed on Myeloid Cells 2 Expression in Tumor-Associated Macrophages Suppresses Cytotoxic T Cell Infiltration and Facilitates Immune Escape in Colorectal Cancer.

BACKGROUND & AIMS: Emerging evidence supports a crucial role for tumor-associated macrophages in shaping the immunosuppressive tumor microenvironment. Furthermore, research has identified that the triggering receptor expressed on myeloid cells 2 has immunomodulatory functions. The present investigated the potential effect of triggering receptor expressed on myeloid cells 2 expression in tumor-associated macrophages on facilitating immune evasion in colorectal cancer. METHODS: Immunohistochemical analysis of clinical specimens, complemented by extensive data mining from The Cancer Genome Atlas, revealed a significant upregulation of triggering receptor expressed on myeloid cells 2 in colorectal cancer-associated tumor-associated macrophages, with this upregulation exhibiting a correlation with poor patient prognosis. RESULTS: Mechanistically, triggering receptor expressed on myeloid cells 2+ tumor-associated macrophages were found to drive fibroblast activation through transforming growth factor-&#x3b2; signaling, inducing fibroblast-activated protein-positive cancer-associated fibroblasts that secrete collagen I/III to establish dense peritumoral barriers. Spatial profiling revealed that these fibrous structures physically impede CD8+ T-cell infiltration, restricting cytotoxic lymphocytes to stromal compartments. Intriguingly, triggering receptor expressed on myeloid cells 2 deficiency enhanced the secretion of matrix metalloproteinase 13 by macrophages, thereby promoting extracellular matrix degradation and improving T-cell penetration. In vivo, Trem2-knockout mice showed a reduction in tumor growth with enhanced intratumoral CD8+ T-cell infiltration compared with wild-type controls. CONCLUSIONS: Our findings establish triggering receptor expressed on myeloid cells 2+ tumor-associated macrophages as central regulators of stromal remodeling and suggest that therapeutic targeting of the triggering receptor expressed on myeloid cells 2/transforming growth factor-&#x3b2;/fibroblast-activated protein pathway may overcome immune resistance in patients with colorectal cancer.

Colorectal Neoplasms↗

Genomic analysis of secretion systems.

Secretion of proteins into the extracellular environment is important to almost all bacteria, and in particular mediates interactions between pathogenic or symbiotic bacteria with their eukaryotic hosts. The accumulation of bacterial genome sequence data in the past few years has provided great insights into the distribution and function of these secretion systems. Three systems are responsible for secretion of proteins across the bacterial cytoplasmic membrane: Sec, SRP and Tat. Many novel examples of systems for transport across the Gram-negative bacterial cell envelope have been discovered through genome sequencing and surveys, including many novel type III secretion systems and autotransporters. Similarly, genomic data mining has revealed many new potential secretion substrates and identified unsuspected domains in secretion-associated proteins. Interestingly, genomic analyses have also hinted at the existence of a dedicated protein secretion system in Gram-positive bacteria, targeting members of the WXG100/ESAT-6 family of proteins, and have revealed an unexpectedly wide distribution of sortase-driven protein-targeting systems.

Bacteria↗

The Plasmodium falciparum sexual development transcriptome: a microarray analysis using ontology-based pattern identification.

The sexual stages of malarial parasites are essential for the mosquito transmission of the disease and therefore are the focus of transmission-blocking drug and vaccine development. In order to better understand genes important to the sexual development process, the transcriptomes of high-purity stage I-V Plasmodium falciparum gametocytes were comprehensively profiled using a full-genome high-density oligonucleotide microarray. The interpretation of this transcriptional data was aided by applying a novel knowledge-based data-mining algorithm termed ontology-based pattern identification (OPI) using current information regarding known sexual stage genes as a guide. This analysis resulted in the identification of a sexual development cluster containing 246 genes, of which approximately 75% were hypothetical, exhibiting highly-correlated, gametocyte-specific expression patterns. Inspection of the upstream promoter regions of these 246 genes revealed putative cis-regulatory elements for sexual development transcriptional control mechanisms. Furthermore, OPI analysis was extended using current annotations provided by the Gene Ontology Consortium to identify 380 statistically significant clusters containing genes with expression patterns characteristic of various biological processes, cellular components, and molecular functions. Collectively, these results, available as part of a web-accessible OPI database (http://carrier.gnf.org/publications/Gametocyte), shed light on the components of molecular mechanisms underlying parasite sexual development and other areas of malarial parasite biology.

Animals↗

A comprehensive comparative analysis of the occurrence of developmental sequences in fungal, plant and animal genomes.

We report a fully comprehensive data-mining exercise, involving an estimated total of 590,000 similarity searches, using agents available on the Internet to search for homologies to polypeptide sequences assigned to the category 'development' in the Gene Ontology Consortium AmiGO database (www.godatabase.org). The results indicate that of 552 such developmental sequences only 78 are shared between all three kingdoms, 72 are shared only between fungi and animals, 58 sequences are shared between plants and fungi, and four sequences were common only to Dictyostelium and fungi. No sequences were strictly fungus specific, but 68 occurred only in plants (Viridiplantae) and 239 occurred only in animals (Metazoa). Although some homology was indicated for a total of 219 fungal sequences, 143 (65%) of the matches returned were assigned E-values of 0.05 and must be categorised as weak similarities at best. The majority of the highly similar matches found in this survey proved to be between sequences involved in basic cell metabolism or essential eukaryotic cell processes (enzymes in common metabolic pathways, transcription regulators, binding proteins, receptors and membrane proteins). What is lacking is cross-kingdom similarity in the management processes that regulate multicellular development. The crown group of eukaryotic kingdoms control and regulate their developmental processes in very different ways. Unfortunately, we know nothing about molecular control of multicellular fungal developmental biology.

Amino Acid Sequence↗

Learning and inference in the brain.

This article is about how the brain data mines its sensory inputs. There are several architectural principles of functional brain anatomy that have emerged from careful anatomic and physiologic studies over the past century. These principles are considered in the light of representational learning to see if they could have been predicted a priori on the basis of purely theoretical considerations. We first review the organisation of hierarchical sensory cortices, paying special attention to the distinction between forward and backward connections. We then review various approaches to representational learning as special cases of generative models, starting with supervised learning and ending with learning based upon empirical Bayes. The latter predicts many features, such as a hierarchical cortical system, prevalent top-down backward influences and functional asymmetries between forward and backward connections that are seen in the real brain. The key points made in article are: (i). hierarchical generative models enable the learning of empirical priors and eschew prior assumptions about the causes of sensory input that are inherent in non-hierarchical models. These assumptions are necessary for learning schemes based on information theory and efficient or sparse coding, but are not necessary in a hierarchical context. Critically, the anatomical infrastructure that may implement generative models in the brain is hierarchical. Furthermore, learning based on empirical Bayes can proceed in a biologically plausible way. (ii). The second point is that backward connections are essential if the processes generating inputs cannot be inverted, or the inversion cannot be parameterised. Because these processes involve many-to-one mappings, are non-linear and dynamic in nature, they are generally non-invertible. This enforces an explicit parameterisation of generative models (i.e. backward connections) to afford recognition and suggests that forward architectures, on their own, are not sufficient for perception. (iii). Finally, non-linearities in generative models, mediated by backward connections, require these connections to be modulatory, so that representations in higher cortical levels can interact to predict responses in lower levels. This is important in relation to functional asymmetries in forward and backward connections that have been demonstrated empirically.

Bayes Theorem↗

Self-organizing neural projections.

The Self-Organizing Map (SOM) algorithm was developed for the creation of abstract-feature maps. It has been accepted widely as a data-mining tool, and the principle underlying it may also explain how the feature maps of the brain are formed. However, it is not correct to use this algorithm for a model of pointwise neural projections such as the somatotopic maps or the maps of the visual field, first of all, because the SOM does not transfer signal patterns: the winner-take-all function at its output only defines a singular response. Neither can the original SOM produce superimposed responses to superimposed stimulus patterns. This presentation introduces a new self-organizing system model related to the SOM that has a linear transfer function for patterns and combinations of patterns all the time. Starting from a randomly interconnected pair of neural layers, and using random mixtures of patterns for training, it creates a pointwise-ordered projection from the input layer to the output layer. If the input layer consists of feature detectors, the output layer forms a feature map of the inputs.

Algorithms↗