Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Interaction graph mining for protein complexes using local clique merging.

While recent technological advances have made available large datasets of experimentally-detected pairwise protein-protein interactions, there is still a lack of experimentally-determined protein complex data. To make up for this lack of protein complex data, we explore the mining of existing protein interaction graphs for protein complexes. This paper proposes a novel graph mining algorithm to detect the dense neighborhoods (highly connected regions) in an interaction graph which may correspond to protein complexes. Our algorithm first locates local cliques for each graph vertex (protein) and then merge the detected local cliques according to their affinity to form maximal dense regions. We present experimental results with yeast protein interaction data to demonstrate the effectiveness of our proposed method. Compared with other existing techniques, our predicted complexes can match or overlap significantly better with the known protein complexes in the MIPS benchmark database. Novel protein complexes were also predicted to help biologists in their search for new protein complexes.

Algorithms↗

Evolutionary dynamics of host-plant use in a genus of leaf-mining moths.

We used nuclear 28S rDNA sequence data to estimate the phylogeny of 77 leaf-mining Phyllonorycter (Gracillariidae) moth species, including all 55 British species, feeding on 44 different plant genera. There was strong support for both the monophyly of Phyllonorycter and the placement of the genus Cameraria as its sister group. Host-plant use was mapped onto the moth phylogeny and investigated statistically in several ways. First, we show that the estimated level of cospeciation between leaf miners and their host plants is not greater than expected by chance, despite the physical intimacy of the association. Nevertheless, the pattern of host-plant use is far from random, with closely related Phyllonorycter species generally feeding on closely related plants. However, although Phyllonorycter species from a given host plant tend to form distinct clades, there is also statistical support for multiple independent colonizations of some host-plant taxa (e.g. the order Rosales and the genus Corylus). Despite numerous host shifts, most Phyllonorycter species feed on trees and the few species that attack shrubs or herbs have mostly acquired these habits independently. There is also limited evidence that host shifts to herbs are more likely from shrubs than from trees. Similarly, most species mine the lower surface of leaves but the few upper-surface miners have each evolved the habit independently. Consequently, these shifts to new adaptive zones have not led to substantial radiations.

Animals↗

[The characteristics of mine-blast wounds in shoals].

The data obtained in experiments on shoal using plastic charges of 100-50-25 g of equivalent power of antipersonnel mines showed that injuring action on shoal was four times greater than that on land and resulted in considerably graver skeletal traumas and distant injuries. Of special significance in pathogenesis of mine-explosive wounds on shoal is pneumonia followed by arterial air embolism and encephalopathy. Although the undermining on land and on shoal have many common etiopathogenetic features, there are substantial differences first of all due to different mechanisms of their appearance. It must be taken into account while performing evacuatory, diagnostic and medical measures in such patients.

Animals↗

Identification and classification of ion-channels across the tree of life provide functional insights into understudied CALHM channels.

The ion channel (IC) genes encoded in the human genome play fundamental roles in cellular functions and disease and are one of the largest classes of druggable proteins. However, limited knowledge of the diverse molecular and cellular functions carried out by ICs presents a major bottleneck in developing selective chemical probes for modulating their functions in disease states. The wealth of sequence data available on ICs from diverse organisms provides a valuable source of untapped information for illuminating the unique modes of channel regulation and functional specialization. However, the extensive diversification of IC sequences and the lack of a unified resource present a challenge in effectively using existing data for IC research. Here, we perform integrative mining of available sequence, structure, and functional data on 419 human ICs across disparate sources, including extensive literature mining by leveraging advances in large language models to annotate and curate the full complement of the "channelome". We employ a well-established orthology inference approach to identify and extend the IC orthologs across diverse organisms to above 48,000. We show that the depth of conservation and taxonomic representation of IC sequences can further be translated to functional similarities by clustering them into functionally relevant groups, which can be used for downstream functional prediction on understudied members. We demonstrate this by delineating co-conserved patterns characteristic of the understudied family of the Calcium Homeostasis Modulator (CALHM) family of ICs. Through mutational analysis of co-conserved residues altered in human diseases and electrophysiological studies, we show that these evolutionarily-constrained residues play an important role in channel gating functions. Thus, by providing new tools and resources for performing large comparative analyses on ICs, this study addresses the unique needs of the IC community and provides the groundwork for accelerating the functional characterization of dark channels for therapeutic intervention.

CALHM1↗

Application of the GA/KNN method to SELDI proteomics data.

SUMMARY: Proteomics technology has shown promise in identifying biomarkers for disease, toxicant exposure and stress. We show by example that the genetic algorithm/k-nearest neighbors method, developed for mining high-dimensional microarray gene expression data, is also capable of mining surface enhanced laser desorption/ionization-time-of-flight proteomics data. AVAILABILITY: The source code of the program and documentation on how to use it are freely available to non-commercial users at http://dir.niehs.nih.gov/dirbb/lifiles/softlic.htm

Algorithms↗

Leveraging process integration in early drug discovery.

Recent advances in new analysis and prediction concepts in informatics, statistics and computational chemistry have drawn attention to mining the enormous flood of information generated from ultra-high-throughput screening (uHTS) and early drug discovery more effectively. This review analyses current infrastructure and process concepts in data analysis, storage and mining, with a particular focus on high-throughput technologies. It also provides examples of how these techniques have been applied successfully together with underlying reasons for these developments.

Computational Biology↗

ArrayTrack--supporting toxicogenomic research at the U.S. Food and Drug Administration National Center for Toxicological Research.

The mapping of the human genome and the determination of corresponding gene functions, pathways, and biological mechanisms are driving the emergence of the new research fields of toxicogenomics and systems toxicology. Many technological advances such as microarrays are enabling this paradigm shift that indicates an unprecedented advancement in the methods of understanding the expression of toxicity at the molecular level. At the National Center for Toxicological Research (NCTR) of the U.S. Food and Drug Administration, core facilities for genomic, proteomic, and metabonomic technologies have been established that use standardized experimental procedures to support centerwide toxicogenomic research. Collectively, these facilities are continuously producing an unprecedented volume of data. NCTR plans to develop a toxicoinformatics integrated system (TIS) for the purpose of fully integrating genomic, proteomic, and metabonomic data with the data in public repositories as well as conventional (Italic)in vitro(/Italic) and (Italic)in vivo(/Italic) toxicology data. The TIS will enable data curation in accordance with standard ontology and provide or interface a rich collection of tools for data analysis and knowledge mining. In this article the design, practical issues, and functions of the TIS are discussed through presenting its prototype version, ArrayTrack, for the management and analysis of DNA microarray data. ArrayTrack is logically constructed of three linked components: a) a library (LIB) that mirrors critical data in public databases; b) a database (MicroarrayDB) that stores microarray experiment information that is Minimal Information About a Microarray Experiment (MIAME) compliant; and c) tools (TOOL) that operate on experimental and public data for knowledge discovery. Using ArrayTrack, we can select an analysis method from the TOOL and apply the method to selected microarray data stored in the MicroarrayDB; the analysis results can be linked directly to gene information in the LIB.

Databases, Factual↗

Image mining for investigative pathology using optimized feature extraction and data fusion.

In many subspecialties of pathology, the intrinsic complexity of rendering accurate diagnostic decisions is compounded by a lack of definitive criteria for detecting and characterizing diseases and their corresponding histological features. In some cases, there exists a striking disparity between the diagnoses rendered by recognized authorities and those provided by non-experts. We previously reported the development of an Image Guided Decision Support (IGDS) system, which was shown to reliably discriminate among malignant lymphomas and leukemia that are sometimes confused with one another during routine microscopic evaluation. As an extension of those efforts, we report here a web-based intelligent archiving subsystem that can automatically detect, image, and index new cells into distributed ground-truth databases. Systematic experiments showed that through the use of robust texture descriptors and density estimation based fusion the reliability and performance of the governing classifications of the system were improved significantly while simultaneously reducing the dimensionality of the feature space.

Diagnosis, Differential↗

Mining Alzheimer disease relevant proteins from integrated protein interactome data.

Huge unrealized post-genome opportunities remain in the understanding of detailed molecular mechanisms for Alzheimer Disease (AD). In this work, we developed a computational method to rank-order AD-related proteins, based on an initial list of AD-related genes and public human protein interaction data. In this method, we first collected an initial seed list of 65 AD-related genes from the OMIM database and mapped them to 70 AD seed proteins. We then expanded the seed proteins to an enriched AD set of 765 proteins using protein interactions from the Online Predicated Human Interaction Database (OPHID). We showed that the expanded AD-related proteins form a highly connected and statistically significant protein interaction sub-network. We further analyzed the sub-network to develop an algorithm, which can be used to automatically score and rank-order each protein for its biological relevance to AD pathways(s). Our results show that functionally relevant AD proteins were consistently ranked at the top: among the top 20 of 765 expanded AD proteins, 19 proteins are confirmed to belong to the original 70 AD seed protein set. Our method represents a novel use of protein interaction network data for Alzheimer disease studies and may be generalized for other disease areas in the future.

Algorithms↗

Integrative analysis of protein interaction data.

We have developed a method for the integrative analysis of protein interaction data. It comprises clustering, visualization and data integration components. The method is generally applicable for all sequenced organisms. Here, we describe in detail the combination of protein interaction data in the yeast Saccharomyces cerevisiae with the functional classification of all yeast proteins. We evaluate the utility of the method by comparison with experimental data and deduce hypotheses about the functional role of so far uncharacterized proteins. Further applications of the integrative analysis method are discussed. The method presented here is powerful and flexible. We show that it is capable of mining large-scale data sets.

Animals↗

Association rule discovery with the train and test approach for heart disease prediction.

Association rules represent a promising technique to improve heart disease prediction. Unfortunately, when association rules are applied on a medical data set, they produce an extremely large number of rules. Most of such rules are medically irrelevant and the time required to find them can be impractical. A more important issue is that, in general, association rules are mined on the entire data set without validation on an independent sample. To solve these limitations, we introduce an algorithm that uses search constraints to reduce the number of rules, searches for association rules on a training set, and finally validates them on an independent test set. The medical significance of discovered rules is evaluated with support, confidence, and lift. Association rules are applied on a real data set containing medical records of patients with heart disease. In medical terms, association rules relate heart perfusion measurements and risk factors to the degree of disease in four specific arteries. Search constraints and test set validation significantly reduce the number of association rules and produce a set of rules with high predictive accuracy. We exhibit important rules with high confidence, high lift, or both, that remain valid on the test set on several runs. These rules represent valuable medical knowledge.

Algorithms↗

Bioinformatics approaches in clinical proteomics.

Protein expression profiling is increasingly being used to discover, validate and characterize biomarkers that can potentially be used for diagnostic purposes and to aid in pharmaceutical development. Correct analysis of data obtained from these experiments requires an understanding of the underlying analytic procedures used to obtain the data, statistical principles underlying high-dimensional data and clinical statistical tools used to determine the utility of the interpreted data. This review summarizes each of these steps, with the goal of providing the nonstatistician proteomics researcher with a working understanding of the various approaches that may be used by statisticians. Emphasis is placed on the process of mining high-dimensional data to identify a specific set of biomarkers that may be used in a diagnostic or other assay setting.

Computational Biology↗

Malaria among gold miners in southern Pará, Brazil: estimates of determinants and individual costs.

As malaria grows more prevalent in the Amazon frontier despite increased expenditures by disease control authorities, national and regional tropical disease control strategies are being called into question. The current crisis involving traditional control/eradication methods has broadened the search for feasible and effective malaria control strategies--a search that necessarily includes an investigation of the roles of a series of individual and community-level socioeconomic characteristics in determining malaria prevalence rates, and the proper methods of estimating these links. In addition, social scientists and policy makers alike know very little about the economic costs associated with malarial infections. In this paper, I use survey data from several Brazilian gold mining areas to (a) test the general reliability of malaria-related questionnaire response data, and suggest categorization methods to minimize the statistical influence of exaggerated responses, (b) estimate three statistical models aimed at detecting the socioeconomic determinants of individual malaria prevalence rates, and (c) calculate estimates of the average cost of a single bout of malaria. The results support the general reliability of survey response data gathered in conjunction with malaria research. Once the effects of vector exposure were controlled for, individual socioeconomic characteristics were only weakly linked to malaria prevalence rates in these very special miners' communities. Moreover, the socioeconomic and exposure links that were significant did not depend on the measure of malaria adopted. Finally, individual costs associated with malarial infections were found to be a significant portion of miners' incomes.

Brazil↗

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans↗

A robust hybrid between genetic algorithm and support vector machine for extracting an optimal feature gene subset.

Development of a robust and efficient approach for extracting useful information from microarray data continues to be a significant and challenging task. Microarray data are characterized by a high dimension, high signal-to-noise ratio, and high correlations between genes, but with a relatively small sample size. Current methods for dimensional reduction can further be improved for the scenario of the presence of a single (or a few) high influential gene(s) in which its effect in the feature subset would prohibit inclusion of other important genes. We have formalized a robust gene selection approach based on a hybrid between genetic algorithm and support vector machine. The major goal of this hybridization was to exploit fully their respective merits (e.g., robustness to the size of solution space and capability of handling a very large dimension of feature genes) for identification of key feature genes (or molecular signatures) for a complex biological phenotype. We have applied the approach to the microarray data of diffuse large B cell lymphoma to demonstrate its behaviors and properties for mining the high-dimension data of genome-wide gene expression profiles. The resulting classifier(s) (the optimal gene subset(s)) has achieved the highest accuracy (99%) for prediction of independent microarray samples in comparisons with marginal filters and a hybrid between genetic algorithm and K nearest neighbors.

Algorithms↗

Integrated transcriptional profiling and linkage analysis for identification of genes underlying disease.

Integration of genome-wide expression profiling with linkage analysis is a new approach to identifying genes underlying complex traits. We applied this approach to the regulation of gene expression in the BXH/HXB panel of rat recombinant inbred strains, one of the largest available rodent recombinant inbred panels and a leading resource for genetic analysis of the highly prevalent metabolic syndrome. In two tissues important to the pathogenesis of the metabolic syndrome, we mapped cis- and trans-regulatory control elements for expression of thousands of genes across the genome. Many of the most highly linked expression quantitative trait loci are regulated in cis, are inherited essentially as monogenic traits and are good candidate genes for previously mapped physiological quantitative trait loci in the rat. By comparative mapping we generated a data set of 73 candidate genes for hypertension that merit testing in human populations. Mining of this publicly available data set is expected to lead to new insights into the genes and regulatory pathways underlying the extensive range of metabolic and cardiovascular disease phenotypes that segregate in these recombinant inbred strains.

Animals↗