Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Suicide in New Zealand.

This paper explores and questions some of the notions associated with suicide including mental illness. On average, about two-thirds of suicide cases do not come into contact with mental health services, therefore, we have no objective assessment of their mental status or their life events. One method of improving our objective understanding of suicide would be to use data mining techniques in order to build life event histories on all deaths due to suicide. Although such an exercise would require major funding, partial case histories became publicly available from a coroner's inquest on cases of suicide during a period of three months in Christchurch, New Zealand. The case histories were accompanied by a newspaper article reporting comments from some of the families involved. A straightforward contextual analysis of this information suggests that (i) only five cases had contact with mental health services, in two of the cases this was due to a previous suicide attempt and in the other three it was due to drug and alcohol dependency; (ii) mental illness as the cause of suicide is fixed in the public mindset, (iii) this in turn makes psychological autopsy type studies that seek information from families and friends questionable; (iv) proportionally more females attempt, but more men tend to complete suicide; and (v) not only is the mental health-suicide relationship tenuous, but suicide also appears to be a process outcome. It is hoped that this will stimulate debate and the collaboration of international experts regardless of their school of thought.

Adolescent↗

GeneLynx mouse: integrated portal to the mouse genome.

GeneLynx Mouse is a meta-database providing an extensive collection of hyperlinks to mouse gene-specific information in diverse databases available via the Internet. The GeneLynx project is based on the simple notion that given any gene-specific identifier (e.g., accession number, gene name, text, or sequence), scientists should be able to access a single location that provides a set of links to all the publicly available information pertinent to the specified gene. The recent climax in the mouse genome and RIKEN cDNA sequencing projects provided the data necessary for the development of a gene-centric mouse information portal based on the GeneLynx ideals. Clusters of RIKEN cDNA sequences were used to define the initial set of mouse genes. Like its human counterpart, GeneLynx Mouse is designed as an extensible relational database with an intuitive and user-friendly Web interface. Data is automatically extracted from diverse resources, using appropriate approaches to maximize the coverage. To promote cross-database interoperability, an indexing utility is provided to facilitate the establishment of hyperlinks in external databases. As a result of the integration of the human and mouse systems, GeneLynx now serves as a powerful comparative genomics data mining resource. GeneLynx Mouse can be freely accessed at http://mouse.genelynx.org.

Animals↗

Characterization of topological structure on complex networks.

Characterizing the topological structure of complex networks is a significant problem especially from the viewpoint of data mining on the World Wide Web. "Page rank" used in the commercial search engine Google is such a measure of authority to rank all the nodes matching a given query. We have investigated the page-rank distribution of the real Web and a growing network model, both of which have directed links and exhibit a power law distributions of in-degree (the number of incoming links to the node) and out-degree (the number of outgoing links from the node), respectively. We find a concentration of page rank on a small number of nodes and low page rank on high degree regimes in the real Web, which can be explained by topological properties of the network, e.g., network motifs, and connectivities of nearest neighbors.

Journal Article↗

Comparative genomics of rice and Arabidopsis. Analysis of 727 cytochrome P450 genes and pseudogenes from a monocot and a dicot.

Data mining methods have been used to identify 356 Cyt P450 genes and 99 related pseudogenes in the rice (Oryza sativa) genome using sequence information available from both the indica and japonica strains. Because neither of these genomes is completely available, some genes have been identified in only one strain, and 28 genes remain incomplete. Comparison of these rice genes with the 246 P450 genes and 26 pseudogenes in the Arabidopsis genome has indicated that most of the known plant P450 families existed before the monocot-dicot divergence that occurred approximately 200 million years ago. Comparative analysis of P450s in the Pinus expressed sequence tag collections has identified P450 families that predated the separation of gymnosperms and flowering plants. Complete mapping of all available plant P450s onto the Deep Green consensus plant phylogeny highlights certain lineage-specific families maintained (CYP80 in Ranunculales) and lineage-specific families lost (CYP92 in Arabidopsis) in the course of evolution.

Arabidopsis↗

Molecular crystal global phase diagrams. I. Method of construction.

A method is described to produce global phase diagrams for single-component molecular crystals with separable internal and external modes. The phase diagrams present the equilibrium crystalline phase as a function of the coefficients of a general intermolecular potential based on rotational symmetry-adapted basis functions. It is assumed that phase transitions are driven by orientational ordering of molecules with a fixed time-averaged shape. The mean-field approximation is utilized and the process begins in a high-temperature disordered reference state, then spontaneous symmetry-breaking phase transitions and phase structure information at lower temperature are sought. The information is mapped onto phase diagrams using the intermolecular expansion coefficients as independent variables. This is illustrated by global phase diagrams for molecules having tetrahedral symmetry (e.g. carbon tetrachloride, adamantane and white phosphorus). Uses of global phase diagrams include crystal structure data mining, guidance for crystal design and enumeration of likely or missing polymorphic structures.

Journal Article↗

Applications of ACORN to data at 1.45 A resolution.

One of the main interests in the molecular biosciences is in understanding structure-function relations and X-ray crystallography plays a major role in this. ACORN can be used as a comprehensive and efficient phasing procedure for the determination of protein structures when atomic resolution data are available. An initial model can automatically be built by ARP/wARP followed by REFMAC for refinement. The alpha helices and beta sheets occurring in many protein structures can be taken as starting fragments for structure solution in ACORN. ACORN, along with ARP/wARP followed by REFMAC, can be an ab initio method for solving protein structure for which data are better than 1.2 A (atomic resolution). Attempts are here made in extending its applications to real data at 1.45 A resolution and also to truncated data at 1.6 A resolution. Two previously known structures, congerin II and alkaline cellulase N257, were resolved using the above approach. Automatic structure solution, phasing and refinement for real data at still lower resolutions for proteins of various complexities are being carried out. Data mining of the secondary structural features using PDB is being carried out for this new approach for 'seed-phasing' to ACORN.

Algorithms↗

Comparative analysis of gene sets in the Gene Ontology space under the multiple hypothesis testing framework.

The Gene Ontology (GO) resource can be used as a powerful tool to uncover the properties shared among, and specific to, a list of genes produced by high-throughput functional genomics studies, such as microarray studies. In the comparative analysis of several gene lists, researchers maybe interested in knowing which GO terms are enriched in one list of genes but relatively depleted in another. Statistical tests such as Fisher's exact test or Chi-square test can be performed to search for such GO terms. However, because multiple GO terms are tested simultaneously, individual p-values from individual tests do not serve as good indicators for picking GO terms. Furthermore, these multiple tests are highly correlated, usual multiple testing procedures that work under an independence assumption are not applicable. In this paper we introduce a procedure, based on False Discovery Rate (FDR), to treat this correlated multiple testing problem. This procedure calculates a moderately conserved estimator of q-value for every GO term. We identify the GO terms with q-values that satisfy a desired level as the significant GO terms. This procedure has been implemented into the GoSurfer software. GoSurfer is a windows based graphical data mining tool. It is freely available at http://www.gosurfer.org.

Algorithms↗

Biclustering algorithms for biological data analysis: a survey.

A large number of clustering approaches have been proposed for the analysis of gene expression data obtained from microarray experiments. However, the results from the application of standard clustering methods to genes are limited. This limitation is imposed by the existence of a number of experimental conditions where the activity of genes is uncorrelated. A similar limitation exists when clustering of conditions is performed. For this reason, a number of algorithms that perform simultaneous clustering on the row and column dimensions of the data matrix has been proposed. The goal is to find submatrices, that is, subgroups of genes and subgroups of conditions, where the genes exhibit highly correlated activities for every condition. In this paper, we refer to this class of algorithms as biclustering. Biclustering is also referred in the literature as coclustering and direct clustering, among others names, and has also been used in fields such as information retrieval and data mining. In this comprehensive survey, we analyze a large number of existing approaches to biclustering, and classify them in accordance with the type of biclusters they can find, the patterns of biclusters that are discovered, the methods used to perform the search, the approaches used to evaluate the solution, and the target applications.

Algorithms↗

Introduction to the special issue on advances in clinical and health-care knowledge management.

Clinical and health-care knowledge management (KM) as a discipline has attracted increasing worldwide attention in recent years. The approach encompasses a plethora of interrelated themes including aspects of clinical informatics, clinical governance, artificial intelligence, privacy and security, data mining, genomic mining, information management, and organizational behavior. This paper introduces key manuscripts which detail health-care and clinical KM cases and applications.

Confidentiality↗

LEAD: a methodology for learning efficient approaches to medical diagnosis.

Determining the most efficient use of diagnostic tests is one of the complex issues facing medical practitioners. With the soaring cost of healthcare, particularly in the US, there is a critical need for cutting costs of diagnostic tests, while achieving a higher level of diagnostic accuracy. This paper develops a learning based methodology that, based on patient information, recommends test(s) that optimize a suitable measure of diagnostic performance. A comprehensive performance measure is developed that accounts for the costs of testing, morbidity, and mortality associated with the tests, and time taken to reach diagnosis. The performance measure also accounts for the diagnostic ability of the tests. The methodology combines tools from the fields of data mining (rough set theory, in particular), utility theory, Markov decision processes (MDP), and reinforcement learning (RL). The rough set theory is used in extracting diagnostic information in the form of rules from the medical databases. Utility theory is used in bringing various nonhomogenous performance measures into one cost based measure. An MDP model together with an RL algorithm facilitates obtaining efficient testing strategies. The methodology is implemented on a sample problem of diagnosing solitary pulmonary nodule (SPN). The results obtained are compared with those from four alternative testing strategies. Our methodology holds significant promise to improve the process of medical diagnosis.

Algorithms↗

Finding unusual medical time-series subsequences: algorithms and applications.

In this work, we introduce the new problem of finding time series discords. Time series discords are subsequences of longer time series that are maximally different to all the rest of the time series subsequences. They thus capture the sense of the most unusual subsequence within a time series. While discords have many uses for data mining, they are particularly attractive as anomaly detectors because they only require one intuitive parameter (the length of the subsequence), unlike most anomaly detection algorithms that typically require many parameters. While the brute force algorithm to discover time series discords is quadratic in the length of the time series, we show a simple algorithm that is three to four orders of magnitude faster than brute force, while guaranteed to produce identical results. We evaluate our work with a comprehensive set of experiments on electrocardiograms and other medical datasets.

Algorithms↗

Enhancing prototype reduction schemes with recursion: a method applicable for "large" data sets.

Most of the prototype reduction schemes (PRS), which have been reported in the literature, process the data in its entirety to yield a subset of prototypes that are useful in nearest-neighbor-like classification. Foremost among these are the prototypes for nearest neighbor classifiers, the vector quantization technique, and the support vector machines. These methods suffer from a major disadvantage, namely, that of the excessive computational burden encountered by processing all the data. In this paper, we suggest a recursive and computationally superior mechanism referred to as adaptive recursive partitioning (ARP)_PRS. Rather than process all the data using a PRS, we propose that the data be recursively subdivided into smaller subsets. This recursive subdivision can be arbitrary, and need not utilize any underlying clustering philosophy. The advantage of ARP_PRS is that the PRS processes subsets of data points that effectively sample the entire space to yield smaller subsets of prototypes. These prototypes are then, in turn, gathered and processed by the PRS to yield more refined prototypes. In this manner, prototypes which are in the interior of the Voronoi spaces, and thus ineffective in the classification, are eliminated at the subsequent invocations of the PRS. We are unaware of any PRS that employs such a recursive philosophy. Although we marginally forfeit accuracy in return for computational efficiency, our experimental results demonstrate that the proposed recursive mechanism yields classification comparable to the best reported prototype condensation schemes reported to-date. Indeed, this is true for both artificial data sets and for samples involving real-life data sets. The results especially demonstrate that a fair computational advantage can be obtained by using such a recursive strategy for "large" data sets, such as those involved in data mining and text categorization applications.

Algorithms↗

Discovery of fuzzy temporal association rules.

We propose a data mining system for discovering interesting temporal patterns from large databases. The mined patterns are expressed in fuzzy temporal association rules which satisfy the temporal requirements specified by the user. Temporal requirements specified by human beings tend to be ill-defined or uncertain. To deal with this kind of uncertainty, a fuzzy calendar algebra is developed to allow users to describe desired temporal requirements in fuzzy calendars easily and naturally. Fuzzy operations are provided and users can define complicated fuzzy calendars to discover the knowledge in the time intervals that are of interest to them. A border-based mining algorithm is proposed to find association rules incrementally. By keeping useful information of the database in a border, candidate itemsets can be computed in an efficient way. Updating of the discovered knowledge due to addition and deletion of transactions can also be done efficiently. The kept information can be used to help save the work of counting and unnecessary scans over the updated database can be avoided. Simulation results show the effectiveness of the proposed system. A performance comparison with other systems is also given.

Algorithms↗

Ontology-based structured cosine similarity in document summarization: with applications to mobile audio-based knowledge management.

Development of algorithms for automated text categorization in massive text document sets is an important research area of data mining and knowledge discovery. Most of the text-clustering methods were grounded in the term-based measurement of distance or similarity, ignoring the structure of the documents. In this paper, we present a novel method named structured cosine similarity (SCS) that furnishes document clustering with a new way of modeling on document summarization, considering the structure of the documents so as to improve the performance of document clustering in terms of quality, stability, and efficiency. This study was motivated by the problem of clustering speech documents (of no rich document features) attained from the wireless experience oral sharing conducted by mobile workforce of enterprises, fulfilling audio-based knowledge management. In other words, this problem aims to facilitate knowledge acquisition and sharing by speech. The evaluations also show fairly promising results on our method of structured cosine similarity.

Algorithms↗

Enhanced automated function prediction using distantly related sequences and contextual association by PFP.

The impetus for the recent development and emergence of automated function prediction methods is an exponentially growing flood of new experimental data, the interpretation of which is hindered by a shortage of reliable annotations for proteins that lack experimental characterization or significant homologs in current databases. Here we introduce PFP, an automated function prediction server that provides the most probable annotations for a query sequence in each of the three branches of the Gene Ontology: biological process, molecular function, and cellular component. Rather than utilizing precise pattern matching to identify functional motifs in the sequences and structures of these proteins, we designed PFP to increase the coverage of function annotation by lowering resolution of predictions when a detailed function is not predictable. To do this we extend a traditional PSI-BLAST search by extracting and scoring annotations (GO terms) individually, including annotations from distantly related sequences, and applying a novel data mining tool, the Function Association Matrix, to score strongly associated pairs of annotations. We show that PFP can correctly assign function using only weakly similar sequences with a significantly better accuracy and coverage than a standard PSI-BLAST search, improving it more than fivefold. The most descriptive annotations predicted by PFP (GO depth > or = 8) can identify a significant subgraph in the GO with > 60% accuracy and approximately 100% coverage for our benchmark set. We also provide examples of the superb performance of PFP in an assessment of automated function prediction servers at the Automated Function Prediction Special Interest Group meeting at ISMB 2005 (AFP-SIG '05).

Algorithms↗

The alcohol-preferring AA and alcohol-avoiding ANA rats: neurobiology of the regulation of alcohol drinking.

The AA (alko, alcohol) and ANA (alko, non-alcohol) rat lines were among the earliest rodent lines produced by bidirectional selection for ethanol preference. The purpose of this review is to highlight the strategies for understanding the neurobiological factors underlying differential alcohol-drinking behavior in these lines. Most early work evaluated functioning of the major neurotransmitter systems implicated in drug reward in the lines. No consistent line differences were found in the dopaminergic system either under baseline conditions or after ethanol challenges. However, increased opioidergic tone in the ventral striatum and a deficiency in endocannabinoid signaling in the prefrontal cortex of AA rats may comprise mechanisms leading to increased ethanol consumption. Because complex behaviors, such as ethanol drinking, are not likely to be controlled by single factors, system-oriented molecular-profiling strategies have been used recently. Microarray based expression analysis of AA and ANA brains and novel data-mining strategies provide a system biological view that allows us to formulate a hypothesis on the mechanism underlying selection for ethanol preference. Two main factors appear active in the selection: a recruitment of signal transduction networks, including mitogen-activated protein kinases and calcium pathways and involving transcription factors such as Creb, Myc and Max, to mediate ethanol reinforcement and plasticity. The second factor acts on the mitochondrion and most likely provides metabolic flexibility for alternative substrate utilization in the presence of low amounts of ethanol.

Alcohol Drinking↗

Reverse immunology approach for the identification of CD8 T-cell-defined antigens: advantages and hurdles.

One of the challenges of tumour immunology remains the identification of strongly immunogenic tumour antigens for vaccination. Reverse immunology, that is, the procedure to predict and identify immunogenic peptides from the sequence of a gene product of interest, has been postulated to be a particularly efficient, high-throughput approach for tumour antigen discovery. Over one decade after this concept was born, we discuss the reverse immunology approach in terms of costs and efficacy: data mining with bioinformatic algorithms, molecular methods to identify tumour-specific transcripts, prediction and determination of proteasomal cleavage sites, peptide-binding prediction to HLA molecules and experimental validation, assessment of the in vitro and in vivo immunogenic potential of selected peptide antigens, isolation of specific cytolytic T lymphocyte clones and final validation in functional assays of tumour cell recognition. We conclude that the overall low sensitivity and yield of every prediction step often requires a compensatory up-scaling of the initial number of candidate sequences to be screened, rendering reverse immunology an unexpectedly complex approach.

Antigens, Neoplasm↗

Microarray analysis of changes in renal phenotype in the ethylene glycol rat model of urolithiasis: potential and pitfalls.

OBJECTIVES: To investigate, in an initial study, the use of microarray analysis (MA) to develop an information base for urolithiasis. MA enables the screening of thousands of genes simultaneously making it the technique of choice for situations where the results are known, but the underlying mechanisms are not. Little is known about the pathological changes occurring in the kidney during urolithiasis and this has severely hampered efforts to develop effective therapeutics. MATERIALS AND METHODS: Male rats were treated with 0.75% ethylene glycol for 2, 4 or 8 weeks; after death the kidneys were processed for RNA isolation and MA, conducted using a rat-based chip (one kidney/chip) and the results confirmed by reverse transcription-polymerase chain reaction (RT-PCR, 21 probe sets; control, four rats; treated, five rats). Targets were defined as different by the software if the fold change (FC) was >or= 2, and sorted into functional categories using a data-mining tool. The repeatability of MA was investigated by subjecting the 4-week samples to MA in two independent runs. RESULTS: The results for targets with a FC of >or= 2 were plotted (y = 1.01x - 0.75; r(2) 0.84). Comparing the results obtained by RT-PCR and MA showed a good qualitative correlation for those targets having a FC of >or= 5 as determined by MA. Changes in the expression of genes associated with tubule function and regulation, oxidative damage, and inflammation were the most common in the functional categories. Changes in the expression of tubule-specific markers indicated that there was damage to the proximal (gamma-adducin, organic anion and cation transporters, sodium-hydrogen exchange protein-isoform 3) and distal tubules (gamma-adducin, kallikrein) at 2 and 4 weeks. Increased expression of mitochondrial uncoupling protein indicated that there were changes to the mitochondria and oxidative stress at 2 and 4 weeks. CONCLUSION: This study shows the power of MA as an exploratory technique, and changes in the expression of several physiologically important genes whose expression has not previously been reported to be affected by hyperoxaluria or calcium oxalate crystalluria.

Animals↗