Search PubMed⌕ Search

PubMed · 12372143

Mining microarray expression data by literature profiling.

Abstract

BACKGROUND: The rapidly expanding fields of genomics and proteomics have prompted the development of computational methods for managing, analyzing and visualizing expression data derived from microarray screening. Nevertheless, the lack of efficient techniques for assessing the biological implications of gene-expression data remains an important obstacle in exploiting this information. RESULTS: To address this need, we have developed a mining technique based on the analysis of literature profiles generated by extracting the frequencies of certain terms from thousands of abstracts stored in the Medline literature database. Terms are then filtered on the basis of both repetitive occurrence and co-occurrence among multiple gene entries. Finally, clustering analysis is performed on the retained frequency values, shaping a coherent picture of the functional relationship among large and heterogeneous lists of genes. Such data treatment also provides information on the nature and pertinence of the associations that were formed. CONCLUSIONS: The analysis of patterns of term occurrence in abstracts constitutes a means of exploring the biological significance of large and heterogeneous lists of genes. This approach should contribute to optimizing the exploitation of microarray technologies by providing investigators with an interface between complex expression data and large literature resources.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Damien Chaussabel, Alan Sher. 2002-09-13. Mining microarray expression data by literature profiling.. https://doi.org/10.1186/gb-2002-3-10-research0055

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Fuzzy species among recombinogenic bacteria.

BACKGROUND: It is a matter of ongoing debate whether a universal species concept is possible for bacteria. Indeed, it is not clear whether closely related isolates of bacteria typically form discrete genotypic clusters that can be assigned as species. The most challenging test of whether species can be clearly delineated is provided by analysis of large populations of closely-related, highly recombinogenic, bacteria that colonise the same body site. We have used concatenated sequences of seven house-keeping loci from 770 strains of 11 named Neisseria species, and phylogenetic trees, to investigate whether genotypic clusters can be resolved among these recombinogenic bacteria and, if so, the extent to which they correspond to named species. RESULTS: Alleles at individual loci were widely distributed among the named species but this distorting effect of recombination was largely buffered by using concatenated sequences, which resolved clusters corresponding to the three species most numerous in the sample, N. meningitidis, N. lactamica and N. gonorrhoeae. A few isolates arose from the branch that separated N. meningitidis from N. lactamica leading us to describe these species as 'fuzzy'. CONCLUSION: A multilocus approach using large samples of closely related isolates delineates species even in the highly recombinogenic human Neisseria where individual loci are inadequate for the task. This approach should be applied by taxonomists to large samples of other groups of closely-related bacteria, and especially to those where species delineation has historically been difficult, to determine whether genotypic clusters can be delineated, and to guide the definition of species.

Cluster Analysis↗

Interlaboratory random amplified polymorphic DNA typing of Yersinia enterocolitica and Y. enterocolitica-like bacteria.

A random amplified polymorphic DNA (RAPD) protocol was developed for interlaboratory use to discriminate food-borne Yersinia enterocolitica O:3 from other serogroups of Y. enterocolitica and from Y. enterocolitica-like species. Factors that were studied regarding the RAPD performance were choice of primers and concentration of PCR reagents (template DNA, MgCl(2), primer and Taq DNA polymerase). A factorial design experiment was performed to identify the optimal concentrations of the PCR reagents. The experiment showed that the concentration of the PCR reagents tested significantly affected the number of distinct RAPD products. The RAPD protocol developed was evaluated regarding its discrimination ability using 70 different Yersinia strains. Cluster analysis of the RAPD patterns obtained revealed three main groups representing (i) Y. pseudotuberculosis, (ii) Y. enterocolitica and (iii) Y. kristensenii, Y. frederiksenii, Y. intermedia and Y. ruckeri. Within the Y. enterocolitica group, the European serovar (O:3) and the North American serovar (O:8) could be clearly separated from each other. All Y. enterocolitica O:3 strains were found in one cluster which could be further divided into two subclusters, representing the geographical origin of the isolates. Thus, one of the subclusters contained Y. enterocolitica O:3 strains originating from Sweden, Finland and Norway, while Danish and English O:3 strains were found in another subcluster together with O:9 and O:5,27 strains. The repeatability (intralaboratory) and reproducibility (interlaboratory) of the RAPD protocol were tested using 15 Yersinia strains representing different RAPD patterns. The intralaboratory and the interlaboratory studies gave similarity coefficients of the same magnitude (generally >70%) for the individual strains. In the present study, it was shown that interreproducible RAPD results could be achieved by appropriate optimisation of the RAPD protocol. Furthermore, the study reflects the heterogeneous genetic diversity of the Y. enterocolitica species.

Cluster Analysis↗