Search PubMed⌕ Search

PubMed · 14959823

Microarray data analysis and mining.

Abstract

DNA microarray is an innovative technology for obtaining information on gene function. Because it is a high-throughput method, computational tools are essential in data analysis and mining to extract the knowledge from experimental results. Filtering procedures and statistical approaches are frequently combined to identify differentially expressed genes. However, obtaining a list of differentially expressed genes is only the starting point because an important step is the integration of differential expression profiles in a biological context, which is a hot topic in data mining. In this chapter an integrated approach of filtering and statistical validation to select trustable differentially expressed genes is described together with a brief introduction on data mining focusing on the classification of co-regulated genes on the basis of their biological function.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Silvia Saviozzi, Giovanni Iazzetti, Enrico Caserta, Alessandro Guffanti, Raffaele A Calogero. 2004. Microarray data analysis and mining.. https://doi.org/10.1385/1-59259-679-7%3A67

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Fuzzy species among recombinogenic bacteria.

BACKGROUND: It is a matter of ongoing debate whether a universal species concept is possible for bacteria. Indeed, it is not clear whether closely related isolates of bacteria typically form discrete genotypic clusters that can be assigned as species. The most challenging test of whether species can be clearly delineated is provided by analysis of large populations of closely-related, highly recombinogenic, bacteria that colonise the same body site. We have used concatenated sequences of seven house-keeping loci from 770 strains of 11 named Neisseria species, and phylogenetic trees, to investigate whether genotypic clusters can be resolved among these recombinogenic bacteria and, if so, the extent to which they correspond to named species. RESULTS: Alleles at individual loci were widely distributed among the named species but this distorting effect of recombination was largely buffered by using concatenated sequences, which resolved clusters corresponding to the three species most numerous in the sample, N. meningitidis, N. lactamica and N. gonorrhoeae. A few isolates arose from the branch that separated N. meningitidis from N. lactamica leading us to describe these species as 'fuzzy'. CONCLUSION: A multilocus approach using large samples of closely related isolates delineates species even in the highly recombinogenic human Neisseria where individual loci are inadequate for the task. This approach should be applied by taxonomists to large samples of other groups of closely-related bacteria, and especially to those where species delineation has historically been difficult, to determine whether genotypic clusters can be delineated, and to guide the definition of species.

Cluster Analysis↗

Genetic epidemiology of cancer: from families to heritable genes.

A reliable determination of familial risks for cancer is important for clinical counseling, prevention and understanding cancer etiology. Family-based gene identification efforts may be targeted if the risks are well characterized and the mode of inheritance is identified. Medically verified data on familial risks have not been available for all types of cancer but they have become available through the use of the nationwide Swedish Family-Cancer Database, which includes all Swedes born in 1932 and later with their parents, totaling over 10 million individuals. Over 150 publications have emanated from this source. The familial risks of cancer have been characterized for all main cancers and the contribution of environmental and heritable effects to the familial aggregation has been assessed. Furthermore, the mode of inheritance has been deduced by comparing risks from parental and sibling probands. Examples are shown on familial clustering of cancers, for which heritable susceptibility genes are yet unknown, such as squamous cell carcinoma of the skin, intestinal carcinoids, thyroid papillary tumors, brain astrocytomas and pituitary adenomas. Some common cancers, such as lung and kidney cancers, appear to show an early-onset recessive component because familial risks among siblings are much higher than those in families where parents are probands. Many of the cancer sites showing high familial risks lack guidelines for clinical counseling or action level. In conclusion, we recommend that any future gene identification efforts, either using linkage or association designs, devise their strategies based on data from family studies. Clinical genetic counseling would benefit from reviewing established familial risks on all main types of cancer.

Cluster Analysis↗