Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Identification of conserved modes of expression profiles during hippocampal development and neuronal differentiation in vitro.

Gene expression profiles can be regarded as sums of simpler modes, analogous to the modes of a vibrating violin string. Decomposition of temporal gene expression profiles into modes by singular value decomposition (SVD) was reported before, but the question as to what degree the SVD modes can be interpreted in terms of biology remains open. We report and compare the results of SVD of published datasets from hippocampal development, neuronal differentiation in vitro, and a control time-series hippocampal dataset. We demonstrate that the first SVD mode reflects the magnitude of expression, interpretable on the Affymetrix platform. In the datasets from gene profiling of hippocampal development and neuronal differentiation, the second mode reflects a monotonous change in expression, either up- or down-regulation, in the time course of experiment. We demonstrate that the top two SVD modes are conserved between datasets and therefore, likely reflect properties of the underlying system (gene expression in hippocampus) rather than of a particular experiment or dataset. Our results also indicate that the magnitude of expression, and the direction of change in expression during hippocampal development, are uncorrelated, suggesting that they are regulated by largely independent mechanisms.

Animals↗

Measurement of strain in the left ventricle during diastole with cine-MRI and deformable image registration.

The assessment of regional heart wall motion (local strain) can localize ischemic myocardial disease, evaluate myocardial viability, and identify impaired cardiac function due to hypertrophic or dilated cardiomyopathies. The objectives of this research were to develop and validate a technique known as hyperelastic warping for the measurement of local strains in the left ventricle from clinical cine-magnetic resonance imaging (MRI) image datasets. The technique uses differences in image intensities between template (reference) and target (loaded) image datasets to generate a body force that deforms a finite element (FE) representation of the template so that it registers with the target image. To validate the technique, MRI image datasets representing two deformation states of a left ventricle were created such that the deformation map between the states represented in the images was known. A beginning diastolic cine-MRI image dataset from a normal human subject was defined as the template. A second image dataset (target) was created by mapping the template image using the deformation results obtained from a forward FE model of diastolic filling. Fiber stretch and strain predictions from hyperelastic warping showed good agreement with those of the forward solution (R2=0.67 stretch, R2=0.76 circumferential strain, R2=0.75 radial strain, and R2=0.70 in-plane shear). The technique had low sensitivity to changes in material parameters (deltaR2= -0.023 fiber stretch, deltaR2=-0.020 circumferential strain, deltaR2=-0.005 radial strain, and deltaR2=0.0125 shear strain with little or no change in rms error), with the exception of changes in bulk modulus of the material. The use of an isotropic hyperelastic constitutive model in the warping analyses degraded the predictions of fiber stretch. Results were unaffected by simulated noise down to a signal-to-noise ratio (SNR) of 4.0 (deltaR2= -0.032 fiber stretch, deltaR2=-0.023 circumferential strain, deltaR2=-0.04 radial strain, and deltaAR2=0.0211 shear strain with little or no increase in rms error). This study demonstrates that warping in conjunction with cine-MRI imaging can be used to determine local ventricular strains during diastole.

Diastole↗

Assessment of analysis-of-variance-based methods to quantify the random variations of observers in medical imaging measurements: guidelines to the investigator.

The random variations of observers in medical imaging measurements negatively affect the outcome of cancer treatment, and should be taken into account during treatment by the application of safety margins that are derived from estimates of the random variations. Analysis-of-variance- (ANOVA-) based methods are the most preferable techniques to assess the true individual random variations of observers, but the number of observers and the number of cases must be taken into account to achieve meaningful results. Our aim in this study is twofold. First, to evaluate three representative ANOVA-based methods for typical numbers of observers and typical numbers of cases. Second, to establish guidelines to the investigator to determine which method, how many observers, and which number of cases are required to obtain the a priori chosen performance. The ANOVA-based methods evaluated in this study are an established technique (pairwise differences method: PWD), a new approach providing additional statistics (residuals method: RES), and a generic technique that uses restricted maximum likelihood (REML) estimation. Monte Carlo simulations were performed to assess the performance of the ANOVA-based methods, which is expressed by their accuracy (closeness of the estimates to the truth), their precision (standard error of the estimates), and the reliability of their statistical test for the significance of a difference in the random variation of an observer between two groups of cases. The highest accuracy is achieved using REML estimation, but for datasets of at least 50 cases or arrangements with 6 or more observers, the differences between the methods are negligible, with deviations from the truth well below +/-3%. For datasets up to 100 cases, it is most beneficial to increase the number of cases to improve the precision of the estimated random variations, whereas for datasets over 100 cases, an improvement in precision is most efficiently achieved by increasing the number of observers. For datasets of at least 50 cases, the standard error ranges between 30% or less with 3 observers down to 10% or less with 8 observers, and the differences in precision between the methods are negligible. The F test (PWD) is very anticonservative and should not be used, while the t test (RES) is reliable for datasets of at least 2 x 50 cases evaluated by 4 or more observers. The likelihood-ratio-test (REML estimation) consistently indicates the significance of a difference in the random variation of an observer between two groups of cases, regardless of the number of cases, and regardless of the number of observers. If a statistical package to perform REML estimation is available, and the investigator feels confident using it, this is the preferred method for studies that involve less than 50 cases evaluated by less than 6 observers. Otherwise, the RES method is an excellent alternative, because of its straightforward implementation, its completeness with respect to the provided statistics, and its overall sufficient accuracy, precision, and reliability of the provided statistical test. If neither the RES method nor REML estimation can provide sufficient performance, either more observers or more cases must be included.

Algorithms↗

Robust three-dimensional object definition in CT and MRI.

This work describes the application of an object definition algorithm to the medical imaging environment for the task of automated detection of anatomical boundaries in three dimensions in the presence of low spatial frequency nonstationarities. We have chosen the Liou-Jain algorithm and have modified it for use with 3D medical image datasets and extended it by including a recruitment operator that corrects for the algorithm's inherent volume underestimation. The algorithm avoids problems in both traditional statistical segmentation and 2D techniques and elegantly bridges the gap between traditional gradient-based edge finding and regression-based segmentation techniques. Results are shown for MRI datasets from the human abdomen and brain and for a CT dataset of a liver tumor, as well as an MRI scan of a glioma in a rat brain. For comparison, the human abdomen dataset was processed by a multivariate, statistical classifier. The results demonstrate the statistical technique's susceptibility to low spatial frequency nonstationarities due to rf field inhomogeneity; the Liou-Jain algorithm is shown to be immune to this effect. Further, the results show spatial consistency as a result of inherent characteristics of the algorithm. Volumes identified by the algorithm are visualized and assessed qualitatively in three dimensions. Quantitative accuracy of the algorithm's volume estimates is assessed by the use of a phantom. This work demonstrates that this technique is effective in automatically detecting anatomical organ and lesion surfaces in 3D medical datasets that are corrupted by low spatial frequency nonstationarity and in obtaining volume estimates.

Algorithms↗

False-positive reduction technique for detection of masses on digital mammograms: global and local multiresolution texture analysis.

We investigated the application of multiresolution global and local texture features to reduce false-positive detection in a computerized mass detection program. One hundred and sixty-eight digitized mammograms were randomly and equally divided into training and test groups. From these mammograms, two datasets were formed. The first dataset (manual) contained four regions of interest (ROIs) selected manually from each of the mammograms. One of the four ROIs contained a biopsy-proven mass and the other three contained normal parenchyma, including dense, mixed dense/fatty, and fatty tissues. The second dataset (hybrid) contained the manually extracted mass ROIs, along with normal tissue ROIs extracted by an automated Density-Weighted Contrast Enhancement (DWCE) algorithm as false-positive detections. A wavelet transform was used to decompose an ROI into several scales. Global texture features were derived from the low-pass coefficients in the wavelet transformed images. Local texture features were calculated from the suspicious object and the peripheral subregions. Linear discriminant models using effective features selected from the global, local, or combined feature spaces were established to maximize the separation between masses and normal tissue. Receiver Operating Characteristic (ROC) analysis was conducted to evaluate the classifier performance. The classification accuracy using global features were comparable to that using local features. With both global and local features, the average area, Az, under the test ROC curve, reached 0.92 for the manual dataset and 0.96 for the hybrid dataset, demonstrating statistically significant improvement over those obtained with global or local features alone. The results indicated the effectiveness of the combined global and local features in the classification of masses and normal tissue for false-positive reduction.

Biophysical Phenomena↗

Bayesian network learning with feature abstraction for gene-drug dependency analysis.

Combined analysis of the microarray and drug-activity datasets has the potential of revealing valuable knowledge about various relations among gene expressions and drug activities in the malignant cell. In this paper, we apply Bayesian networks, a tool for compact representation of the joint probability distribution, to such analysis. For the alleviation of data dimensionality problem, the huge datasets were condensed using a feature abstraction technique. The proposed analysis method was applied to the NCI60 dataset (http://discover.nci.nih.gov) consisting of gene expression profiles and drug activity patterns on human cancer cell lines. The Bayesian networks, learned from the condensed dataset, identified most of the salient pairwise correlations and some known relationships among several features in the original dataset, confirming the effectiveness of the proposed feature abstraction method. Also, a survey of the recent literature confirms the several relationships appearing in the learned Bayesian network to be biologically meaningful.

Algorithms↗

Corticosteroid-regulated genes in rat kidney: mining time series array data.

Kidney is a major target for adverse effects associated with corticosteroids. A microarray dataset was generated to examine changes in gene expression in rat kidney in response to methylprednisolone. Four control and 48 drug-treated animals were killed at 16 times after drug administration. Kidney RNA was used to query 52 individual Affymetrix chips, generating data for 15,967 different probe sets for each chip. Mining techniques applicable to time series data that identify drug-regulated changes in gene expression were applied. Four sequential filters eliminated probe sets that were not expressed in the tissue, not regulated by drug, or did not meet defined quality control standards. These filters eliminated 14,890 probe sets (94%) from further consideration. Application of judiciously chosen filters is an effective tool for data mining of time series datasets. The remaining data can then be further analyzed by clustering and mathematical modeling. Initial analysis of this filtered dataset identified a group of genes whose pattern of regulation was highly correlated with prototype corticosteroid enhanced genes. Twenty genes in this group, as well as selected genes exhibiting either downregulation or no regulation, were analyzed for 5' GRE half-sites conserved across species. In general, the results support the hypothesis that the existence of conserved DNA binding sites can serve as an important adjunct to purely analytic approaches to clustering genes into groups with common mechanisms of regulation. This dataset, as well as similar datasets on liver and muscle, are available online in a format amenable to further analysis by others.

Animals↗

Association analysis of genetic polymorphisms in the CDC2 gene with late-onset Alzheimer disease.

BACKGROUND: Alzheimer disease (AD) is a complex neurodegenerative disorder resulting from multiple genetic and non-genetic factors. Linkage studies indicated that chromosome 10 has at least one locus for this disease. The cell division cycle 2 (CDC2) gene, which is close to one of the linkage regions, has previously been associated with the risk of AD with an odds ratio of 1.78. Biologically, CDC2, which is involved in paired helical filament-tau formation, is thought as a candidate gene in AD. METHODS: In this study, six single nucleotide polymorphisms spanning the entire gene were selected and examined for association for late-onset AD (LOAD) in two large independent datasets. A family-based dataset including 1,337 Caucasian discordant sib pairs and an independent dataset of 745 Caucasian cases and 998 controls for LOAD were used. Family-based association tests and logistic regression conditional on the apolipoprotein E genotype and sex were applied to association study in family-based and case-control datasets, respectively. RESULTS: Neither dataset demonstrated any association with LOAD in our samples with all p values >0.16. CONCLUSION: Our results suggest that if any contribution of common genetic variants in CDC2 to the risk of developing AD exists, it is likely to be very small.

Aged↗

A combinational feature selection and ensemble neural network method for classification of gene expression data.

BACKGROUND: Microarray experiments are becoming a powerful tool for clinical diagnosis, as they have the potential to discover gene expression patterns that are characteristic for a particular disease. To date, this problem has received most attention in the context of cancer research, especially in tumor classification. Various feature selection methods and classifier design strategies also have been generally used and compared. However, most published articles on tumor classification have applied a certain technique to a certain dataset, and recently several researchers compared these techniques based on several public datasets. But, it has been verified that differently selected features reflect different aspects of the dataset and some selected features can obtain better solutions on some certain problems. At the same time, faced with a large amount of microarray data with little knowledge, it is difficult to find the intrinsic characteristics using traditional methods. In this paper, we attempt to introduce a combinational feature selection method in conjunction with ensemble neural networks to generally improve the accuracy and robustness of sample classification. RESULTS: We validate our new method on several recent publicly available datasets both with predictive accuracy of testing samples and through cross validation. Compared with the best performance of other current methods, remarkably improved results can be obtained using our new strategy on a wide range of different datasets. CONCLUSIONS: Thus, we conclude that our methods can obtain more information in microarray data to get more accurate classification and also can help to extract the latent marker genes of the diseases for better diagnosis and treatment.

Acute Disease↗

A comparative study of discriminating human heart failure etiology using gene expression profiles.

BACKGROUND: Human heart failure is a complex disease that manifests from multiple genetic and environmental factors. Although ischemic and non-ischemic heart disease present clinically with many similar decreases in ventricular function, emerging work suggests that they are distinct diseases with different responses to therapy. The ability to distinguish between ischemic and non-ischemic heart failure may be essential to guide appropriate therapy and determine prognosis for successful treatment. In this paper we consider discriminating the etiologies of heart failure using gene expression libraries from two separate institutions. RESULTS: We apply five new statistical methods, including partial least squares, penalized partial least squares, LASSO, nearest shrunken centroids and random forest, to two real datasets and compare their performance for multiclass classification. It is found that the five statistical methods perform similarly on each of the two datasets: it is difficult to correctly distinguish the etiologies of heart failure in one dataset whereas it is easy for the other one. In a simulation study, it is confirmed that the five methods tend to have close performance, though the random forest seems to have a slight edge. CONCLUSIONS: For some gene expression data, several recently developed discriminant methods may perform similarly. More importantly, one must remain cautious when assessing the discriminating performance using gene expression profiles based on a small dataset; our analysis suggests the importance of utilizing multiple or larger datasets.

Data Interpretation, Statistical↗

Reproducible clusters from microarray research: whither?

MOTIVATION: In cluster analysis, the validity of specific solutions, algorithms, and procedures present significant challenges because there is no null hypothesis to test and no 'right answer'. It has been noted that a replicable classification is not necessarily a useful one, but a useful one that characterizes some aspect of the population must be replicable. By replicable we mean reproducible across multiple samplings from the same population. Methodologists have suggested that the validity of clustering methods should be based on classifications that yield reproducible findings beyond chance levels. We used this approach to determine the performance of commonly used clustering algorithms and the degree of replicability achieved using several microarray datasets. METHODS: We considered four commonly used iterative partitioning algorithms (Self Organizing Maps (SOM), K-means, Clutsering LARge Applications (CLARA), and Fuzzy C-means) and evaluated their performances on 37 microarray datasets, with sample sizes ranging from 12 to 172. We assessed reproducibility of the clustering algorithm by measuring the strength of relationship between clustering outputs of subsamples of 37 datasets. Cluster stability was quantified using Cramer's v2 from a kXk table. Cramer's v2 is equivalent to the squared canonical correlation coefficient between two sets of nominal variables. Potential scores range from 0 to 1, with 1 denoting perfect reproducibility. RESULTS: All four clustering routines show increased stability with larger sample sizes. K-means and SOM showed a gradual increase in stability with increasing sample size. CLARA and Fuzzy C-means, however, yielded low stability scores until sample sizes approached 30 and then gradually increased thereafter. Average stability never exceeded 0.55 for the four clustering routines, even at a sample size of 50. These findings suggest several plausible scenarios: (1) microarray datasets lack natural clustering structure thereby producing low stability scores on all four methods; (2) the algorithms studied do not produce reliable results and/or (3) sample sizes typically used in microarray research may be too small to support derivation of reliable clustering results. Further research should be directed towards evaluating stability performances of more clustering algorithms on more datasets specially having larger sample sizes with larger numbers of clusters considered.

Cluster Analysis↗

The effect of oligonucleotide microarray data pre-processing on the analysis of patient-cohort studies.

BACKGROUND: Intensity values measured by Affymetrix microarrays have to be both normalized, to be able to compare different microarrays by removing non-biological variation, and summarized, generating the final probe set expression values. Various pre-processing techniques, such as dChip, GCRMA, RMA and MAS have been developed for this purpose. This study assesses the effect of applying different pre-processing methods on the results of analyses of large Affymetrix datasets. By focusing on practical applications of microarray-based research, this study provides insight into the relevance of pre-processing procedures to biology-oriented researchers. RESULTS: Using two publicly available datasets, i.e., gene-expression data of 285 patients with Acute Myeloid Leukemia (AML, Affymetrix HG-U133A GeneChip) and 42 samples of tumor tissue of the embryonal central nervous system (CNS, Affymetrix HuGeneFL GeneChip), we tested the effect of the four pre-processing strategies mentioned above, on (1) expression level measurements, (2) detection of differential expression, (3) cluster analysis and (4) classification of samples. In most cases, the effect of pre-processing is relatively small compared to other choices made in an analysis for the AML dataset, but has a more profound effect on the outcome of the CNS dataset. Analyses on individual probe sets, such as testing for differential expression, are affected most; supervised, multivariate analyses such as classification are far less sensitive to pre-processing. CONCLUSION: Using two experimental datasets, we show that the choice of pre-processing method is of relatively minor influence on the final analysis outcome of large microarray studies whereas it can have important effects on the results of a smaller study. The data source (platform, tissue homogeneity, RNA quality) is potentially of bigger importance than the choice of pre-processing method.

Algorithms↗

Mining gene expression data by interpreting principal components.

BACKGROUND: There are many methods for analyzing microarray data that group together genes having similar patterns of expression over all conditions tested. However, in many instances the biologically important goal is to identify relatively small sets of genes that share coherent expression across only some conditions, rather than all or most conditions as required in traditional clustering; e.g. genes that are highly up-regulated and/or down-regulated similarly across only a subset of conditions. Equally important is the need to learn which conditions are the decisive ones in forming such gene sets of interest, and how they relate to diverse conditional covariates, such as disease diagnosis or prognosis. RESULTS: We present a method for automatically identifying such candidate sets of biologically relevant genes using a combination of principal components analysis and information theoretic metrics. To enable easy use of our methods, we have developed a data analysis package that facilitates visualization and subsequent data mining of the independent sources of significant variation present in gene microarray expression datasets (or in any other similarly structured high-dimensional dataset). We applied these tools to two public datasets, and highlight sets of genes most affected by specific subsets of conditions (e.g. tissues, treatments, samples, etc.). Statistically significant associations for highlighted gene sets were shown via global analysis for Gene Ontology term enrichment. Together with covariate associations, the tool provides a basis for building testable hypotheses about the biological or experimental causes of observed variation. CONCLUSION: We provide an unsupervised data mining technique for diverse microarray expression datasets that is distinct from major methods now in routine use. In test uses, this method, based on publicly available gene annotations, appears to identify numerous sets of biologically relevant genes. It has proven especially valuable in instances where there are many diverse conditions (10's to hundreds of different tissues or cell types), a situation in which many clustering and ordering algorithms become problematic. This approach also shows promise in other topic domains such as multi-spectral imaging datasets.

Algorithms↗

The effects of multiple features of alternatively spliced exons on the K(A)/K(S) ratio test.

BACKGROUND: The evolution of alternatively spliced exons (ASEs) is of primary interest because these exons are suggested to be a major source of functional diversity of proteins. Many exon features have been suggested to affect the evolution of ASEs. However, previous studies have relied on the KA/KS ratio test without taking into consideration information sufficiency (i.e., exon length > 75 bp, cross-species divergence > 5%) of the studied exons, leading to potentially biased interpretations. Furthermore, which exon feature dominates the results of the KA/KS ratio test and whether multiple exon features have additive effects have remained unexplored. RESULTS: In this study, we collect two different datasets for analysis - the ASE dataset (which includes lineage-specific ASEs and conserved ASEs) and the ACE dataset (which includes only conserved ASEs). We first show that information sufficiency can significantly affect the interpretation of relationship between exons features and the KA/KS ratio test results. After discarding exons with insufficient information, we use a Boolean method to analyze the relationship between test results and four exon features (namely length, protein domain overlapping, inclusion level, and exonic splicing enhancer (ESE) frequency) for the ASE dataset. We demonstrate that length and protein domain overlapping are dominant factors, and they have similar impacts on test results of ASEs. In addition, despite the weak impacts of inclusion level and ESE motif frequency when considered individually, combination of these two factors still have minor additive effects on test results. However, the ACE dataset shows a slightly different result in that inclusion level has a marginally significant effect on test results. Lineage-specific ASEs may have contributed to the difference. Overall, in both ASEs and ACEs, protein domain overlapping is the most dominant exon feature while ESE frequency is the weakest one in affecting test results. CONCLUSION: The proposed method can easily find additive effects of individual or multiple factors on the KA/KS ratio test results of exons. Therefore, the system can analyze complex conditions in evolution where multiple features are involved. More factors can also be added into the system to extend the scope of evolutionary analysis of exons. In addition, our method may be useful when orthologous exons can not be found for the KA/KS ratio test.

Alternative Splicing↗

Improving the specificity of high-throughput ortholog prediction.

BACKGROUND: Orthologs (genes that have diverged after a speciation event) tend to have similar function, and so their prediction has become an important component of comparative genomics and genome annotation. The gold standard phylogenetic analysis approach of comparing available organismal phylogeny to gene phylogeny is not easily automated for genome-wide analysis; therefore, ortholog prediction for large genome-scale datasets is typically performed using a reciprocal-best-BLAST-hits (RBH) approach. One problem with RBH is that it will incorrectly predict a paralog as an ortholog when incomplete genome sequences or gene loss is involved. In addition, there is an increasing interest in identifying orthologs most likely to have retained similar function. RESULTS: To address these issues, we present here a high-throughput computational method named Ortholuge that further evaluates previously predicted orthologs (including those predicted using an RBH-based approach) - identifying which orthologs most closely reflect species divergence and may more likely have similar function. Ortholuge analyzes phylogenetic distance ratios involving two comparison species and an outgroup species, noting cases where relative gene divergence is atypical. It also identifies some cases of gene duplication after species divergence. Through simulations of incomplete genome data/gene loss, we show that the vast majority of genes falsely predicted as orthologs by an RBH-based method can be identified. Ortholuge was then used to estimate the number of false-positives (predominantly paralogs) in selected RBH-predicted ortholog datasets, identifying approximately 10% paralogs in a eukaryotic data set (mouse-rat comparison) and 5% in a bacterial data set (Pseudomonas putida - Pseudomonas syringae species comparison). Higher quality (more precise) datasets of orthologs, which we term "ssd-orthologs" (supporting-species-divergence-orthologs), were also constructed. These datasets, as well as Ortholuge software that may be used to characterize other species' datasets, are available at http://www.pathogenomics.ca/ortholuge/ (software under GNU General Public License). CONCLUSION: The Ortholuge method reported here appears to significantly improve the specificity (precision) of high-throughput ortholog prediction for both bacterial and eukaryotic species. This method, and its associated software, will aid those performing various comparative genomics-based analyses, such as the prediction of conserved regulatory elements upstream of orthologous genes.

Algorithms↗

MathDAMP: a package for differential analysis of metabolite profiles.

BACKGROUND: With the advent of metabolomics as a powerful tool for both functional and biomarker discovery, the identification of specific differences between complex metabolite profiles is becoming a major challenge in the data analysis pipeline. The task remains difficult, given the datasets' size, complexity, and common shifts in migration (elution/retention) times between samples analyzed by hyphenated mass spectrometry methods. RESULTS: We present a Mathematica (Wolfram Research, Inc.) package MathDAMP (Mathematica package for Differential Analysis of Metabolite Profiles), which highlights differences between raw datasets acquired by hyphenated mass spectrometry methods by applying arithmetic operations to all corresponding signal intensities on a datapoint-by-datapoint basis. Peak identification and integration is thus bypassed and the results are displayed graphically. To facilitate direct comparisons, the raw datasets are automatically preprocessed and normalized in terms of both migration times and signal intensities. A combination of dynamic programming and global optimization is used for the alignment of the datasets along the migration time dimension. The processed datasets and the results of direct comparisons between them are visualized using density plots (axes represent migration time and m/z values while peaks appear as color-coded spots) providing an intuitive overall view. Various forms of comparisons and statistical tests can be applied to highlight subtle differences. Overlaid electropherograms (chromatograms) corresponding to the vicinities of the candidate differences from any result may be generated in a descending order of significance for visual confirmation. Additionally, a standard library table (a list of m/z values and migration times for known compounds) may be aligned and overlaid on the plots to allow easier identification of metabolites. CONCLUSION: Our tool facilitates the visualization and identification of differences between complex metabolite profiles according to various criteria in an automated fashion and is useful for data-driven discovery of biomarkers and functional genomics.

Biomarkers↗

Biclustering of gene expression data by Non-smooth Non-negative Matrix Factorization.

BACKGROUND: The extended use of microarray technologies has enabled the generation and accumulation of gene expression datasets that contain expression levels of thousands of genes across tens or hundreds of different experimental conditions. One of the major challenges in the analysis of such datasets is to discover local structures composed by sets of genes that show coherent expression patterns across subsets of experimental conditions. These patterns may provide clues about the main biological processes associated to different physiological states. RESULTS: In this work we present a methodology able to cluster genes and conditions highly related in sub-portions of the data. Our approach is based on a new data mining technique, Non-smooth Non-Negative Matrix Factorization (nsNMF), able to identify localized patterns in large datasets. We assessed the potential of this methodology analyzing several synthetic datasets as well as two large and heterogeneous sets of gene expression profiles. In all cases the method was able to identify localized features related to sets of genes that show consistent expression patterns across subsets of experimental conditions. The uncovered structures showed a clear biological meaning in terms of relationships among functional annotations of genes and the phenotypes or physiological states of the associated conditions. CONCLUSION: The proposed approach can be a useful tool to analyze large and heterogeneous gene expression datasets. The method is able to identify complex relationships among genes and conditions that are difficult to identify by standard clustering algorithms.

Algorithms↗

FM-test: a fuzzy-set-theory-based approach to differential gene expression data analysis.

BACKGROUND: Microarray techniques have revolutionized genomic research by making it possible to monitor the expression of thousands of genes in parallel. As the amount of microarray data being produced is increasing at an exponential rate, there is a great demand for efficient and effective expression data analysis tools. Comparison of gene expression profiles of patients against those of normal counterpart people will enhance our understanding of a disease and identify leads for therapeutic intervention. RESULTS: In this paper, we propose an innovative approach, fuzzy membership test (FM-test), based on fuzzy set theory to identify disease associated genes from microarray gene expression profiles. A new concept of FM d-value is defined to quantify the divergence of two sets of values. We further analyze the asymptotic property of FM-test, and then establish the relationship between FM d-value and p-value. We applied FM-test to a diabetes expression dataset and a lung cancer expression dataset, respectively. Within the 10 significant genes identified in diabetes dataset, six of them have been confirmed to be associated with diabetes in the literature and one has been suggested by other researchers. Within the 10 significantly overexpressed genes identified in lung cancer data, most (eight) of them have been confirmed by the literatures which are related to the lung cancer. CONCLUSION: Our experiments on synthetic datasets show that FM-test is effective and robust. The results in diabetes and lung cancer datasets validated the effectiveness of FM-test. FM-test is implemented as a Web-based application and is available for free at http://database.cs.wayne.edu/bioinformatics.

Algorithms↗