Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Oligonucleotide Array Sequence Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

OligoArray: genome-scale oligonucleotide design for microarrays.

SUMMARY: OligoArray is a program that computes gene specific and secondary structure free oligonucleotides for genome-scale oligonucleotide microarray construction or other applications. AVAILABILITY: The program code is distributed under the GNU General Public License and is freely available for non-profit use via request from the authors.

Computer-Aided Design↗

A comparative review of statistical methods for discovering differentially expressed genes in replicated microarray experiments.

MOTIVATION: A common task in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Recently several statistical methods have been proposed to accomplish this goal when there are replicated samples under each condition. However, it may not be clear how these methods compare with each other. Our main goal here is to compare three methods, the t-test, a regression modeling approach (Thomas et al., Genome Res., 11, 1227-1236, 2001) and a mixture model approach (Pan et al., http://www.biostat.umn.edu/cgi-bin/rrs?print+2001,2001a,b) with particular attention to their different modeling assumptions. RESULTS: It is pointed out that all the three methods are based on using the two-sample t-statistic or its minor variation, but they differ in how to associate a statistical significance level to the corresponding statistic, leading to possibly large difference in the resulting significance levels and the numbers of genes detected. In particular, we give an explicit formula for the test statistic used in the regression approach. Using the leukemia data of Golub et al. (Science, 285, 531-537, 1999), we illustrate these points. We also briefly compare the results with those of several other methods, including the empirical Bayesian method of Efron et al. (J. Am. Stat. Assoc., to appear, 2001) and the Significance Analysis of Microarray (SAM) method of Tusher et al. (PROC: Natl Acad. Sci. USA, 98, 5116-5121, 2001).

Acute Disease↗

A heuristic managing errors for DNA sequencing.

MOTIVATION: A new heuristic algorithm for solving DNA sequencing by hybridization problem with positive and negative errors. RESULTS: A heuristic algorithm providing better solutions than algorithms known from the literature based on tabu search method.

Algorithms↗

A multivariate approach applied to microarray data for identification of genes with cell cycle-coupled transcription.

We have analyzed microarray data using a modeling approach based on the multivariate statistical method partial least squares (PLS) regression to identify genes with periodic fluctuations in expression levels coupled to the cell cycle in the budding yeast, Saccharomyces cerevisiae. PLS has major advantages for analyzing microarray data since it can model data sets with large numbers of variables and with few observations. A response model was derived describing the expression profile over time expected for periodically transcribed genes, and was used to identify budding yeast transcripts with similar profiles. PLS was then used to interpret the importance of the variables (genes) for the model, yielding a ranking list of how well the genes fitted the generated model. Application of an appropriate cutoff value, calculated from randomized data, allows the identification of genes whose expression appears to be synchronized with cell cycling. Our approach also provides information about the stage in the cell cycle where their transcription peaks. Three synchronized yeast cell microarray data sets were analyzed, both separately and combined. Cell cycle-coupled periodicity was suggested for 455 of the 6,178 transcripts monitored in the combined data set, at a significance level of 0.5%. Among the candidates, 85% of the known periodic transcripts were included. Analysis of the three data sets separately yielded similar ranking lists, showing that the method is robust.

Algorithms↗

ROSO: optimizing oligonucleotide probes for microarrays.

UNLABELLED: ROSO is software to design optimal oligonucleotide probe sets for microarrays. Selected probes show no significant cross-hybridization, no stable secondary structures and their Tm are chosen to minimize the Tm variability of the probe set. AVAILABILITY: The program is available on the internet. Sources are freely available, for non-profit use, on request to the authors. SUPPLEMENTARY INFORMATION: http://pbil.univ-lyon1.fr/roso

Algorithms↗

SEPON, a Selection and Evaluation Pipeline for OligoNucleotides based on ESTs with a non-target Tm algorithm for reducing cross-hybridization in microarray gene expression experiments.

SEPON, Selection and Evaluation Pipeline for OligoNucleotide generates n-mer oligonucleotide sequences from expressed sequence tags of non-annotated genomes for microarray gene-expression profiling. A non-target melting temperature (T(m)) algorithm will reduce cross-hybridization by estimating T(m) of oligonucleotide hybridization to non-specific targets (non-target T(m)) and discard oligonucleotides with non-target T(m) estimate above user-defined threshold. SEPON allows user-defined filtering, predicts exon location, assigns penalty based on 3' distance, GC content, secondary structure T(m) and non-target T(m) and ranks oligonucleotides for optimal selection.

Algorithms↗

Rainbow: a toolbox for phylogenetic supertree construction and analysis.

UNLABELLED: Rainbow is a program that provides a graphic user interface to construct supertrees using different methods. It also provides tools to analyze the quality of the supertrees produced. Rainbow is available for Mac OS X, Windows and Linux. AVAILABILITY: Rainbow is a free open-source software. Its binary files, source code, and manual can be downloaded from the Rainbow web page: http://genome.cs.iastate.edu/Rainbow/

Algorithms↗

GoArrays: highly dynamic and efficient microarray probe design.

MOTIVATION: The use of oligonucleotide microarray technology requires a very detailed attention to the design of specific probes spotted on the solid phase. These problems are far from being commonplace since they refer to complex physicochemical constraints. Whereas there are more and more publicly available programs for microarray oligonucleotide design, most of them use the same algorithm or criteria to design oligos, with only little variation. RESULTS: We show that classical approaches used in oligo design software may be inefficient under certain experimental conditions, especially when dealing with complex target mixtures. Indeed, our biological model is a human obligate parasite, the microsporidia Encephalitozoon cuniculi. Targets that are extracted from biological samples are composed of a mixture of pathogen transcripts and host cell transcripts. We propose a new approach to design oligonucleotides which combines good specificity with a potentially high sensitivity. This approach is original in the biological point of view as well as in the algorithmic point of view. We also present an experimental validation of this new strategy by comparing results obtained with standard oligos and with our composite oligos. A specific E.cuniculi microarray will overcome the difficulty to discriminate the parasite mRNAs from the host cell mRNAs demonstrating the power of the microarray approach to elucidate the lifestyle of an intracellular pathogen using mix mRNAs.

Algorithms↗

Fusing microarray experiments with multivariate regression.

MOTIVATION: It is widely acknowledged that microarray data are subject to high noise levels and results are often platform dependent. Therefore, microarray experiments should be replicated several times and in several laboratories before the results can be relied upon. To make the best use of such extensive datasets, methods for microarray data fusion are required. Ideally, the fused data should distil important aspects of the data while suppressing unwanted sources of variation and be amenable to further informal and formal methods of analysis. Also, the variability in the quality of experimentation should be taken into account. RESULTS: We present such an approach to data fusion, based on multivariate regression. We apply our methodology to data from a previous study on cell-cycle control in Schizosaccharomyces pombe. AVAILABILITY: The algorithm implemented in R is freely available from the authors on request.

Algorithms↗

YODA: selecting signature oligonucleotides.

MOTIVATION: Selecting oligonucleotide probes for use in microarray design, and other applications requiring signature sequences, involves identifying sequences which will bind strongly to their intended target, while binding only weakly (or preferably, not at all) to non-target sequences which may be present in the hybridization reaction. While many tools to assist in selection of such sequences exist, all the ones we examined lack important oligo design and software features. RESULTS: YODA is an application for assisting biological researchers in selecting signature sequences. It incorporates a custom sequence similarity search to find potential cross-hybridizing non-target sequences. For this task, most oligo design tools rely on BLAST, which is ill suited for it due to an unacceptable risk of false negatives. YODA supports multiple probe design goals including single-genome, multiple-genome, pathogen-host and species/strain-identification. A graphical interface is provided as well as a command-line interface, both of which support many user-controlled parameters. YODA is easy to install and use and runs on Windows, Mac OS X and Linux platforms. AVAILABILITY: Freely available (LGLP) along with source code and additional documentation at http://pathport.vbi.vt.edu/YODA CONTACT: enordber@vbi.vt.edu.

Algorithms↗

MADE4: an R package for multivariate analysis of gene expression data.

SUMMARY: MADE4, microarray ade4, is a software package that facilitates multivariate analysis of microarray gene-expression data. MADE4 accepts a wide variety of gene-expression data formats. MADE4 takes advantage of the extensive multivariate statistical and graphical functions in the R package ade4, extending these for application to microarray data. In addition, MADE4 provides new graphical and visualization tools that aid in interpretation of multivariate analysis of microarray data.

Algorithms↗

Hotelling's T2 multivariate profiling for detecting differential expression in microarrays.

The most widely used statistical methods for finding differentially expressed genes (DEGs) are essentially univariate. In this study, we present a new T(2) statistic for analyzing microarray data. We implemented our method using a multiple forward search (MFS) algorithm that is designed for selecting a subset of feature vectors in high-dimensional microarray datasets. The proposed T2 statistic is a corollary to that originally developed for multivariate analyses and possesses two prominent statistical properties. First, our method takes into account multidimensional structure of microarray data. The utilization of the information hidden in gene interactions allows for finding genes whose differential expressions are not marginally detectable in univariate testing methods. Second, the statistic has a close relationship to discriminant analyses for classification of gene expression patterns. Our search algorithm sequentially maximizes gene expression difference/distance between two groups of genes. Including such a set of DEGs into initial feature variables may increase the power of classification rules. We validated our method by using a spike-in HGU95 dataset from Affymetrix. The utility of the new method was demonstrated by application to the analyses of gene expression patterns in human liver cancers and breast cancers. Extensive bioinformatics analyses and cross-validation of DEGs identified in the application datasets showed the significant advantages of our new algorithm.

Algorithms↗

Increased power of microarray analysis by use of an algorithm based on a multivariate procedure.

MOTIVATION: The power of microarray analyses to detect differential gene expression strongly depends on the statistical and bioinformatical approaches used for data analysis. Moreover, the simultaneous testing of tens of thousands of genes for differential expression raises the 'multiple testing problem', increasing the probability of obtaining false positive test results. To achieve more reliable results, it is, therefore, necessary to apply adjustment procedures to restrict the family-wise type I error rate (FWE) or the false discovery rate. However, for the biologist the statistical power of such procedures often remains abstract, unless validated by an alternative experimental approach. RESULTS: In the present study, we discuss a multiplicity adjustment procedure applied to classical univariate as well as to recently proposed multivariate gene-expression scores. All procedures strictly control the FWE. We demonstrate that the use of multivariate scores leads to a more efficient identification of differentially expressed genes than the widely used MAS5 approach provided by the Affymetrix software tools (Affymetrix Microarray Suite 5 or GeneChip Operating Software). The practical importance of this finding is successfully validated using real time quantitative PCR and data from spike-in experiments. AVAILABILITY: The R-code of the statistical routines can be obtained from the corresponding author. CONTACT: Schuster@imise.uni-leipzig.de

Algorithms↗

Oligonucleotide arrays: information from replication and spatial structure.

MOTIVATION: The introduction of oligonucleotide DNA arrays has resulted in much debate concerning appropriate models for the measurement of gene expression. By contrast, little account has been taken of the possibility of identifying the physical imperfections in the raw data. RESULTS: This paper demonstrates that, with the use of replicates and an awareness of the spatial structure, deficiencies in the data can be identified, the possibility of their correction can be ascertained and correction can be effected (by use of local scaling) where possible. The procedures were motivated by data from replicates of Arabidopsis thaliana using the GeneChip ATH1-121501 microarray. Similar problems are illustrated for GeneChip Human Genome U133 arrays and for the newer and larger GeneChip Wheat Genome microarray. AVAILABILITY: R code is freely available on request.

Algorithms↗

OligoFaktory: a visual tool for interactive oligonucleotide design.

SUMMARY: The OligoFaktory is a set of tools for the design, on an arbitrary number of target sequences, of high-quality long oligonucleotide for micro-array, of primer pair for PCR, of siRNA and more. The user-centered interface exists in two flavours: a web portal and a standalone software for Mac OS X Tiger. A unified presentation of results provides overviews with distribution charts and relative location bar graphs, as well as detailed features for each oligonucleotide. Input and output files conform to a common XML interchange file format to allow both automatic generation of input data, archiving, and post-processing of results. The design pipeline can use BLAST servers to evaluate specificity of selected oligonucleotides. AVAILABILITY: The web portal http://ueg.ulb.ac.be/oligofaktory/; the software for Macintosh: http://www.oligofaktory.org/

Algorithms↗

Decoding non-unique oligonucleotide hybridization experiments of targets related by a phylogenetic tree.

MOTIVATION: The reliable identification of presence or absence of biological agents ("targets"), such as viruses or bacteria, is crucial for many applications from health care to biodiversity. If genomic sequences of targets are known, hybridization reactions between oligonucleotide probes and targets performed on suitable DNA microarrays will allow to infer presence or absence from the observed pattern of hybridization. Targets, for example all known strains of HIV, are often closely related and finding unique probes becomes impossible. The use of non-unique oligonucleotides with more advanced decoding techniques from statistical group testing allows to detect known targets with great success. Of great relevance, however, is the problem of identifying the presence of previously unknown targets or of targets that evolve rapidly. RESULTS: We present the first approach to decode hybridization experiments using non-unique probes when targets are related by a phylogenetic tree. Using a Bayesian framework and a Markov chain Monte Carlo approach we are able to identify over 94% of known targets and assign up to 70% of unknown targets to their correct clade in hybridization simulations on biological and simulated data. AVAILABILITY: Software implementing the method described in this paper and datasets are available from http://algorithmics.molgen.mpg.de/probetrees.

Algorithms↗

A multivariate approach for integrating genome-wide expression data and biological knowledge.

MOTIVATION: Several statistical methods that combine analysis of differential gene expression with biological knowledge databases have been proposed for a more rapid interpretation of expression data. However, most such methods are based on a series of univariate statistical tests and do not properly account for the complex structure of gene interactions. RESULTS: We present a simple yet effective multivariate statistical procedure for assessing the correlation between a subspace defined by a group of genes and a binary phenotype. A subspace is deemed significant if the samples corresponding to different phenotypes are well separated in that subspace. The separation is measured using Hotelling's T(2) statistic, which captures the covariance structure of the subspace. When the dimension of the subspace is larger than that of the sample space, we project the original data to a smaller orthonormal subspace. We use this method to search through functional pathway subspaces defined by Reactome, KEGG, BioCarta and Gene Ontology. To demonstrate its performance, we apply this method to the data from two published studies, and visualize the results in the principal component space.

Algorithms↗

PYCHEM: a multivariate analysis package for python.

UNLABELLED: We have implemented a multivariate statistical analysis toolbox, with an optional standalone graphical user interface (GUI), using the Python scripting language. This is a free and open source project that addresses the need for a multivariate analysis toolbox in Python. Although the functionality provided does not cover the full range of multivariate tools that are available, it has a broad complement of methods that are widely used in the biological sciences. In contrast to tools like MATLAB, PyChem 2.0.0 is easily accessible and free, allows for rapid extension using a range of Python modules and is part of the growing amount of complementary and interoperable scientific software in Python based upon SciPy. One of the attractions of PyChem is that it is an open source project and so there is an opportunity, through collaboration, to increase the scope of the software and to continually evolve a user-friendly platform that has applicability across a wide range of analytical and post-genomic disciplines. AVAILABILITY: http://sourceforge.net/projects/pychem

Algorithms↗