Search PubMed⌕ Search

PubMed · 12499309

Microarray data assembler.

Abstract

SUMMARY: Large volumes of microarray data are generated and deposited in public databases. Most of this data is in the form of tab-delimited text files or Excel spreadsheets. Combining data from several of these files to reanalyze these data sets is time consuming. Microarray Data Assembler is specifically designed to simplify this task. The program can list files and data sources, convert selected text files into Excel files and assemble data across multiple Excel worksheets and workbooks. This program thus makes data assembling easy, saves time and helps avoid manual error. AVAILABILITY: The program is freely available for non-profit use, via email request from the author, after signing a Material Transfer Agreement with Johns Hopkins University.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ramswamy Anbazhagan. 2003. Microarray data assembler.. https://doi.org/10.1093/bioinformatics%2F19.1.157

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The vaccine data link in Nha Trang, Vietnam: a progress report on the implementation of a database to detect adverse events related to vaccinations.

Real, perceived and unknown adverse events secondary to vaccinations are a source of concern for care providers of children. In the USA large linked databases have provided helpful information regarding the safety of vaccines. Very little prospectively collected data on vaccine safety is available from resource poor countries, but safety concerns may be even more relevant in such settings. Vaccine manufacturers do not have to pass the same rigorous safety standards as vaccine manufacturers in rich countries. Vaccines, which protect against cholera, Japanese encephalitis, rabies or typhoid fever are predominantly used in resource poor, tropical countries and frequently do not undergo vigorous post marketing surveillance. New vaccines specifically suited for resource poor countries are sometimes marketed without the scrutiny of vigilant, independent regulatory authorities. We describe here the design and implementation of a large linked database for a semi-rural province in central Vietnam. The design overcomes several problems inherent in data bases of medical events and vaccinations in developing countries. Assigning a permanent identification (ID) number to each resident avoids the ambiguities of ID numbers based on the address. The distribution and use of medical identification cards with a permanent ID number assists in the unambiguous identification of vaccinees and patients. Medical records of all admissions are coded according to International Classification of Diseases (ICD-10) and transcribed into a computer system. Because these processes are novel the data collected by the study will be validated. Project staff will check records on vaccinations and hospital admissions through household visits at regular intervals. Data describing vaccinations and medical events are linked to the data collected by the project staff in a computer system. Based on the validation of the data we hope to optimize this model. Once we find the model working it is planned export this vaccine data safety link to other settings of similar economic status.

Database Management Systems↗

Gene expression data preprocessing.

We present an interactive web tool for preprocessing microarray gene expression data. It analyses the data, suggests the most appropriate transformations and proceeds with them after user agreement. The normal preprocessing steps include scale transformations, management of missing values, replicate handling, flat pattern filtering and pattern standardization and they are required before performing any pattern analysis. The processed data set can be sent to other pattern analysis tools.

Database Management Systems↗

TRAIT (TRAnscript Integrated Table): a knowledgebase of human skeletal muscle transcripts.

TRAIT is a knowledgebase integrating information on transcripts with related data from genome, proteins, ortholog genes and diseases. It was initially built as a system to manage an EST-based gene discovery project on human skeletal muscle, which yielded over 4500 independent sequence clusters. Transcripts are annotated using automatic as well as manual procedures, linking known transcripts to public databases and unknown transcripts to tables of predicted features. Data are stored in a MySQL database. Complex queries are automatically built by means of a user-friendly web interface that allows the concurrent selection of many fields such as ontology, expression level, map position and protein domains. The results are parsed by the system and returned in a ranked order, in respect to the number of satisfied criteria.

Database Management Systems↗