Search PubMed⌕ Search

Biomedical subjects

Joseph White

Publications and source records attributed to Joseph White.

13 recordsLinked to original sources

The MGED Ontology: a resource for semantics-based description of microarray experiments.

MOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: Stoeckrt@pcbi.upenn.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

Physiogenomic resources for rat models of heart, lung and blood disorders.

Cardiovascular disorders are influenced by genetic and environmental factors. The TIGR rodent expression web-based resource (TREX) contains over 2,200 microarray hybridizations, involving over 800 animals from 18 different rat strains. These strains comprise genetically diverse parental animals and a panel of chromosomal substitution strains derived by introgressing individual chromosomes from normotensive Brown Norway (BN/NHsdMcwi) rats into the background of Dahl salt sensitive (SS/JrHsdMcwi) rats. The profiles document gene-expression changes in both genders, four tissues (heart, lung, liver, kidney) and two environmental conditions (normoxia, hypoxia). This translates into almost 400 high-quality direct comparisons (not including replicates) and over 100,000 pairwise comparisons. As each individual chromosomal substitution strain represents on average less than a 5% change from the parental genome, consomic strains provide a useful mechanism to dissect complex traits and identify causative genes. We performed a variety of data-mining manipulations on the profiles and used complementary physiological data from the PhysGen resource to demonstrate how TREX can be used by the cardiovascular community for hypothesis generation.

Animals↗

TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets.

TGICL is a pipeline for analysis of large Expressed Sequence Tags (EST) and mRNA databases in which the sequences are first clustered based on pairwise sequence similarity, and then assembled by individual clusters (optionally with quality values) to produce longer, more complete consensus sequences. The system can run on multi-CPU architectures including SMP and PVM.

Cluster Analysis↗

Within the fold: assessing differential expression measures and reproducibility in microarray assays.

BACKGROUND: 'Fold-change' cutoffs have been widely used in microarray assays to identify genes that are differentially expressed between query and reference samples. More accurate measures of differential expression and effective data-normalization strategies are required to identify high-confidence sets of genes with biologically meaningful changes in transcription. Further, the analysis of a large number of expression profiles is facilitated by a common reference sample, the construction of which must be carefully addressed. RESULTS: We carried out a series of 'self-self' hybridizations in which aliquots of the same RNA sample were labeled separately with Cy3 and Cy5 fluorescent dyes and co-hybridized to the same microarray. From this, we can analyze the intensity-dependent behavior of microarray data, define a statistically significant measure of differential expression that exploits the structure of the fluorescent signals, and measure the inherent reproducibility of the technique. We also devised a simple procedure for identifying and eliminating low-quality data for replicates within and between slides. We examine the properties required of a universal reference RNA sample and show how pooling a small number of samples with a diverse representation of expressed genes can outperform more complex mixtures as a reference sample. CONCLUSION: Analysis of cell-line samples can identify systematic structure in measured gene-expression levels. A general procedure for analyzing cDNA microarray data is proposed and validated. We show that pooled reference samples should be based not only on the expression of individual genes in each cell line but also on the expression levels of genes within cell lines.

Brain Neoplasms↗

Porcine gene discovery by normalized cDNA-library sequencing and EST cluster assembly.

Genetic and environmental factors affect the efficiency of pork production by influencing gene expression during porcine reproduction, tissue development, and growth. The identification and functional analysis of gene products important to these processes would be greatly enhanced by the development of a database of expressed porcine gene sequence. Two normalized porcine cDNA libraries (MARC 1PIG and MARC 2PIG), derived respectively from embryonic and reproductive tissues, were constructed, sequenced, and analyzed. A total of 66,245 clones from these two libraries were 5?-end sequenced and deposited in GenBank. Cluster analysis revealed that within-library redundancy is low, and comparison of all porcine ESTs with the human database suggests that the sequences from these two libraries represent portions of a significant number of independent pig genes. A Porcine Gene Index (PGI), comprising 15,616 tentative consensus sequences and 31,466 singletons, includes all sequences in public repositories and has been developed to facilitate further comparative map development and characterization of porcine genes (http://www.tigr.org/tdb/ssgi/). The clones and sequences from these libraries provide a catalog of expressed porcine genes and a resource for development of high-density hybridization arrays for transcriptional profiling of porcine tissues. In addition, comparison of porcine ESTs with sequences from other species serves as a valuable resource for comparative map development. Both arrayed cDNA libraries are available for unrestricted public use.

Animals↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Cross-referencing eukaryotic genomes: TIGR Orthologous Gene Alignments (TOGA).

Comparative genomics promises to rapidly accelerate the identification and functional classification of biologically important human genes. We developed the TIGR Orthologous Gene Alignment (TOGA; ) database to provide a cross-reference between fully and partially sequenced eukaryotic transcribed sequences. Starting with the assembled expressed sequence tag (EST) and gene sequences that comprise the 28 TIGR Gene Indices, we used high-stringency pair-wise sequence searches and a reflexive, transitive closure process to associate sequence-specific best hits, generating 32,652 tentative ortholog groups (TOGs). This has allowed us to identify putative orthologs and paralogs for known genes, as well as those that exist only as uncharacterized ESTs and to provide links to additional information including genome sequence and mapping data. TOGA provides an important new resource for the analysis of gene function in eukaryotes. In addition, an analysis of the most widely represented sequences can begin to provide insight into eukaryotic biological processes.

Algorithms↗

Three meanings of capacity; or, why the federal government is most likely to lead on insurance access issues.

This essay considers on what health policy issues the federal government is best able to lead. Positive leadership requires knowledge, power, and will. The federal government has different supplies of each for different aspects of quality of, cost of, and access to health care. Here I review technical capacity to attain desired ends, define the institutional strengths and weaknesses of the federal government, and outline current dynamics of the national political process. This analysis suggests both prospects for and some characteristics of successful policy. The federal government is more likely to lead on insurance than on other health policy issues because its supply of relevant knowledge and power is relatively high on insurance issues and the political barriers are lower than conventional wisdom suggests. But that leadership could take the form of either the expanding or contracting of access to insurance.

Decision Making, Organizational↗