Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Catalogs, Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

A comprehensive catalog of CpG islands methylated in human lung adenocarcinomas for the identification of tumor suppressor genes.

CpG island methylation is an important mechanism in gene silencing and is a key epigenetic event in cancer development. As yet, the number and identities of the genes that are inactivated in cancer cells has not been determined. In order to address this issue, we have performed a comprehensive isolation of CpG islands that are methylated in human lung adenocarcinomas. We have isolated approximately 200 CpG islands that are methylated in tumor DNA including those of known tumor-associated genes such as the HOXA5 gene. As the library contains the CpG islands of a number of known tumor suppressor genes it is highly likely that additional, previously unidentified tumor suppressor genes, will be present. On average, 1-2% of CpG islands were methylated specifically in tumors although this figure differed greatly between patients. This study provides an important resource in the search for genes inactivated in tumors and for the investigation of epigenetic dysregulation of gene expression by CpG island methylation.

Adenocarcinoma↗

Hospital-based patient education programs and the role of the hospital librarian.

This paper examines current advances in hospital-based patient education, and delineates the role of the hospital librarian in these programs. Recently, programs of planned patient education have been recognized by health care personnel and the public as being an integral part of health care delivery. Various key elements, including legislative action, the advent of audiovisual technology, and rising health care costs have contributed to the development of patient education programs in hospitals. As responsible members of the hospital organization, hospital librarians should contribute their expertise to patient education programs. They are uniquely trained with skills in providing information on other health education programs; in assembling, cataloging, and managing collections of patient education materials; and in providing documentation of their use. In order to demonstrate the full range of their skills and to contribute to patient care, education, and research, hospital librarians should actively participate in programs of planned patient education.

Cost-Benefit Analysis↗

A molecular compendium of genes expressed in multiple myeloma.

We have created a molecular resource of genes expressed in primary malignant plasma cells using a combination of cDNA library construction, 5' end single-pass sequencing, bioinformatics, and microarray analysis. In total, we identified 9732 nonredundant expressed genes. This dataset is available as the Myeloma Gene Index (www.uhnres.utoronto.ca/akstewart_lab).Predictably, the sequenced profile of myeloma cDNAs mirrored the known function of immunoglobulin-producing, high-respiratory rate, low-cycling, terminally differentiated plasma cells. Nevertheless, approximately 10% of myeloma-expressed sequences matched only entries in the database of Expressed Sequence Tags (dbEST) or the high-throughput genomic sequence (htgs) database. Numerous novel genes of potential biologic significance were identified. We therefore spotted 4300 sequenced cDNAs on glass slides creating a myeloma-enriched microarray. Several of the most highly expressed genes identified by sequencing, such as a novel putative disulfide isomerase (MGC3178), tumor rejection antigen TRA1, heat shock 70-kDa protein 5, and annexin A2, were also differentially expressed between myeloma and B lymphoma cell lines using this myeloma-enriched microarray. Furthermore, a defined subset of 34 up-regulated and 18 down-regulated genes on the array were able to differentiate myeloma from nonmyeloma cell lines. These not only include genes involved in B-cell biology such as syndecan, BCMA, PIM2, MUM1/IRF4, and XBP1, but also novel uncharacterized genes matching sequences only in the public databases. In summary, our expressed gene catalog and myeloma-enriched microarray contains numerous genes of unknown function and may complement other commercially available arrays in defining the molecular portrait of this hematopoietic malignancy. GenBank Accession numbers include BF169967-BF176369, BF185966-BF185969, and BF177280-BF177455.

Amino Acid Sequence↗

TRIPLES: a database of gene function in Saccharomyces cerevisiae.

Using a novel multipurpose mini-transposon, we have generated a collection of defined mutant alleles for the analysis of disruption phenotypes, protein localization, and gene expression in Saccharomyces cerevisiae. To catalog this unique data set, we have developed TRIPLES, a Web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces. Encompassing over 250 000 data points, TRIPLES provides convenient access to information from nearly 7800 transposon-mutagenized yeast strains; within TRIPLES, complete data reports of each strain may be viewed in table format, or if desired, downloaded as tab-delimited text files. Each report contains external links to corresponding entries within the Saccharomyces Genome Database and International Nucleic Acid Sequence Data Library (GenBank). Unlike other yeast databases, TRIPLES also provides on-line order forms linked to each clone report; users may immediately request any desired strain free-of-charge by submitting a completed form. In addition to presenting a wealth of information for over 2300 open reading frames, TRIPLES constitutes an important medium for the distribution of useful reagents throughout the yeast scientific community. Maintained by the Yale Genome Analysis Center, TRIPLES may be accessed at http://ycmi.med.yale.edu/ygac/triples.htm

DNA Transposable Elements↗

Identification of alternate polyadenylation sites and analysis of their tissue distribution using EST data.

Alternate polyadenylation affects a large fraction of higher eucaryote mRNAs, producing mature transcripts with 3' ends of variable length. This variation is poorly represented in the current transcript catalogs derived from whole genome sequences, mostly because such posttranscriptional events are not detectable directly at the DNA level. Alternate polyadenylation of an mRNA is better understood by comparison to EST databases. Comparing ESTs to mRNAs, however, is a difficult task subjected to the pitfalls of internal priming, presence of intron sequences, repeated elements, chimerical ESTs or matches with EST from paralogous genes. We present here a computer program that addresses these problems and displays ESTs matches to a query mRNA sequence to predict alternate polyadenylation and to suggest library-specific forms. The output highlights effective polyadenylation signals, possible sources of artifacts such as A-rich stretches in the mRNA sequences, and allows for a direct visualization of EST libraries using color codes. Statistical biases in the distribution of alternative mRNA forms among EST libraries were systematically sought. About 1450 human and 200 mouse mRNAs displayed such biases, suggesting in each case a tissue- or disease-specific regulation of polyadenylation.

3' Untranslated Regions↗

A hierarchical clustering approach for large compound libraries.

A modified version of the k-means clustering algorithm was developed that is able to analyze large compound libraries. A distance threshold determined by plotting the sum of radii of leaf clusters was used as a termination criterion for the clustering process. Hierarchical trees were constructed that can be used to obtain an overview of the data distribution and inherent cluster structure. The approach is also applicable to ligand-based virtual screening with the aim to generate preferred screening collections or focused compound libraries. Retrospective analysis of two activity classes was performed: inhibitors of caspase 1 [interleukin 1 (IL1) cleaving enzyme, ICE] and glucocorticoid receptor ligands. The MDL Drug Data Report (MDDR) and Collection of Bioactive Reference Analogues (COBRA) databases served as the compound pool, for which binary trees were produced. Molecules were encoded by all Molecular Operating Environment 2D descriptors and topological pharmacophore atom types. Individual clusters were assessed for their purity and enrichment of actives belonging to the two ligand classes. Significant enrichment was observed in individual branches of the cluster tree. After clustering a combined database of MDDR, COBRA, and the SPECS catalog, it was possible to retrieve MDDR ICE inhibitors with new scaffolds using COBRA ICE inhibitors as seeds. A Java implementation of the clustering method is available via the Internet (http://www.modlab.de).

Algorithms↗

MiST: a microbial signal transduction database.

Signal transduction pathways control most cellular activities in living cells ranging from regulation of gene expression to fine-tuning enzymatic activity and controlling motile behavior in response to extracellular and intracellular signals. Because of their extreme sequence variability and extensive domain shuffling, signal transduction proteins are difficult to identify, and their current annotation in most leading databases is often incomplete or erroneous. To overcome this problem, we have developed the microbial signal transduction (MiST) database (http://genomics.ornl.gov/mist), a comprehensive library of the signal transduction proteins from completely sequenced bacterial and archaeal genomes. By searching for domain profiles that implicate a particular protein as participating in signal transduction, we have systematically identified 69 270 two- and one-component proteins in 365 bacterial and archaeal genomes. We have designed a user-friendly website to access and browse the predicted signal transduction proteins within various organisms. Further capabilities include gene/protein sequence retrieval, visualized domain architectures, interactive chromosomal views for exploring gene neighborhood, advanced querying options and cross-species comparison. Newly available, complete genomes are loaded into the database each month. MiST is the only comprehensive and up-to-date electronic catalog of the signaling machinery in microbial genomes.

Archaeal Proteins↗

CISMeF: cataloque and index of French speaking health resources.

In 1999, the Internet has become a major source of health information. The objective of CISMeF is to catalogue and index the main French-speaking sites and documents concerning health. Currently, the number of resources already totalled over 6,100 with a mean of 75 new sites each week. CISMeF contains a thematic index, including medical specialities and an alphabetic index. CISMeF uses two standard tools for organising information: the MeSH (Medical Subject Heading) thesaurus from the Medline bibliographic database (National Library of Medicine) and the Dublin Core metadata format. A brief description of the site is systematically added. CISMeF respects the Net Scoring, criteria to assess the quality of health information on the Internet. The CISMeF project fulfils a valuable tool for the French-speaking health community: 2,500 machines visit the Web site each working day.

Abstracting and Indexing↗

NIPALSTREE: a new hierarchical clustering approach for large compound libraries and its application to virtual screening.

A hierarchical clustering algorithm--NIPALSTREE--was developed that is able to analyze large data sets in high-dimensional space. The result can be displayed as a dendrogram. At each tree level the algorithm projects a data set via principle component analysis onto one dimension. The data set is sorted according to this one dimension and split at the median position. To avoid distortion of clusters at the median position, the algorithm identifies a potentially more suited split point left or right of the median. The procedure is recursively applied on the resulting subsets until the maximal distance between cluster members exceeds a user-defined threshold. The approach was validated in a retrospective screening study for angiotensin converting enzyme (ACE) inhibitors. The resulting clusters were assessed for their purity and enrichment in actives belonging to this ligand class. Enrichment was observed in individual branches of the dendrogram. In further retrospective virtual screening studies employing the MDL Drug Data Report (MDDR), COBRA, and the SPECS catalog, NIPALSTREE was compared with the hierarchical k-means clustering approach. Results show that both algorithms can be used in the context of virtual screening. Intersecting the result lists obtained with both algorithms improved enrichment factors while losing only few chemotypes.

Algorithms↗

Catalog of gene expression in adult neural stem cells and their in vivo microenvironment.

Stem cells generally reside in a stem cell microenvironment, where cues for self-renewal and differentiation are present. However, the genetic program underlying stem cell proliferation and multipotency is poorly understood. Transcriptome analysis of stem cells and their in vivo microenvironment is one way of uncovering the unique stemness properties and provides a framework for the elucidation of stem cell function. Here, we characterize the gene expression profile of the in vivo neural stem cell microenvironment in the lateral ventricle wall of adult mouse brain and of in vitro proliferating neural stem cells. We have also analyzed an Lhx2-expressing hematopoietic-stem-cell-like cell line in order to define the transcriptome of a well-characterized and pure cell population with stem cell characteristics. We report the generation, assembly and annotation of 50,792 high-quality 5'-end expressed sequence tag sequences. We further describe a shared expression of 1065 transcripts by all three stem cell libraries and a large overlap with previously published gene expression signatures for neural stem/progenitor cells and other multipotent stem cells. The sequences and cDNA clones obtained within this framework provide a comprehensive resource for the analysis of genes in adult stem cells that can accelerate future stem cell research.

Animals↗

Massively parallel characterization of adolescent idiopathic scoliosis risk variants.

Adolescent idiopathic scoliosis (AIS) is a common pediatric musculoskeletal disorder characterized by lateral spinal curvature, often leading to chronic pain and deformity. Although a significant genetic component to AIS is recognized, the functional impact of most associated genetic variants, particularly those in noncoding regions, remains largely unknown. Using massively parallel reporter assays, we characterize 1664 variant positions in linkage disequilibrium with 26 AIS lead variants identified by genome-wide association studies (GWASs) in chondrocytes, a major cell type implicated in AIS pathogenesis. Using a library of 7173 candidate regulatory sequences, we compare the 1664 reference alleles against 4708 alternate alleles in two human chondrocyte cell lines (TC28a2 and SW1353). Our analysis identifies 92 variants that exhibit significant differential regulatory activity between their reference and alternate alleles, 79 of which are predicted to disrupt transcription factor binding sites, often correlating with their observed regulatory effect. Notably, we validate rs9496392, a single-nucleotide variant near the ADGRG6 locus, which shows consistent differential regulatory activity in both cell lines. ADGRG6 is a key regulator of cartilage homeostasis, and its cartilage-specific knockout in mice results in a scoliosis-like phenotype. The AIS risk allele of rs9496392 (T) is predicted to strongly disrupt several TFBSs, including SP1. This study provides a foundational catalog of functional AIS-associated regulatory variants active in chondrocytes, offering crucial insights into the perturbed gene regulatory networks in AIS. These findings lay the groundwork for identifying biomarkers and potential therapeutic targets for this complex childhood disease.

Journal Article↗

A protein linkage map of the P2 nonstructural proteins of poliovirus.

The yeast two-hybrid system was used to catalog all detectable interactions among the P2 nonstructural cleavage products of poliovirus type 1 (Mahoney). Evidence has been obtained for specific associations among 2A(pro), 2BC, 2C, and 2B. Specifically, 2A(pro) can interact with itself and 2BC and its cleavage products (2B and 2C) interact in all possible combinations, with the exception of 2C/2C. Detected interactions were confirmed in vitro by a glutathione S-transferase pulldown assay, which allowed us to detect 2C/2C association. transdominant-negative mutants of 2B (K. Johnson and P. J. Sarnow, J. Virol. 65:4341-4349, 1991) were examined and were found to retain interaction with wild-type 2B, perhaps reflecting a need for 2B multimerization in viral RNA replication. The multimerization of 2B was examined further by screening a mutagenized library for 2B variants that have lost the ability to bind wild-type 2B. The screen identified two nonconservative missense mutations within a central hydrophobic region, as well as truncations and frameshifts that implicate the C terminus in homointeraction. Introduction of the missense mutations into the genome of the virus conferred a quasi-infectious phenotype, an observation strongly suggesting that the 2B/2B interaction is required for replication of the viral genome.

Carrier Proteins↗

Prediction of unidentified human genes on the basis of sequence similarity to novel cDNAs from cynomolgus monkey brain.

BACKGROUND: The complete assignment of the protein-coding regions of the human genome is a major challenge for genome biology today. We have already isolated many hitherto unknown full-length cDNAs as orthologs of unidentified human genes from cDNA libraries of the cynomolgus monkey (Macaca fascicularis) brain (parietal lobe and cerebellum). In this study, we used cDNA libraries of three other parts of the brain (frontal lobe, temporal lobe and medulla oblongata) to isolate novel full-length cDNAs. RESULTS: The entire sequences of novel cDNAs of the cynomolgus monkey were determined, and the orthologous human cDNA sequences were predicted from the human genome sequence. We predicted 29 novel human genes with putative coding regions sharing an open reading frame with the cynomolgus monkey, and we confirmed the expression of 21 pairs of genes by the reverse transcription-coupled polymerase chain reaction method. The hypothetical proteins were also functionally annotated by computer analysis. CONCLUSIONS: The 29 new genes had not been discovered in recent explorations for novel genes in humans, and the ab initio method failed to predict all exons. Thus, monkey cDNA is a valuable resource for the preparation of a complete human gene catalog, which will facilitate post-genomic studies.

Animals↗

The gene-protein database of Escherichia coli: edition 5.

The gene-protein database of Escherichia coli is both an index relating a gene to its protein product on two-dimensional gels, and a catalog of information about the function, regulation, and genetics of individual proteins obtained from two-dimensional gel analysis or collated from the literature. Edition 5 has 102 new entries--a 15% increase in the number of annotated two-dimensional gel spots. The large increase in this edition was accomplished in part by the use of a new method for expression analysis of ordered segments of the E. coli genome, which has resulted in linking 50 gel spots to their genes (or open reading frames) and another 45 to specific regions of the chromosome awaiting the availability of DNA sequence information. Communication of information from the scientific community resulted in additional identifications and regulatory information. To increase accessibility of the database it has been placed in the repository at the National Center for Biotechnology Information (NCBI) at the National Library of Medicine under the name ECO2DBASE. It will be updated twice yearly. This edition of the gene-protein database is estimated to contain entries for one-sixth of the protein-encoding genes of E. coli.

Bacterial Proteins↗

SAGE identification of gene transcripts with profiles unique to pluripotent mouse R1 embryonic stem cells.

The identification of signals that regulate pluripotentiality and self-renewal is fundamental to the understanding of stem cell biology. To quantify the functionally active genome of pluripotent R1 embryonic stem (ES) cells, we used the method of serial analysis of gene expression (SAGE) to sequence a total of 140,313 SAGE tags. Of 44,569 unique transcripts, 9% matched known genes in the nonredundant GenBank database, whereas >35% of the unique tags did not match any known mouse sequence. Comparisons of relatively abundant (> or = 20) tags in the ES cell SAGE catalog with publicly available SAGE data sets identified 16 transcripts with an abundance profile unique to pluripotent R1 ES cells. We confirmed 12 by RT-PCR including those encoding KLF2, a transcription factor; galanin, a hypothalamic neurohormone; BAX, a proapoptotic signaling factor; and CDK4 and PAL31, cell cycle progression associated proteins. The data from this study provide a starting point for detailed transcriptome analyses of stem cells.

Animals↗

Characterization and repeat analysis of the compact genome of the freshwater pufferfish Tetraodon nigroviridis.

Tetraodon nigroviridis is a freshwater pufferfish 20-30 million years distant from Fugu rubripes. The genome of both tetraodontiforms is compact, mostly because intergenic and intronic sequences are reduced in size compared to other vertebrate genomes. The previously uncharacterized Tetraodon genome is described here together with a detailed analysis of its repeat content and organization. We report the sequencing of 46 megabases of bacterial artificial chromosome (BAC) end sequences, which represents a random DNA sample equivalent to 13% of the genome. The sequence and location of rRNA gene clusters, centromeric and subtelocentric satellite sequences have been determined. Minisatellites and microsatellites have been cataloged and notable differences were observed in comparison with microsatellites from Fugu. The genome contains homologies to all known families of transposable elements, including Ty3-gypsy, Ty1-copia, Line retrotransposons, DNA transposons, and retroviruses, although their overall abundance is <1%. This structural analysis is an important prerequisite to sequencing the Tetraodon genome.

Animals↗

Toward a catalog of human genes and proteins: sequencing and analysis of 500 novel complete protein coding human cDNAs.

With the complete human genomic sequence being unraveled, the focus will shift to gene identification and to the functional analysis of gene products. The generation of a set of cDNAs, both sequences and physical clones, which contains the complete and noninterrupted protein coding regions of all human genes will provide the indispensable tools for the systematic and comprehensive analysis of protein function to eventually understand the molecular basis of man. Here we report the sequencing and analysis of 500 novel human cDNAs containing the complete protein coding frame. Assignment to functional categories was possible for 52% (259) of the encoded proteins, the remaining fraction having no similarities with known proteins. By aligning the cDNA sequences with the sequences of the finished chromosomes 21 and 22 we identified a number of genes that either had been completely missed in the analysis of the genomic sequences or had been wrongly predicted. Three of these genes appear to be present in several copies. We conclude that full-length cDNA sequencing continues to be crucial also for the accurate identification of genes. The set of 500 novel cDNAs, and another 1000 full-coding cDNAs of known transcripts we have identified, adds up to cDNA representations covering 2%--5 % of all human genes. We thus substantially contribute to the generation of a gene catalog, consisting of both full-coding cDNA sequences and clones, which should be made freely available and will become an invaluable tool for detailed functional studies.

3' Untranslated Regions↗

Use of molecular variation in the NCBI dbSNP database.

While high quality information regarding variation in genes is currently available in locus-specific or specialized mutation databases, the need remains for a general catalog of genome variation to address the large-scale sampling designs required by association studies, gene mapping, and evolutionary biology. In response to this need, the National Center for Biotechnology Information (NCBI) has established the dbSNP database http://ncbi. nlm.nih.gov/SNP/ to serve as a generalized, central variation database. Submissions to dbSNP will be integrated with other sources of information at NCBI such as GenBank, PubMed, LocusLink, and the Human Genome Project data, and the complete contents of dbSNP are available to the public via anonymous FTP. Hum Mutat 15:68-75, 2000. Published 2000 Wiley-Liss, Inc.

Databases, Factual↗