Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

High-throughput proteomic analysis of human infiltrating ductal carcinoma of the breast.

Large-scale proteomics will play a critical role in the rapid display, identification and validation of new protein targets, and elucidation of the underlying molecular events that are associated with disease development, progression and severity. However, because the proteome of most organisms are significantly more complex than the genome, the comprehensive analysis of protein expression changes will require an analytical effort beyond the capacity of standard laboratory equipment. We describe the first high-throughput proteomic analysis of human breast infiltrating ductal carcinoma (IDCA) using OCT (optimal cutting temperature) embedded biopsies, two-dimensional difference gel electrophoresis (2-D DIGE) technology and a fully automated spot handling workstation. Total proteins from four breast IDCAs (Stage I, IIA, IIB and IIIA) were individually compared to protein from non-neoplastic tissue obtained from a female donor with no personal or family history of breast cancer. We detected differences in protein abundance that ranged from 14.8% in stage I IDCA versus normal, to 30.6% in stage IIB IDCA versus normal. A total of 524 proteins that showed > or = three-fold difference in abundance between IDCA and normal tissue were picked, processed and identified by mass spectrometry. Out of the proteins picked, approximately 80% were unambiguously assigned identities by matrix-assisted laser desorbtion/ionization-time of flight mass spectrometry or liquid chromatography-tandem mass spectrometry in the first pass. Bioinformatics tools were also used to mine databases to determine if the identified proteins are involved in important pathways and/or interact with other proteins. Gelsolin, vinculin, lumican, alpha-1-antitrypsin, heat shock protein-60, cytokeratin-18, transferrin, enolase-1 and beta-actin, showed differential abundance between IDCA and normal tissue, but the trend was not consistent in all samples. Out of the proteins with database hits, only heat shock protein-70 (more abundant) and peroxiredoxin-2 (less abundant) displayed the same trend in all the IDCAs examined. This preliminary study demonstrates quantitative and qualitative differences in protein abundance between breast IDCAs and reveals 2-D DIGE portraits that may be a reflection of the histological and pathological status of breast IDCA.

Adult↗

Identification and analysis of the mouse basic/Helix-Loop-Helix transcription factor family.

The basic/Helix-Loop-Helix (bHLH) proteins are a family of transcription factors that regulates a variety of biological processes. Based on a previously defined consensus motif, we identified the complete set of bHLH protein family from the mouse proteome databases and carried out a series of bioinformatics analysis. As results, 124 mouse bHLH proteins were identified in this study, and 28 of them were additional bHLH proteins beyond the previous report. These 124 mouse bHLH proteins were classified into groups from A to F by the nomenclature and phylogenetic analysis. Statistic analysis of the Gene Ontology annotation of these proteins showed that the bHLH proteins tend to perform functions related to cell differentiation and development. Gene function enrichment analysis among six groups illuminated that the proteins in certain group tend to have special biology functions, so that the molecular function of the uncharacterized proteins in groups could be inferred.

Amino Acid Sequence↗

Characterisation of Teladorsagia circumcincta microsatellites and their development as population genetic markers.

There is a need to develop tools to study the genetics of parasitic nematodes. This is particularly urgent for those species in which anthelmintic resistance is common such as the sheep parasite Teladorsagia (Ostertagia) circumcincta. The lack of information on the population genetics of such parasites severely limits our ability to study the genetic basis of anthelmintic resistance. This paper presents the results of three approaches used to isolate microsatellite markers from T. circumcincta and the development of a panel of markers suitable for population genetic analysis. Hybridisation screening of small insert genomic libraries and interspecies PCR amplification of Haemonchus contortus microsatellites were used to identify CA/GT microsatellites. Many of these loci were associated with a 146bp tandem repeat, named TecRep, that is related to a repetitive element previously identified in other trichostrongylid nematode genomes but apparently absent from other nematode groups. A large proportion of the loci isolated were problematic for use as population genetic markers, predominantly due to a high frequency of null alleles or the association with the TecRep repeat. Bioinformatic screening of a T. circumcincta EST database identified both di- and tri-nucleotide microsatellite repeats and a greater proportion of these turned out to be more robust markers than those derived from genomic sequence. A panel of seven markers has been selected and characterised which are sufficiently robust and polymorphic to be valuable population genetic markers for this parasite.

Animals↗

Efficiency and limits of the Serial Analysis of Gene Expression (SAGE) method: discussions based on first results in bovine trypanotolerance.

Post genomic biotechnologies, such as transcriptome analysis, are now efficient enough to characterize the full complement of genes involved in the expression of specific biological functions. One of them is the Serial Analysis of Gene Expression (SAGE) technique. SAGE involves the construction of transcript libraries for a quantitative analysis of the entire set of genes expressed or inactivated at particular stages of cellular activation. Bioinformatic comparisons in hosts and pathogens genomic databases allow the identification of several up- and down-regulated genes, ESTs and unknown transcripts directly involved in the host-pathogen immunological interaction mechanisms. Based on the first results obtained during an experimental Trypanosoma congolense infection in trypanotolerant cattle, the efficiency and limits of such a technique, from the data acquisition level to the data analysis level, is discussed in this analysis.

Animals↗

'Harvester': a fast meta search engine of human protein resources.

SUMMARY: We have developed a Web-based tool named 'Harvester' that bulk-collects bioinformatic data on human proteins from various databases and prediction servers. The information on every single protein is assembled on a single HTML page as a combination of database screen-shots and plain text. A full text meta search engine, similar to Google trade mark, allows screening of the whole genome proteome for current protein functions and predictions in a few seconds. With Harvester it is now possible to compare and check the quality of different database entries and prediction algorithms on a single page. A feedback forum allows users to comment on Harvester and to report database inconsistencies. AVAILABILITY: The service is freely available to the academic community at http://harvester.embl.de.

Database Management Systems↗

SIMAP--the similarity matrix of proteins.

MOTIVATION: Sequence similarity searches are of great importance in bioinformatics. Exhaustive searches for homologous proteins in databases are computationally expensive and can be replaced by a database of pre-calculated homologies in many cases. Retrieving similarities from an incrementally updated database instead of repeatedly recalculating them should provide homologs much faster and frees computational resources for other purposes. RESULTS: We have implemented SIMAP-a database containing the similarity space formed by almost all amino acid sequences from public databases and completely sequenced genomes. The database is capable of handling very large datasets and allows incremental updates. We have implemented a powerful backbone for similarity computation, which is based on FASTA heuristics. By providing WWW interfaces as well as web services, we make our data accessible to the worldwide community. We have also adapted procedures to detect putative orthologs as example applications. AVAILABILITY: The SIMAP portal page providing links to SIMAP services is publicly available: http://mips.gsf.de/services/analysis/simap/. The web services can be accessed under http://mips.gsf.de/proj/hobitws/services/RPCSimapService?wsdl and http://mips.gsf.de/proj/hobitws/services/DocSimapService?wsdl

Algorithms↗

ApiDB: integrated resources for the apicomplexan bioinformatics resource center.

ApiDB (http://ApiDB.org) represents a unified entry point for the NIH-funded Apicomplexan Bioinformatics Resource Center (BRC) that integrates numerous database resources and multiple data types. The phylum Apicomplexa comprises numerous veterinary and medically important parasitic protozoa including human pathogenic species of the genera Cryptosporidium, Plasmodium and Toxoplasma. ApiDB serves not only as a database in its own right, but as a single web-based point of entry that unifies access to three major existing individual organism databases (PlasmoDB.org, ToxoDB.org and CryptoDB.org), and integrates these databases with data available from additional sources. Through the ApiDB site, users may pose queries and search all available apicomplexan data and tools, or they may visit individual component organism databases.

Animals↗

Text mining biomedical literature for discovering gene-to-gene relationships: a comparative study of algorithms.

Partitioning closely related genes into clusters has become an important element of practically all statistical analyses of microarray data. A number of computer algorithms have been developed for this task. Although these algorithms have demonstrated their usefulness for gene clustering, some basic problems remain. This paper describes our work on extracting functional keywords from MEDLINE for a set of genes that are isolated for further study from microarray experiments based on their differential expression patterns. The sharing of functional keywords among genes is used as a basis for clustering in a new approach called BEA-PARTITION in this paper. Functional keywords associated with genes were extracted from MEDLINE abstracts. We modified the Bond Energy Algorithm (BEA), which is widely accepted in psychology and database design but is virtually unknown in bioinformatics, to cluster genes by functional keyword associations. The results showed that BEA-PARTITION and hierarchical clustering algorithm outperformed k-means clustering and self-organizing map by correctly assigning 25 of 26 genes in a test set of four known gene groups. To evaluate the effectiveness of BEA-PARTITION for clustering genes identified by microarray profiles, 44 yeast genes that are differentially expressed during the cell cycle and have been widely studied in the literature were used as a second test set. Using established measures of cluster quality, the results produced by BEA-PARTITION had higher purity, lower entropy, and higher mutual information than those produced by k-means and self-organizing map. Whereas BEA-PARTITION and the hierarchical clustering produced similar quality of clusters, BEA-PARTITION provides clear cluster boundaries compared to the hierarchical clustering. BEA-PARTITION is simple to implement and provides a powerful approach to clustering genes or to any clustering problem where starting matrices are available from experimental observations.

Abstracting and Indexing↗

Membrane-associated and secreted genes in breast cancer.

The identification of membrane-associated and secreted genes that are differentially expressed is a useful step in defining new targets for the diagnosis and treatment of cancer. Extracting information on the subcellular localization of genes represented on DNA microarrays is difficult and is limited by the incomplete sequence and annotation that is available in existing databases. Here we combine a biochemical and bioinformatic approach to identify membrane-associated and secreted genes expressed in the MCF-7 breast cancer cell line. Our approach is based on the analysis of differential hybridization levels of RNAs that have been physically separated by virtue of their association with polysomes on the endoplasmic reticulum. This approach is specifically applicable to oligonucleotide microarrays such as Affymetrix, which use single-color hybridization instead of dual-color competitive hybridizations. Assignment to membrane-associated and secreted class membership is based on both the differential hybridization levels and an expression threshold, which are calculated empirically from data collected on a reference set of known cytoplasmic and membrane proteins. This method enabled the identification of 755 membrane-associated and secreted probe sets expressed in MCF-7 cells for which this annotation did not previously exist. The data were used to filter a previously reported expression dataset to identify membrane-associated and secreted genes which are associated with poor prognosis in breast cancer and represent potential targets for diagnosis and treatment. The approach reported here should provide a useful tool for the analysis of gene expression patterns, identifying membrane-associated or secreted genes with biological relevance that have the potential for clinical applications in diagnosis or treatment.

Breast Neoplasms↗

Safety evaluation of chemical mixtures and combinations of chemical and non-chemical stressors.

Recent developments in hazard identification and risk assessment of chemical mixtures are reviewed. Empirical, descriptive approaches to study and characterize the toxicity of mixtures have dominated during the past two decades, but an increasing number of mechanistic approaches have made their entry into mixture toxicology. A series of empirical studies with simple chemical mixtures in rats is described in some detail because of the important lessons from this work. The development of regulatory guidelines for the toxicological evaluation of chemical mixtures is discussed briefly. Current issues in mixture toxicology include the adverse health effects of ambient air pollution; the application of such modern, sophisticated methodologies as genomics, bioinformatics, and physiologically based pharmacokinetic modeling; and databases for mixture toxicity. Finally, the state of the art of our knowledge on the potential adverse health effects of combined exposures to chemicals and non-chemical stressors (noise, heat/cold, microorganisms, immobilization, restraint, or transportation), research initiatives in these fields, and the development of an indicator for the cumulative health impact of multiple environmental exposures are discussed.

Dose-Response Relationship, Drug↗

Identification and characterization of Bombyx mori eIF5A gene through bioinformatics approaches.

As the genome of B. mori is available in GenBank and the EST database of B. mori is expanding, identification of novel genes of B. mori was conceivable by data-mining techniques and bioinformatics tools. In this study, we used the in silico cloning method to identify eukaryotic initiation factor 5A (eIF5A) gene in B. mori. With the hypusine formation, eIF5A is involved in the regulation of cell proliferation and apoptosis. Using the computer program MEGA3, we conducted a search for homologs of eIF5A among many eukaryotic species and confirmed that the eIF5A was conserved in all organisms investigated. This gene has been registered in GenBank under the accession number DQ104412.

Amino Acid Sequence↗

An integrated genetic data environment (GDE)-based LINUX interface for analysis of HIV-1 and other microbial sequences.

MOTIVATION: Sequence databases encode a wealth of information needed to develop improved vaccination and treatment strategies for the control of HIV and other important pathogens. To facilitate effective utilization of these datasets, we developed a user-friendly GDE-based LINUX interface that reduces input/output file formatting. DESIGN AND RESULTS: GDE was adapted to the Linux operating system, bioinformatics tools were integrated with microbe-specific databases, and up-to-date GDE menus were developed for several clinically important viral, bacterial and parasitic genomes. Each microbial interface was designed for local access and contains Genbank, BLAST-formatted and phylogenetic databases. AVAILABILITY: GDE-Linux is available for research purposes by direct application to the corresponding author. Application-specific menus and support files can be downloaded from (http://www.bioafrica.net).

Database Management Systems↗

Expanding the organismal scope of proteomics: cross-species protein identification by mass spectrometry and its implications.

Due to the limited applicability of conventional protein identification methods to the proteomes of organisms with unsequenced genomes, researchers have developed approaches to identify proteins using mass spectrometry and sequence similarity database searches. Both the integration of mass spectrometry with bioinformatics and genomic sequencing drive the expanding organismal scope of proteomics.

Amino Acid Sequence↗

Computational analysis of protein tyrosine phosphatases: practical guide to bioinformatics and data resources.

The exponential growth of sequence data has become a challenge to database curators and end-users alike and biologists seeking to utilize the data effectively are faced with numerous analysis methods. Here, with practical examples from our bioinformatics analysis of the protein tyrosine phosphatases (PTPs), we show how computational analysis can be exploited to fuel hypothesis-driven experimental research through the exploration of online databases. We cover the following elements: (i) similarity searches and strategies to collect a non-redundant database of tyrosine-specific PTP domains; (ii) utilization of this database to classify human, fly, and worm PTPs (based on alignments and phylogenetic analysis); (iii) three-dimensional structural analysis to identify conserved regions (structure-function) and non-conserved selectivity-determining regions (substrate specificity); and (iv) genomic analysis, including mapping of exon structure, identification of pseudogenes, and exploration of disease databases. We discuss the importance of manual curation, illustrating examples in which pseudogenes give rise to predicted proteins in GenBank and note that domain servers, such as PFAM and SMART, erroneously include dual-specificity and lipid phosphatases in their collection of tyrosine-specific PTPs. To capitalize on our annotated set of 402 PTP domains (from 47 species and five phyla), we identify sequence conservation across taxonomic categories and explore structure-function relationships among tandem domain receptor-like PTPs. We define three Src homology 2 domain-containing PTP genes in stingray, zebrafish, and fugu and speculate on their evolutionary relationship with human pseudogenes. Our annotated sequences, along with a web service for phylogenetic classification of PTP domains, are available online (http://ptp.cshl.edu and http://science.novonordisk.com/ptp).

Amino Acid Sequence↗

CSRDB: a small RNA integrated database and browser resource for cereals.

Plant small RNAs (smRNAs), which include microRNAs (miRNAs), short interfering RNAs (siRNAs) and trans-acting siRNAs (ta-siRNAs), are emerging as significant components of epigenetic processes and of gene networks involved in development and in homeostasis. Here we present a bioinformatics resource for cereal crops, the Cereal Small RNA Database (CSRDB), consisting of large-scale datasets of maize and rice smRNA sequences generated by high-throughput pyrosequencing. The smRNA sequences have been mapped to the rice genome and to the available maize genome sequence and these results are presented in two genome browser datasets using the Generic Genome Browser. Potential RNA targets for the smRNAs have been predicted and access to the resulting smRNA/RNA target pair dataset has been made available through a MySQL based relational database. Various ways to access the data are provided including links from the genome browser to the target database. Data linking and integration are the main focus for this interface, and internal as well as external links are present. The resource is available at http://sundarlab.ucdavis.edu/smrnas/ and will be updated as more sequences become available.

Databases, Nucleic Acid↗

A cSNP map and database for human chromosome 21.

Single nucleotide polymorphisms (SNPs) are likely to contribute to the study of complex genetic diseases. The genomic sequence of human chromosome 21q was recently completed with 225 annotated genes, thus permitting efficient identification and precise mapping of potential cSNPs by bioinformatics approaches. Here we present a human chromosome 21 (HC21) cSNP database and the first chromosome-specific cSNP map. Potential cSNPs were generated using three approaches: (1) Alignment of the complete HC21 genomic sequence to cognate ESTs and mRNAs. Candidate cSNPs were automatically extracted using a novel program for context-dependent SNP identification that efficiently discriminates between true variation, poor quality sequencing, and paralogous gene alignments. (2) Multiple alignment of all known HC21 genes to all other human database entries. (3) Gene-targeted cSNP discovery. To date we have identified 377 cSNPs averaging ~1 SNP per 1.5 kb of transcribed sequence, covering 65% of known genes in the chromosome. Validation of our bioinformatics approach was demonstrated by a confirmation rate of 78% for the predicted cSNPs, and in total 32% of the cSNPs in our database have been confirmed. The database is publicly available at http://csnp.unige.ch or http://csnp.isb-sib.ch. These SNPs provide a tool to study the contribution of HC21 loci to complex diseases such as bipolar affective disorder and allele-specific contributions to Down syndrome phenotypes.

Base Composition↗

trEST, trGEN and Hits: access to databases of predicted protein sequences.

High throughput genome (HTG) and expressed sequence tag (EST) sequences are currently the most abundant nucleotide sequence classes in the public database. The large volume, high degree of fragmentation and lack of gene structure annotations prevent efficient and effective searches of HTG and EST data for protein sequence homologies by standard search methods. Here, we briefly describe three newly developed resources that should make discovery of interesting genes in these sequence classes easier in the future, especially to biologists not having access to a powerful local bioinformatics environment. trEST and trGEN are regularly regenerated databases of hypothetical protein sequences predicted from EST and HTG sequences, respectively. Hits is a web-based data retrieval and analysis system providing access to precomputed matches between protein sequences (including sequences from trEST and trGEN) and patterns and profiles from Prosite and Pfam. The three resources can be accessed via the Hits home page (http://hits. isb-sib.ch).

Amino Acid Sequence↗

Company strategies for using bioinformatics.

Bioinformatics enables biotechnology companies to access and analyse their growing databases of experimental results, and to exploit public data from genome programmes and other sources. Traditionally occupying the domain of a 'guru' supplying answers to infrequent research questions, corporate bioinformatics is breaking down under the flood of data. New, more robust, professional and expandable systems will give scientists effective access to new tools. This review outlines how companies have evolved beyond the 'guru', and have organized their bioinformatics by acquiring or developing bioinformatics resources. It also describes why the biologist must be central to this process, and why this is a problem for computer professionals to solve, not for 'gurus'.

Biotechnology↗