Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Biosphere: the interoperation of web services in microarray cluster analysis.

UNLABELLED: The growing use of DNA microarrays in biomedical research has led to the proliferation of analysis tools. These software programs address different aspects of analysis (e.g. normalisation and clustering within and across individual arrays) as well as extended analysis methods (e.g. clustering, annotation and mining of multiple datasets). Therefore, microarray data analysis typically requires the interoperability of multiple software programs involving different analysis types and methods. Such interoperation is often hampered by the heterogeneity inherent in the software tools (which may function by implementing different interfaces and using different programming languages). To address this problem, we employed the simple object access protocol (SOAP)-based web service approach that provides a uniform programmatic interface to these heterogeneous software components. To demonstrate this approach in the microarray context, we created a web server application, Biosphere, which interoperates a number of web services that are geographically widely distributed. These web services include a clustering web service, which is a suite of different clustering algorithms for analysing microarray data; XEMBL, developed at the European Bioinformatics Institute (EBI) for retrieving EMBL Nucleotide Sequence Database sequence data; and three gene annotation web services: GetGO, GetHAPI and GetUMLS. GetGO allows retrieval of Gene Ontology (GO) annotation, and the other two web services retrieve annotation from the biomedical literature that is indexed based on the Medical Subject Headings (MeSH) terms. With these web services, Biosphere allows the users to do the following: (i) cluster gene expression data using seven different algorithms; (ii) visualise the clustering results that are grouped statistically in colour; and (iii) retrieve sequence, annotation and citation data for the genes of interest. AVAILABILITY: Biosphere and its web services described in Web Service Description Language (WSDL) can be accessed at http://rook.cecid.hku.hk:8280/BiosphereServer.

Cluster Analysis↗

[Cloning full-length homologous cDNAs of pollen allergens in Humulus Scandens(Lour.) Merr by degenerate primer].

AIM: To establish a stable and reliable method for fast cloning homologous genes of pollen allergens in allergen-containing plants. METHODS: Degenerate primers were designed based on the bioinformatic analysis of numerous allergens available from the database. Subsequent amplification of the allergen genes was conducted in the weed pollen cDNA pool by a selective PCR profile. Following the truncated gene cloning, RACE method was used to isolate full-length cDNA. Gene function was deduced by sequence alignment in GenBank database. The degenerate ability of the primer was compared with the full-length cDNA sequences. RESULTS: Three full-length cDNAs were obtained. Sequence analysis showed that these new genes shared as high as 79%-85% homology with a large amount of known allergen profilins and were hence regarded as members of panallergen profilin family. Comparing these genes with the degenerate primers that were initially used in truncated gene cloning revealed that alternative nucleotide degeneracy occurred beyond the degenerate site predesigned, suggesting that further degeneracy was expanded by Touchdown-gradient PCR. CONCLUSION: Cloning of homologous genes or allergen genes can be efficiently achieved by using the combination of degenerate primer with Touchdown-gradient RT-PCR in the species such as Humulus scandens that has not yet been investigated.

Allergens↗

[Introduction to genome databases].

A brief introduction to the genome databases GDB, GenoList and Ensembl is given. These databases, mirrored and maintained at the Centre of Bioinformatics, Peking University, provide useful information for genome research.

English Abstract↗

Molecular evolution of enolase.

Enolase (EC 4.2.1.11) is an enzyme of the glycolytic pathway catalyzing the dehydratation reaction of 2-phosphoglycerate. In vertebrates the enzyme exists in three isoforms: alpha, beta and gamma. The amino-acid and nucleotide sequences deposited in the GenBank and SwissProt databases were subjected to analysis using the following bioinformatic programs: ClustalX, GeneDoc, MEGA2 and S.I.F.T. (sort intolerant from tolerant). Phylogenetic trees of enolases created with the use of the MEGA2 program show evolutionary relationships and functional diversity of the three isoforms of enolase in vertebrates. On the basis of calculations and the phylogenetic trees it can be concluded that vertebrate enolase has evolved according to the "birth and death" model of evolution. An analysis of amino acid sequences of enolases: non-neuronal (NNE), neuron specific (NSE) and muscle specific (MSE) using the S.I.F.T. program indicated non-uniform number of possible substitutions. Tolerated substitutions occur most frequently in alpha-enolase, while the lowest number of substitutions has accumulated in gamma-enolase, which may suggest that it is the most recently evolved isoenzyme of enolase in vertebrates.

Animals↗

The RESID Database of Protein Modifications as a resource and annotation tool.

The RESID Database of Protein Modifications is a comprehensive collection of annotations and structures for protein modifications and cross-links including pre-, co-, and post-translational modifications. The database provides: systematic and alternate names, atomic formulas and masses, enzymatic activities that generate the modifications, keywords, literature citations, Gene Ontology (GO) cross-references, protein sequence database feature table annotations, structure diagrams, and molecular models. This database is freely accessible on the Internet through resources provided by the European Bioinformatics Institute (http://www.ebi.ac.uk/RESID), and by the National Cancer Institute--Frederick Advanced Biomedical Computing Center (http://www.ncifcrf.gov/RESID). Each RESID Database entry presents a chemically unique modification and shows how that modification is currently annotated in the protein sequence databases, Swiss-Prot and the Protein Information Resource (PIR). The RESID Database provides a table of corresponding equivalent feature annotations that is used in the UniProt project, an international effort to combine the resources of the Swiss-Prot, TrEMBL and PIR. As an annotation tool, the RESID Database is used in standardizing and enhancing modification descriptions in the feature tables of Swiss-Prot entries. As an Internet resource, the RESID Database assists researchers in high-throughput proteomics to search monoisotopic masses and mass differences and identify known and predicted protein modifications.

Databases, Factual↗

Current status of the Asthma and Allergy Database.

The database provides an online resource for access to data on the genetics of asthma and allergy. This report describes the present status of the site. Currently, a detailed description of 88 linkage studies (7164 linkage positions) and 72 mutation studies are available. The results can be accessed in table form or graphically. The database also contains mouse asthma studies and human homology relationships, gene expression studies and links to relevant patents. Technical details about the server architecture, database installation, database construction, database structure and the user interface are explained elsewhere [Wjst and Immervoll (1998) Bioinformatics, 14, 827-828]. The URL is http://cooke.gsf.de

Animals↗

NemaFootPrinter: a web based software for the identification of conserved non-coding genome sequence regions between C. elegans and C. briggsae.

BACKGROUND: NemaFootPrinter (Nematode Transcription Factor Scan Through Philogenetic Footprinting) is a web-based software for interactive identification of conserved, non-exonic DNA segments in the genomes of C. elegans and C. briggsae. It has been implemented according to the following project specifications:a) Automated identification of orthologous gene pairs. b) Interactive selection of the boundaries of the genes to be compared. c) Pairwise sequence comparison with a range of different methods. d) Identification of putative transcription factor binding sites on conserved, non-exonic DNA segments. RESULTS: Starting from a C. elegans or C. briggsae gene name or identifier, the software identifies the putative ortholog (if any), based on information derived from public nematode genome annotation databases. The investigator can then retrieve the genome DNA sequences of the two orthologous genes; visualize graphically the genes' intron/exon structure and the surrounding DNA regions; select, through an interactive graphical user interface, subsequences of the two gene regions. Using a bioinformatics toolbox (Blast2seq, Dotmatcher, Ssearch and connection to the rVista database) the investigator is able at the end of the procedure to identify and analyze significant sequences similarities, detecting the presence of transcription factor binding sites corresponding to the conserved segments. The software automatically masks exons. DISCUSSION: This software is intended as a practical and intuitive tool for the researchers interested in the identification of non-exonic conserved sequence segments between C. elegans and C. briggsae. These sequences may contain regulatory transcriptional elements since they are conserved between two related, but rapidly evolving genomes. This software also highlights the power of genome annotation databases when they are conceived as an open resource and the possibilities offered by seamless integration of different web services via the http protocol. AVAILABILITY: The program is freely available at http://bio.ifom-firc.it/NTFootPrinter.

Animals↗

A global representation of the carbohydrate structures: a tool for the analysis of glycan.

Glycan resources have been developed of late, such as carbohydrate databases, analysis tools, and algorithms for analysis of carbohydrate features. With this background, bioinformatics approaches to carbohydrate research have recently begun using a large amount of protein and carbohydrate data. This paper introduces one of these projects that elucidates the range of carbohydrate structures. In this study, the variety of carbohydrate structures have been enumerated in a global tree structure called variation trees, using the KEGG GLYCAN database, which is a public-domain glycan resource for bioinformatics analysis. Additionally, a glycosyltransferase mapping list of glycosyltransferases and their catalyzing glycosidic linkages was constructed. From this, we present the composite structure map (CSM), which is a structural variation map integrating its variation trees and glycosyltransferase map list. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner, illustrating its versatility as a new bioinformatics resource and tool capable of analyzing carbohydrate structures on a global scale. These resources are available at http://www.genome.jp/kegg/glycan/.

Carbohydrate Conformation↗

A combined in vitro/bioinformatic investigation of redox regulatory mechanisms governing cell cycle progression.

The intracellular reduction-oxidation (redox) environment influences cell cycle progression; however, underlying mechanisms are poorly understood. To examine potential mechanisms, the intracellular redox environment was characterized per cell cycle phase in Chinese hamster ovary fibroblasts via flow cytometry by measuring reduced glutathione (GSH), reactive oxygen species (ROS), and DNA content with monochlorobimane, 2',7'-dichlorohydrofluorescein diacetate (H2DCFDA), and DRAQ5, respectively. GSH content was significantly greater in G2/M compared with G1 phase cells, whereas GSH was intermediate in S phase cells. ROS content was similar among phases. Together, these data demonstrate that G2/M cells are more reduced than G1 cells. Conventional approaches to define regulatory mechanisms are subjective in nature and focus on single proteins/pathways. Proteome databases provide a means to overcome these inherent limitations. Therefore, a novel bioinformatic approach was developed to exhaustively identify putative redox-regulated cell cycle proteins containing redox-sensitive protein motifs. Using the InterPro (http://www.ebi.ac.uk/interpro/) database, we categorized 536 redox-sensitive motifs as: 1) active/functional-site cysteines, 2) electron transport, 3) heme, 4) iron binding, 5) zinc binding, 6) metal binding (non-Fe/Zn), and 7) disulfides. Comparing this list with 1,634 cell cycle-associated proteins from Swiss-Prot and SpTrEMBL (http://us.expasy.org/sprot/) revealed 92 candidate proteins. Three-fourths (69 of 92) of the candidate proteins function in the central cell cycle processes of transcription, nucleotide metabolism, (de)phosphorylation, and (de)ubiquitinylation. The majority of oxidant-sensitive candidate proteins (68.9%) function during G2/M phase. As the G2/M phase is more reduced than the G1 phase, oxidant-sensitive proteins may be temporally regulated by oscillation of the intracellular redox environment. Combined with evidence of intracellular redox compartmentalization, we propose a spatiotemporal mechanism that functionally links an oscillating intracellular redox environment with cell cycle progression.

Amino Acid Motifs↗

Screening and identification of key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods.

OBJECTIVE: Radiotherapy (RT) plays a crucial role in the comprehensive treatment of rectal cancer. However, the impact of radiotherapy on the tumor microenvironment (TME), especially its effect on immune cell infiltration and immune-related gene expression, has not been fully studied. This study aims to screen and analyze key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods for the purpose of identifying potential biomarkers and providing new insights for the personalized therapy of rectal cancer. METHODS: Using data from the Public Gene Expression Database (GEO) and the Cancer Genomics Database (TCGA), the impact of radiotherapy on the immune microenvironment of rectal cancer was explored using bioinformatics tools. Through screening differentially expressed genes (DEGs), correlation analysis, TIMER database analysis, immune infiltration score, and correlation analysis between key genes and prognosis, the effects of radiotherapy on the immune microenvironment of rectal cancer were investigated. RESULTS: Totally 7 upregulated and 4 downregulated differentially expressed genes were identified, among which MASP1, LTK, SLC9A3R2 were negatively correlated with myeloid suppressor cell infiltration (MDSCs), while ZP2 was positively correlated. The expression of MASP1 and SLC9A3R2 was closely related to the level of immune cell infiltration and played significant roles in the immune microenvironment. High expression of MASP1 was significantly correlated with survival benefits from immune checkpoint inhibitor therapy, while SLC9A3R2 was closely related to the efficacy of PD-L1 inhibitors and CTLA4 inhibitors. CONCLUSIONS: MASP1 and SLC9A3R2, as two key genes that may be related to the immune microenvironment of rectal cancer radiotherapy, deserve further exploration of their roles in the mechanism. The combination of radiotherapy and immunotherapy holds promising prospects in the treatment of rectal cancer, and exploration of related mechanisms will provide new strategies and targets for the treatment of various tumors and rectal cancer.

Bioinformatics↗

Lack of cross-reactivity between the Bacillus thuringiensis derived protein Cry1F in maize grain and dust mite Der p7 protein with human sera positive for Der p7-IgE.

Cry1F protein, derived from Bacillus thuringiensis, is effective at controlling lepidopteran pests and a synthetic Cry1F transgene was transferred into maize. For the safety assessment of genetically modified food crops, the allergenic potential of the introduced novel trait(s) is evaluated. Because no single parameter is currently predictive of allergic potential, a 'weight of evidence' approach has been proposed. As part of this assessment, the amino acid (aa) sequence of the Cry1F protein was compared to a database of known allergens using recommended criteria. The Cry1F protein did not show significant similarity or a match of eight contiguous identical aa with any allergen. However, a single six contiguous aa match was identified between Cry1F and the Der p7 protein of the dust mite, Dermatophagoides pteronyssinus. To investigate whether Cry1F was cross-reactive with Der p7, sera from 10 dust mite allergic patients containing Der p 7-specific IgE antibody were used to compare IgE-specific binding. No evidence of cross-reactivity was observed between Cry1F and Der p7. This study provides in vitro IgE sera screening data, that when considered in the context of other bioinformatic data [Hileman R.E., Silvanovich, A., Goodman R.E., Rice E.A., Holleschak G., Astwood J.D., Hefle S.L., 2002. Bioinformatic methods for allergenicity assessment using a comprehensive allergen database. Int. Arch. Allergy Immunol. 128, 280-291; Stadler, M.B., Stadler, B.M., 2003. Allergenicity prediction by protein sequence. FASEB J. 17, 1141-1143.], adds further evidence arguing against the use of a six contiguous identical amino acid search to identify potential cross-reactive allergens. Cry1F is heat labile, rapidly hydrolyzed in an in vitro pepsin resistance assay, not glycosylated and not from an allergenic source. Taken together, these data indicate a lack of allergenic concern for Cry1F.

Allergens↗

Molecular profiling techniques and bioinformatics in cancer research.

AIMS: Our aim was to describe the commonly used molecular profiling techniques in cancer research, to examine their limitations and to discuss the challenges of bioinformatics. METHODS: A literature search was performed using the PubMed database to identify publications relevant to this review. Citations from these articles were also examined to yield further relevant publications. RESULTS: We describe the use of DNA microarrays, comparative genomic hybridisation, tissue microarrays and digital differential display. The limitations of these technologies, their contribution to cancer research and the challenges of bioinformatics are also discussed. CONCLUSIONS: Although these high throughput technologies each have their own limitations they are rapidly developing and contributing significantly to our understanding of cancer genetics. They have also led to the emergence of bioinformatics as a rapidly developing and vital field.

Biomarkers, Tumor↗

The RESID Database of Protein Modifications: 2003 developments.

The RESID Database is a comprehensive collection of annotations and structures for protein pre-, co- and post-translational modifications including amino-terminal, carboxyl-terminal and peptide chain cross-link modifications. The RESID Database includes: systematic and alternate names, atomic formulas and masses, enzyme activities generating the modifications, keywords, literature citations, Gene Ontology cross-references, Protein Information Resource (PIR) and SWISS-PROT protein sequence database feature table annotations, structure diagrams and molecular models. This database is freely accessible on the Internet through the European Bioinformatics Institute at http://srs.ebi.ac.uk/srs6bin/cgi-bin/wgetz?-page+LibInfo+-lib+RESID, through the National Cancer Institute - Frederick Advanced Biomedical Computing Center at http://www.ncifcrf.gov/RESID, or through the Protein Information Resource at http://pir.georgetown.edu/pirwww/dbinfo/resid.html.

Animals↗

Munich information center for protein sequences plant genome resources: a framework for integrative and comparative analyses 1(W).

With several plant genomes sequenced, the power of comparative genome analysis can now be applied. However, genome-scale cross-species analyses are limited by the effort for data integration. To develop an integrated cross-species plant genome resource, we maintain comprehensive databases for model plant genomes, including Arabidopsis (Arabidopsis thaliana), maize (Zea mays), Medicago truncatula, and rice (Oryza sativa). Integration of data and resources is emphasized, both in house as well as with external partners and databases. Manual curation and state-of-the-art bioinformatic analysis are combined to achieve quality data. Easy access to the data is provided through Web interfaces and visualization tools, bulk downloads, and Web services for application-level access. This allows a consistent view of the model plant genomes for comparative and evolutionary studies, the transfer of knowledge between species, and the integration with functional genomics data.

Computational Biology↗

The iProClass integrated database for protein functional analysis.

Increasingly, scientists have begun to tackle gene functions and other complex regulatory processes by studying organisms at the global scales for various levels of biological organization, ranging from genomes to metabolomes and physiomes. Meanwhile, new bioinformatics methods have been developed for inferring protein function using associative analysis of functional properties to complement the traditional sequence homology-based methods. To fully exploit the value of the high-throughput system biology data and to facilitate protein functional studies requires bioinformatics infrastructures that support both data integration and associative analysis. The iProClass database, designed to serve as a framework for data integration in a distributed networking environment, provides comprehensive descriptions of all proteins, with rich links to over 50 databases of protein family, function, pathway, interaction, modification, structure, genome, ontology, literature, and taxonomy. In particular, the database is organized with PIRSF family classification and maps to other family, function, and structure classification schemes. Coupled with the underlying taxonomic information for complete genomes, the iProClass system (http://pir.georgetown.edu/iproclass/) supports associative studies of protein family, domain, function, and structure. A case study of the phosphoglycerate mutases illustrates a systematic approach for protein family and phylogenetic analysis. Such studies may serve as a basis for further analysis of protein functional evolution, and its relationship to the co-evolution of metabolic pathways, cellular networks, and organisms.

Amino Acid Sequence↗