Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

The European Bioinformatics Institute's data resources: towards systems biology.

Genomic and post-genomic biological research has provided fine-grain insights into the molecular processes of life, but also threatens to drown biomedical researchers in data. Moreover, as new high-throughput technologies are developed, the types of data that are gathered en masse are diversifying. The need to collect, store and curate all this information in ways that allow its efficient retrieval and exploitation is greater than ever. The European Bioinformatics Institute's (EBI's) databases and tools have evolved to meet the changing needs of molecular biologists: since we last wrote about our services in the 2003 issue of Nucleic Acids Research, we have launched new databases covering protein-protein interactions (IntAct), pathways (Reactome) and small molecules (ChEBI). Our existing core databases have continued to evolve to meet the changing needs of biomedical researchers, and we have developed new data-access tools that help biologists to move intuitively through the different data types, thereby helping them to put the parts together to understand biology at the systems level. The EBI's data resources are all available on our website at http://www.ebi.ac.uk.

Computational Biology↗

Combining multiplex reverse transcription-PCR and a diagnostic microarray to detect and differentiate enterovirus 71 and coxsackievirus A16.

Cluster A enteroviruses, including enterovirus 71 (EV71) and coxsackievirus A16 (CA16), are known to cause hand-foot-and-mouth disease (HFMD). Despite the close genetic relationship between these two viruses, EV71 is generally known to be a more perpetuating pathogen involved in severe clinical manifestations and deaths. While the serotyping of enteroviruses is mostly done by conventional immunological methods, many clinical isolates remain unclassifiable due to the limited number of antibodies against enterovirus surface proteins. Array-based assays are able to detect several serotypes with high accuracy. We combined an enterovirus microarray with multiplex reverse transcription-PCR to try to develop a method of sensitively and accurately detecting and differentiating EV71 and CA16. In an effort to design serotype-specific probes for detection of the virus, we first did an elaborate bioinformatic analysis of the sequence database derived from different enterovirus serotypes. We then constructed a microarray using 60-mer degenerate oligonucleotide probes covalently bound to array slides. Using this enterovirus microarray to study 144 clinical specimens from patients infected with HFMD or suspected to have HFMD, we found that it had a diagnostic accuracy of 92.0% for EV71 and 95.8% for CA16. Diagnostic accuracy for other enteroviruses (non-EV71 or -CA16) was 92.0%. All specimens were analyzed in parallel by real-time PCR and subsequently confirmed by neutralization tests. This highly sensitive array-based assay may become a useful alternative in clinical diagnostics of EV71 and CA16.

DNA Probes↗

tacg--a grep for DNA.

BACKGROUND: Pattern matching is the core of bioinformatics; it is used in database searching, restriction enzyme mapping, and finding open reading frames. It is done repeatedly over increasingly long sequences, thus codes must be efficient and insensitive to sequence length. Such patterns of interest include simple motifs with IUPAC degeneracies, regular expressions, patterns allowing mismatches, and probability matrices. RESULTS: I describe a small application which allows searching for all the above pattern types individually, which further allows these atomic motifs to be assembled into logical rules for more sophisticated analysis. CONCLUSION: tacg is small, portable, faster and more capable than most alternatives, relatively easy to modify, and freely available in source code.

Algorithms↗

Bovine transcriptome analysis by SAGE technology during an experimental Trypanosoma congolense infection.

In central and sub-Saharan Africa, trypanosomosis is a tsetse fly-transmitted disease, which is considered as the most important impediment to livestock production in the region. However, several indigenous West African taurine breeds (Bos taurus) present remarkable tolerance to the infection. This genetic capability, named trypanotolerance, results from numerous biological mechanisms most probably under multigenic dependences, among which are control of the trypanosome infection by limitation of parasitemia and control of severe anemia due to the pathogenic effects. Today, some postgenomic biotechnologies, such as transcriptome analyses, allow characterization of the full expressed genes involved in the majority of animal diseases under genetic control. One of them is serial analysis of gene expression (SAGE) technology, which consists of the construction of mRNA transcript libraries for qualitative and quantitative analysis of the entire genes expressed or inactivated at a particular step of cellular activation. We developed four different mRNA transcript libraries from white blood cells on a N'Dama trypanotolerant animal during an experimental Trypanosoma congolense (T. congolense) infection: one before experimental infection (ND0), one at the parasitemia peak (NDm), one at the minimal packed cell volume (NDa), and the last one at the end of the experiment after normalization (NDf). Bioinformatic comparisons in bovine genomic databases allowed us to obtain more than 75,000 sequences, among which are several known genes, some others are already described as expressed sequence tags (ESTs), and the last are completely new, but probably functional in trypanotolerance. The knowledge of all identified named or unnamed genes involved in trypanotolerance characteristics will allow us to use them in a field marker-assisted selections strategy and in microarrays prediction sets for bovine trypanotolerance.

Animals↗

Regulation of N-myristoyltransferase by novel inhibitor proteins.

N-myristoylation ensures the proper function and intracellular trafficking of proteins. Many proteins involved in a wide variety of signaling, including cellular transformation and oncogenesis, are myristoylated. The myristoylation of proteins is catalyzed by the ubiquitously distributed eukaryotic enzyme N-myristoyltransferase (NMT). Previously, we reported that NMT activity is higher in colonic epithelial neoplasms than in normal-appearing colonic tissue and that the increase in NMT activity appears at an early stage in colonic carcinogenesis. Furthermore, we observed that NMT expression is elevated in colorectal and gallbladder carcinoma. In our laboratory, an endogenous NMT inhibitor protein (NIP71) was discovered from bovine brain that inhibited NMT activity in rat colonic tumors. Very recently we have demonstrated that the protein NIP71, which is a potential inhibitor of NMT, is homologous to heat-shock cognate protein (HSC70). In addition, we have discovered that enolase is a potent inhibitor of NMT. Further work may elucidate the role of HSC70 and/or enolase in the regulation of NMT, which may lead to the development of a gene-based therapy of colorectal cancer. The interaction of oncoproteomic and oncogenomic data sets through powerful bioinformatics will yield a comprehensive database of protein properties, which will serve as an invaluable tool for cancer researchers to understand the progress of tumorigenesis.

Acyltransferases↗

Molecular Cloning of Human beta-arrestin1 cDNA, Expression and Functional Study.

Complete coding sequences of beta-arrestin1 (1A and 1B) were cloned through application of bioinformatics analysis to the dbEST database. beta-arrestin1A was overexpressed in E.coli with partial expression products as inclusion body. Anti-beta-arrestin1 antibodies were prepared by using purified inclusion body. Results also demonstrate that activation of inhibitory G protein mediated by delta and kappa pioid receptors was strongly attenuated by overexpression of beta-arrestin1A in co-transfected 293 cells.

Journal Article↗

Identification through bioinformatics of cDNAs encoding human thymic shared Ag-1/stem cell Ag-2. A new member of the human Ly-6 family.

The Ly-6 family of cell surface molecules includes many members that have been characterized in the mouse. Until recently, very few Ly-6 family members had been described in the human. A significant development with important implications for novel gene discovery has been the growth of the public Expressed Sequence Tag (EST) database. Here we report that, through the application of bioinformatics analysis to the dbEST database, we obtained the sequence of human TSA-1/SCA-2, a new member of the human Ly-6 family. In addition, we identified full-length clones encoding this molecule as well as expression data in various tissues. Sequencing of the clones identified this way confirmed the sequence predicted through bioinformatics. This study constitutes an example of the application of bioinformatics to the analysis of the recently expanded databases for the identification of genes of potential importance in the immune system.

Amino Acid Sequence↗

Clinical bioinformatics.

Clinical bioinformatics provides biological and medical information to allow for individualized healthcare. In this review, we describe the uses of clinical bioinformatics. After the analysis of the complete human genome sequences, clinical bioinformatics enables researchers to search online biological databases and use the biological information in their medical practices. The data obtained from using microarray is extremely complicated. In clinical bioinformatics, selecting appropriate software to analyze the microarray data for medical decision making is crucial. Proteomics strategy tools usually focus on similarity searches, structure prediction, and protein modeling. In clinical bioinformatics, the proteomic data only have meaning if they are integrated with clinical data. In pharmacogenomics, clinical bioinformatics includes elaborate studies of bioinformatics tools and various facets of proteomics related to drug target identification and clinical validation. Using clinical bioinformatics, researchers apply computational and high-throughput experimental techniques to cancer research and systems biology. Meanwhile, researchers of bioinformatics and medical information have incorporated clinical bioinformatics to improve health care, using biological and medical information. Using the high volume of biological information from clinical bioinformatics will contribute to changes in practice standards in the healthcare system. We believe that clinical bioinformatics provides benefits of improving healthcare, disease prevention and health maintenance as we move toward the era of personalized medicine.

Computational Biology↗

Evolving strategies for the incorporation of bioinformatics within the undergraduate cell biology curriculum.

Recent advances in genomics and structural biology have resulted in an unprecedented increase in biological data available from Internet-accessible databases. In order to help students effectively use this vast repository of information, undergraduate biology students at Drake University were introduced to bioinformatics software and databases in three courses, beginning with an introductory course in cell biology. The exercises and projects that were used to help students develop literacy in bioinformatics are described. In a recently offered course in bioinformatics, students developed their own simple sequence analysis tool using the Perl programming language. These experiences are described from the point of view of the instructor as well as the students. A preliminary assessment has been made of the degree to which students had developed a working knowledge of bioinformatics concepts and methods. Finally, some conclusions have been drawn from these courses that may be helpful to instructors wishing to introduce bioinformatics within the undergraduate biology curriculum.

Biology↗

Bioinformatics-based discovery of a novel factor with apparent specificity to colon cancer.

In a previous study, a data mining tool called Digital Differential Display (DDD) from the Cancer Genome Anatomy Project (CGAP) was used to predict solid tumor- and organ-specific genes from the expressed sequence tag (EST) database. To validate the use of bioinformatics approaches in gene discovery, one of the ESTs, which was predicted to be colon tumor-specific, was chosen for further study. Reverse Transcriptase-Polymerase Chain Reaction (RT-PCR) analysis of matched sets of cDNAs from normal and colon tumor tissues indicated that the EST was specifically expressed in the majority of colon tumors. Expression was also detected in early adenomas. Among other normal tissues, EST expression was detected only in the small intestine. The colon tumor specificity of this EST was inferred from the lack of expression in carcinomas of the breast, lung, ovary, pancreas and prostate. To validate the computational prediction of specificity, a full-length cDNA encompassing the entire open reading frame was cloned and, in view of its apparent specificity to the colon tumors, this gene was termed Colon Carcinoma Related Gene (CCRG). CCRG encodes a novel cysteine-rich motif and a putative signal peptide sequence. Supernatant from COS cells transfected with the CCRG expression vector stimulated proliferation of colon cancer cells. Immunoreactive CCRG was also detected in the paraffin sections of colon tumor samples. CCRG belongs to a new class of growth factors and may be important in the diagnosis and treatment of colon cancers. Identification of CCRG using bioinformatics approaches validates gene discovery using computational approaches.

Amino Acid Sequence↗

A strategy for database interoperation.

To realize the full potential of biological databases (DBs) requires more than the interactive, hypertext flavor of database interoperation that is now so popular in the bioinformatics community. Interoperation based on declarative queries to multiple network-accessible databases will support analyses and investigations that are orders of magnitude faster and more powerful than what can be accomplished through interactive navigation. I present a vision of the capabilities that a query-based interoperation infrastructure should provide, and identify assumptions underlying, and requirements of, this vision. I then propose an architecture for query-based interoperation that includes a number of novel components of an information infrastructure for molecular biology. These components include a knowledge base that describes relationships among the conceptualizations used in different biological databases, a module that can determine the DBs that are relevant to a particular query, a module that can translate a query and its results from one conceptualization to another, a collection of DB drivers that provide uniform physical access to different database management systems, a suite of translators that can interconvert among different database schema languages, and a database that describes the network location and access methods for biological databases. A number of the components are translators that bridge the heterogeneities that exist between biological DBs at several different levels, including the conceptual level, the data model, the query language, and data formats.

Artificial Intelligence↗

A Web-based data warehouse on gene expression in human malignant melanoma.

The identification of melanoma-specific dysregulated genes could identify new molecular markers. By applying bioinformatic tools for screening of biomedical databases, a melanoma-specific gene expression profile "data warehouse" was constructed. Utilizable data sets of global gene expression analyses were available from nine studies that applied different technology platforms. A single study used cell lines, five investigations analyzed cell lines and tissues obtained from patients, two studies used exclusively specimens obtained from patients, and one study analyzed blood cells prepared from patients. The total number of investigated patients was 116. From 815 differential-regulated genes, 772 (95%) were identified merely in a single study, 37 in at least two studies, five (RAB33A, ERBB3, ADRB2, MERTK, SNF1LK, and ITPKB) in at least three studies, and a single gene, RAB33A, in four studies. These data show that the accuracy, reproducibility, and comparability among different gene expression profile studies are low in melanoma. In conclusion, the study demonstrates the high diversity of gene expression profiles associated with melanoma, the necessity to include a sufficient number of samples regarding clinical standards, for the design of standardized sample collecting and preparation, for the development of common standards for microarray data processing, and for developing standardized bioinformatic tools.

Cell Line↗

Current Comparative Table (CCT) automates customized searches of dynamic biological databases.

The Current Comparative Table (CCT) software program enables working biologists to automate customized bioinformatics searches, typically of remote sequence or HMM (hidden Markov model) databases. CCT currently supports BLAST, hmmpfam and other programs useful for gene and ortholog identification. The software is web based, has a BioPerl core and can be used remotely via a browser or locally on Mac OS X or Linux machines. CCT is particularly useful to scientists who study large sets of molecules in today's evolving information landscape because it color-codes all result files by age and highlights even tiny changes in sequence or annotation. By empowering non-bioinformaticians to automate custom searches and examine current results in context at a glance, CCT allows a remote database submission in the evening to influence the next morning's bench experiment. A demonstration of CCT is available at http://orb.public.stolaf.edu/CCTdemo and the open source software is freely available from http://sourceforge.net/projects/orb-cct.

Computational Biology↗

Bioinformatics for medical diagnostics: assessment of microarray data in the context of clinical databases.

MOTIVATION: To identify genes suitable for medical diagnostics microarray data is assessed in the context of clinical databases, which store complex information about the patient phenotype. The wealth of data and lacking standards make it difficult to analyse this kind of data. RESULTS: We present a workflow for exploratory analysis of microarray data together with clinical data consisting of four steps: definition of clinically meaningful research questions in a masterfile, generation of analysis files, selection and characterization of differentially expressed genes, and estimation of classification accuracy. We applied this workflow to large data sets from the field of cardiology and oncology (n~500 patients). Systematic data management of microarray data and clinical data helps to make results more transparent and comparable.

Cardiology↗

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics↗

PCOGR: phylogenetic COG ranking as an online tool to judge the specificity of COGs with respect to freely definable groups of organisms.

BACKGROUND: The rapidly increasing number of completely sequenced genomes led to the establishment of the COG-database which, based on sequence homologies, assigns similar proteins from different organisms to clusters of orthologous groups (COGs). There are several bioinformatic studies that made use of this database to determine (hyper)thermophile-specific proteins by searching for COGs containing (almost) exclusively proteins from (hyper)thermophilic genomes. However, public software to perform individually definable group-specific searches is not available. RESULTS: The tool described here exactly fills this gap. The software is accessible at http://www.uni-wh.de/pcogr and is linked to the COG-database. The user can freely define two groups of organisms by selecting for each of the (current) 66 organisms to belong either to groupA, to the reference groupB or to be ignored by the algorithm. Then, for all COGs a specificity index is calculated with respect to the specificity to groupA, i. e. high scoring COGs contain proteins from the most of groupA organisms while proteins from the most organisms assigned to groupB are absent. In addition to ranking all COGs according to the user defined specificity criteria, a graphical visualization shows the distribution of all COGs by displaying their abundance as a function of their specificity indexes. CONCLUSIONS: This software allows detecting COGs specific to a predefined group of organisms. All COGs are ranked in the order of their specificity and a graphical visualization allows recognizing (i) the presence and abundance of such COGs and (ii) the phylogenetic relationship between groupA- and groupB-organisms. The software also allows detecting putative protein-protein interactions, novel enzymes involved in only partially known biochemical pathways, and alternate enzymes originated by convergent evolution.

Escherichia coli↗

Individual metabolism should guide agriculture toward foods for improved health and nutrition.

Genomics and bioinformatics have the vast potential to identify genes that cause disease by investigating whole-genome databases. Comparison of an individual's geno-type with a genomic database will allow the prescription of drugs to be tailored to an individual's genotype. This same bioinformatic approach, applied to the study of human metabolites, has the potential to identify and validate targets to improve personalized nutritional health and thus serve to define the added value for the next generation of foods and crops. Advances in high-throughput analytic chemistry and computing technologies make the creation of a vast database of metabolites possible for several subsets of metabolites, including lipids and organic acids. In creating integrative databases of metabolites for bioinformatic investigation, the current concept of measuring single biomarkers must be expanded to 3 dimensions to 1) include a highly comprehensive set of metabolite measurements (a profile) by multiparallel analyses, 2) measure the metabolic profile of individuals over time rather than simply in the fasted state, and 3) integrate these metabolic profiles with genomic, expression, and proteomic databases. Application of the knowledge of individual metabolism will revolutionize the ability of nutrition to deliver health benefits through food in the same way that knowledge of genomics will revolutionize individual treatment of dis-ease with pharmaceuticals.

Biomarkers↗

The PEDANT genome database in 2005.

The PEDANT genome database (http://pedant.gsf.de) contains pre-computed bioinformatics analyses of publicly available genomes. Its main mission is to provide robust automatic annotation of the vast majority of amino acid sequences, which have not been subjected to in-depth manual curation by human experts in high-quality protein sequence databases. By design PEDANT annotation is genome-oriented, making it possible to explore genomic context of gene products, and evaluate functional and structural content of genomes using a category-based query mechanism. At present, the PEDANT database contains exhaustive annotation of over 1,240,000 proteins from 270 eubacterial, 23 archeal and 41 eukaryotic genomes.

Computational Biology↗