Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Ontology-based knowledge representation for bioinformatics.

Much of biology works by applying prior knowledge ('what is known') to an unknown entity, rather than the application of a set of axioms that will elicit knowledge. In addition, the complex biological data stored in bioinformatics databases often require the addition of knowledge to specify and constrain the values held in that database. One way of capturing knowledge within bioinformatics applications and databases is the use of ontologies. An ontology is the concrete form of a conceptualisation of a community's knowledge of a domain. This paper aims to introduce the reader to the use of ontologies within bioinformatics. A description of the type of knowledge held in an ontology will be given.The paper will be illustrated throughout with examples taken from bioinformatics and molecular biology, and a survey of current biological ontologies will be presented. From this it will be seen that the use to which the ontology is put largely determines the content of the ontology. Finally, the paper will describe the process of building an ontology, introducing the reader to the techniques and methods currently in use and the open research questions in ontology development.

Artificial Intelligence↗

Evolving research trends in bioinformatics.

The cross-disciplinary nature of bioinformatics entails co-evolution with other biomedical disciplines, whereby some bioinformatics applications become popular in certain disciplines and, in turn, these disciplines influence the focus of future bioinformatics development efforts. We observe here that the growth of computational approaches within various biomedical disciplines is not merely a reflection of a general extended usage of computers and the Internet, but due to the production of useful bioinformatics databases and methods for the rest of the biomedical scientific community. We have used the abstracts stored both in the MEDLINE database of biomedical literature and in NIH-funded project grants, to quantify two effects. First, we examine the biomedical literature as a whole and find that the use of computational methods has become increasingly prevalent across biomedical disciplines over the past three decades, while use of databases and the Internet have been rapidly increasing over the past decade. Second, we study the recent trends in the use of bioinformatics topics. We observe that molecular sequence databases are a widely adopted contribution in biomedicine from the field of bioinformatics, and that microarray analysis is one of the major new topics engaged by the bioinformatics community. Via this analysis, we were able to identify areas of rapid growth in the use of informatics to aid in curriculum planning, development of computational infrastructure and strategies for workforce education and funding.

Computational Biology↗

Evolving from bioinformatics in-the-small to bioinformatics in-the-large.

We argue the significance of a fundamental shift in bioinformatics, from in-the-small to in-the-large. Adopting a large-scale perspective is a way to manage the problems endemic to the world of the small-constellations of incompatible tools for which the effort required to assemble an integrated system exceeds the perceived benefit of the integration. Where bioinformatics in-the-small is about data and tools, bioinformatics in-the-large is about metadata and dependencies. Dependencies represent the complexities of large-scale integration, including the requirements and assumptions governing the composition of tools. The popular make utility is a very effective system for defining and maintaining simple dependencies, and it offers a number of insights about the essence of bioinformatics in-the-large. Keeping an in-the-large perspective has been very useful to us in large bioinformatics projects. We give two fairly different examples, and extract lessons from them showing how it has helped. These examples both suggest the benefit of explicitly defining and managing knowledge flows and knowledge maps (which represent metadata regarding types, flows, and dependencies), and also suggest approaches for developing bioinformatics database systems. Generally, we argue that large-scale engineering principles can be successfully adapted from disciplines such as software engineering and data management, and that having an in-the-large perspective will be a key advantage in the next phase of bioinformatics development.

Computational Biology↗

The International Gene Trap Consortium Website: a portal to all publicly available gene trap cell lines in mouse.

Gene trapping is a method of generating murine embryonic stem (ES) cell lines containing insertional mutations in known and novel genes. A number of international groups have used this approach to create sizeable public cell line repositories available to the scientific community for the generation of mutant mouse strains. The major gene trapping groups worldwide have recently joined together to centralize access to all publicly available gene trap lines by developing a user-oriented Website for the International Gene Trap Consortium (IGTC). This collaboration provides an impressive public informatics resource comprising approximately 45 000 well-characterized ES cell lines which currently represent approximately 40% of known mouse genes, all freely available for the creation of knockout mice on a non-collaborative basis. To standardize annotation and provide high confidence data for gene trap lines, a rigorous identification and annotation pipeline has been developed combining genomic localization and transcript alignment of gene trap sequence tags to identify trapped loci. This information is stored in a new bioinformatics database accessible through the IGTC Website interface. The IGTC Website (www.genetrap.org) allows users to browse and search the database for trapped genes, BLAST sequences against gene trap sequence tags, and view trapped genes within biological pathways. In addition, IGTC data have been integrated into major genome browsers and bioinformatics sites to provide users with outside portals for viewing this data. The development of the IGTC Website marks a major advance by providing the research community with the data and tools necessary to effectively use public gene trap resources for the large-scale characterization of mammalian gene function.

Animals↗

Designing XML schemas for bioinformatics.

Data interchange bioinformatics databases will, in the future, most likely take place using extensible markup language (XML). The document structure will be described by an XML Schema rather than a document type definition (DTD). To ensure flexibility, the XML Schema must incorporate aspects of Object-Oriented Modeling. This impinges on the choice of the data model, which, in turn, is based on the organization of bioinformatics data by biologists. Thus, there is a need for the general bioinformatics community to be aware of the design issues relating to XML Schema. This paper, which is aimed at a general bioinformatics audience, uses examples to describe the differences between a DTD and an XML Schema and indicates how Unified Modeling Language diagrams may be used to incorporate Object-Oriented Modeling in the design of schema.

Biotechnology↗

DCCP and DICP: construction and analyses of databases for copper- and iron-chelating proteins.

Copper and iron play important roles in a variety of biological processes, especially when being chelated with proteins. The proteins involved in the metal binding, transporting and metabolism have aroused much interest. To facilitate the study on this topic, we constructed two databases (DCCP and DICP) containing the known copper- and iron-chelating proteins, which are freely available from the website http://sdbi.sdut.edu.cn/en. Users can conveniently search and browse all of the entries in the databases. Based on the two databases, bioinformatic analyses were performed, which provided some novel insights into metalloproteins.

Animals↗

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

MYH16 upregulation is associated with lung adenocarcinoma aggressiveness and immune infiltration.

Myosin heavy chain 16 (MYH16) may significantly affect cell cycle progression. Nevertheless, there is a lack of evidence about the clinical relevance of MYH16 upregulation in pan cancers, including lung adenocarcinoma (LUAD). MYH16 expression patterns were evaluated in various bioinformatics databases using The Cancer Genome Atlas data set. Clinical and pathological factor data were employed to risk-stratify patients. The Kaplan-Meier plotter approach was used to estimate survival rates. Tumor immune infiltration was explored via the TIMER tool, and gene set enrichment analysis (GSEA) was used to identify the pathways involved in MYH16 upregulation. The results showed that MYH16 was abnormally upregulated in pan cancers, including LUAD. MYH16 expression induction in LUAD was found to be related to the tumor stage. Furthermore, MYH16 upregulation was correlated with LUAD development and worse overall survival, particularly in women. Notably, MYH16 overexpression in LUAD tissues corresponded to the amount of immune infiltration in the tumor. Additionally, univariate Cox hazard regression analysis revealed that MYH16 may be an independent prognostic indicator for LUAD. Furthermore, a nomogram was constructed according to MYH16 expression and clinical characteristics. BMP6 expression deficiency may be a key factor contributing to MYH16 upregulation in LUAD. Finally, GSEA demonstrated that MYH16 might mediate meiosis and gene silencing through RNA signaling pathways. This study, for the first time, showed that MYH16 upregulation in LUAD is associated with various risk factors, increased cancer aggressiveness, enhanced infiltration of tumor immune cells, and reduced survival rates.

Female↗

Proteomic analysis of proteins expressed by Helicobacter pylori under oxidative stress.

Helicobacter pylori is a spiral, slow growing gram-negative microaerophilic bacterium. It has been shown to be the etiological agent of gastroduodenal diseases, such as chronic gastritis, gastric and duodenal ulcers, and gastric cancer. To address the influence of oxidative stress and its underlying mechanisms, we have compared proliferation, urease activity and protein expression profile of H. pylori incubated under normal microaerophilic (5% O2) and aerobic stress (20% O2) conditions. Oxidative-stress cells displayed coccoid morphology and time-dependent decrease in proliferation. The urease activity was completely abrogated after 32 h. We have further compared the protein expression profiles of H. pylori under normal growing and oxidative-stress conditions by a global proteomic analysis, which includes high-resolution 2-DE followed by MALDI-TOF-MS and bioinformatic databases search/peptide-mass comparison. The results revealed that more than ten proteins were differentially expressed under oxidative stress. Most notably, the protein expression levels of urease accessory protein E (UreE, an essential metallochaperone for urease activity) and alkylhydroperoxide reductase (AhpC) with antioxidant potential are greatly decreased under stress conditions. Measurements of messenger RNA transcription level by performing RT-PCR on total mRNA also confirmed that gene expressions for these two proteins are consistently repressed under oxygen tension. These changes form a firm basis to account for the loss of urease activity and anti-oxidative ability of H. pylori after long-term exposure to reactive oxygen. Conceivably, UreE and AhpC may thus be listed as potential targets for the development of therapeutic drugs against H. pylori.

Animals↗

Functional proteomics and correlated signaling pathway of the thermophilic bacterium Bacillus stearothermophilus TLS33 under cold-shock stress.

The thermophilic bacterium Bacillus stearothermophilus TLS33 was examined under cold-shock stress by a proteomic approach to gain a better understanding of the protein synthesis and complex regulatory pathways of bacterial adaptation. After downshift in the temperature from 65 degrees C, the optimal growth temperature for this bacterium, to 37 degrees C and 25 degrees C for 2 h, we used the high-throughput techniques of proteomic analysis combining 2-DE and MS to identify 53 individual proteins including differentially expressed proteins. The bioinformatics database was used to search the biological functions of proteins and correlate these with gene homology and metabolic pathways in cell protection and adaptation. Eight cold-shock-induced proteins were shown to have markedly different protein expression: glucosyltransferase, anti-sigma B (sigma(B)) factor, Mrp protein homolog, dihydroorthase, hypothetical transcriptional regulator in FeuA-SigW intergenic region, RibT protein, phosphoadenosine phosphosulfate reductase and prespore-specific transcriptional activator RsfA. Interestingly, six of these cold-shock-induced proteins are correlated with the signal transduction pathway of bacterial sporulation. This study aims to provide a better understanding of the functional adaptation of this bacterium to environmental cold-shock stress.

Amino Acid Sequence↗

Genomic structure and genitourinary expression of mouse cytosolic prostaglandin E(2) synthase gene.

Prostaglandin E(2) (PGE(2)) plays an important role in genitourinary function. Multiple enzymes are involved in its biosynthesis. Here we report the genomic structure and tissue-selective expression of cytosolic PGE(2) synthase (cPGES) in genitourinary tissues. Full-length mouse cPGES cDNA was cloned by reverse transcript-polymerase chain reaction (RT-PCR) and 5'- and 3'-rapid amplification of cDNA ends (RACE). Analysis of a cPGES cDNA with partially sequenced cPGES genomic clones and bioinformatic databases demonstrates that the murine cPGES gene spans approximately 22 kb and consists of eight exons. The cPGES gene promoter is GC-rich and contains many SP1 sites but lacks an obvious TATA box motif. RNase protection assay revealed constitutive expression of cPGES was greatest in the testis with lower levels in the ovary, kidney, bladder and uterus. In situ hybridization studies demonstrated that cPGES mRNA was most highly expressed in the epithelial cells of seminiferous tubules in the testis. In the female reproductive tissues, cPGES was mainly localized in ovarian primary and secondary follicles and oviductal epithelial cells with less expression in uterine endometrium. In the kidney cPGES expression was diffusely expressed. In urinary bladder, cPGES expression was restricted to the transitional epithelial cells. This expression pattern is consistent with an important role for cPGES-mediated PGE(2) in urogenital tissue function.

Amino Acid Sequence↗

Expression of mouse membrane-associated prostaglandin E2 synthase-2 (mPGES-2) along the urogenital tract.

Prostaglandin E(2) (PGE(2)) is the most common prostanoid and has a variety of bioactivities including a crucial role in urogenital function. Multiple enzymes are involved in its biosynthesis. Among 3 PGE(2) terminal synthetic enzymes, membrane-associated PGE(2) synthase-2 (mPGES-2) is the most recently identified, and its role remains uncharacterized. In previous studies, membrane-associated PGE(2) synthase-1 (mPGES-1) and cytosolic PGE(2) synthase (cPGES) were reported to be expressed along the urogenital tracts. Here we report the genomic structure and tissue distribution of mPGES-2 in the urogenital system. Analysis of several bioinformatic databases demonstrated that mouse mPGES-2 spans 7 kb and consists of 7 exons. The mPGES-2 promoter contains multiple Sp1 sites and a GC box without a TATA box motif. Real-time quantitative PCR revealed that constitutive mPGES-2 mRNA was most abundant in the heart, brain, kidney and small intestine. In the urogenital system, mPGES-2 was highly expressed in the renal cortex, followed by the renal medulla and ovary, with lower levels in the ureter, bladder and uterus. Immunohistochemistry studies indicated that mPGES-2 was ubiquitously expressed along the nephron, with much lower levels in the glomeruli. In the ureter and bladder, mPGES-2 was mainly localized to the urothelium. In the reproductive system, mPGES-2 was restricted to the epithelial cells of the testis, epididymis, vas deferens and seminal vesicle in males, and oocytes, stroma cells and corpus luteum of the ovary and epithelial cells of the oviduct and uterus in females. This expression pattern is consistent with an important role for mPGES-2-mediated PGE(2) in urogenital function.

Animals↗

Protein expression profiling of the shrimp cellular response to white spot syndrome virus infection.

To better understand the pathogenesis of white spot syndrome virus (WSSV) and to determine which cell pathways might be affected after WSSV infection, two-dimensional gel electrophoresis (2-DE) was used to produce protein expression profiles from samples taken at 48 h post-infection (hpi) from the stomachs of Litopenaeus vannamei (also called Penaeus vannamei) that were either specific pathogen free or else infected with WSSV. Seventy-five protein spots that consistently showed either a marked change (>50%) in accumulated levels or else were highly expressed throughout the course of WSSV infection were selected for further study. After in-gel trypsin digestion followed by LC-nanoESI-MS/MS, bioinformatics databases were searched for matches. A total of 53 proteins were identified, with functions that included energy production, calcium homeostasis, nucleic acid synthesis, signaling/communication, oxygen carrier/transportation, and SUMO-related modification. 2-DE results were shown to be consistent with relative EST database data from a previously developed EST database of two Penaeus monodon cDNA libraries. For seven selected genes, 2-DE and EST data were also compared with transcriptional time-course RT-PCR data. This study is the first global analysis of differentially expressed proteins in WSSV-infected shrimp, and in addition to increasing our understanding of the molecular pathogenesis of this virus-associated shrimp disease, the results presented here should be useful both for identifying potential biomarkers and for developing antiviral measures.

Animals↗

Bioinformatic analysis of neuropeptide and receptor expression profiles during midgut metamorphosis in Drosophila melanogaster.

Neuropeptides are important messenger molecules in invertebrates, serving as neuromodulators in the nervous system and as regulatory hormones released into the circulation. Understanding the function of neuropeptides will require the integration of genetic, biochemical, physiological and behavioral information. The advent of DNA microarrays and bioinformatic databases provides a wealth of data describing the expression profiles of thousands of genes during biological processes. One such array catalogs the developmental patterns of gene expression during the metamorphic transformation of the Drosophila midgut. We have mined the data from this experiment to explore changes of expression in genes coding for known neuropeptides, peptide hormones, and their receptors during the metamorphosis of the midgut. We found small but significant changes in the expression of the peptides diuretic hormone, FGLa-type allatostatins, myoinhibiting peptide, ecdysis-triggering hormone, drosokinin and the burs subunit of bursicon, as well as the receptors DAR-2, NPFR1, ALCR-2, Lkr and DH-R. Just as advances have been made in understanding the molecular basis of invertebrate neuropeptide action by analysis of genome projects, data mining of gene expression databases can help to integrate molecular, biochemical and physiological knowledge of biological processes.

Animals↗

Development of a Web site for the genetic epidemiology of endometriosis.

OBJECTIVE: Endometriosis is a complex trait, in which genetic and environmental factors act together to produce the phenotype. So far, research into candidate genes has largely been based on biological and clinical hypotheses. Results of these studies and the wealth of gene and marker sequence information from the Human Genome Project could--when brought together--provide the researcher with new etiological avenues to explore. DESIGN: Online review. SETTING: The Web site being developed draws together evidence of genetic variants associated with endometriosis and new etiological hypotheses. It incorporates links to up-to-date genomic information relevant to the candidates from a range of bioinformatics databases. PATIENT(S): Endometriosis cases and controls in association studies. INTERVENTION(S): None. MAIN OUTCOME MEASURE(S): Allele and genotype frequencies. RESULT(S): The Web site summarizes the main hypotheses for endometriosis etiology that provide the basis for the search for genes involved, together with [1] the existing evidence of associations with candidate genes, with links to the relevant publications; [2] details of these candidate genes and the surrounding chromosomal area (location, function, polymorphisms, marker maps); [3] molecular biological findings, from studies of aberrant gene and protein expression in relevant tissues; and [4] chromosomal regions that have been implicated. CONCLUSION(S): This Web site should provide a useful information tool for the endometriosis researcher. We encourage researchers worldwide to use it, contribute to it, and share their knowledge about the condition.

Alleles↗

Linking publication, gene and protein data.

The computational reconstruction of biological systems, 'systems biology', is necessarily dependent on the existence of well-annotated data sets defining and describing the components of these systems, especially genes and the proteins they encode. Information about these components can be accessed either through structured bioinformatics databases, which store basic chemical and functional information abstracted from (or supplementing) the scientific literature, or through the literature itself, which is richer in content but essentially unstructured.

Amino Acid Sequence↗

The characterization of pregnancy associated plasma protein-E and the identification of an alternative splice variant.

We have performed differential display and bioinformatic database mining of the placenta, in an attempt to find novel diagnostic markers of pathological pregnancies. We have identified a full-length cDNA encoding the preproprotein of pregnancy associated plasma protein-E (PAPP-E); a putative metalloprotease, of 1790-residues with a putative 21-residue signal peptide. An alternatively spliced mRNA was found to encode an 826-residue precursor protein corresponding to the N-terminus of PAPP-E. Both PAPP-E variants were found to be co-expressed abundantly in the placenta and non-pregnant mammary gland with low expression in the kidney, foetal brain and pancreas. Analysis of the predicted proteins suggests that the longer variant be targeted to the nucleus while the shorter variant is secreted extracellularly. Gene structure analysis revealed that PAPP-E was encoded on 23 exons on chromosome 1 and its splice variant on the first five same exons. The discovery of the PAPP-E variants will help in the deciphering of the physiology of this new family of metzincins in not only the placenta during pregnancy but also the mammary gland in breast cancer. The new PAPP-E variants could have the potential for the diagnosis of pathological pregnancies including trisomies such as Down's syndrome.

Alternative Splicing↗