Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Computer identification of snoRNA genes using a Mammalian Orthologous Intron Database.

Based on comparative genomics, we created a bioinformatic package for computer prediction of small nucleolar RNA (snoRNA) genes in mammalian introns. The core of our approach was the use of the Mammalian Orthologous Intron Database (MOID), which contains all known introns within the human, mouse and rat genomes. Introns from orthologous genes from these three species, that have the same position relative to the reading frame, are grouped in a special orthologous intron table. Our program SNO.pl searches for conserved snoRNA motifs within MOID and reports all cases when characteristic snoRNA-like structures are present in all three orthologous introns of human, mouse and rat sequences. Here we report an example of the SNO.pl usage for searching a particular pattern of conserved C/D-box snoRNA motifs (canonical C- and D-boxes and the 6 nt long terminal stem). In this computer analysis, we detected 57 triplets of snoRNA-like structures in three mammals. Among them were 15 triplets that represented known C/D-box snoRNA genes. Six triplets represented snoRNA genes that had only been partially characterized in the mouse genome. One case represented a novel snoRNA gene, and another three cases, putative snoRNAs. Our programs are publicly available and can be easily adapted and/or modified for searching any conserved motifs within mammalian introns.

Algorithms↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

The impact of microbial genomics on antimicrobial drug development.

There is an urgent need to develop novel classes of antibiotics to counter the threat of the spread of multiply resistant bacterial pathogens. The availability of the complete genome sequence of many pathogenic microbes provides information on every potential drug target and is an invaluable resource in the search for novel compounds. Here, we review the approaches being taken to exploit the genome databases through a combination of bioinformatics, transcriptional analysis, and a further understanding of the molecular basis of the disease process. The emphasis is changing from compound screening to target hunting, as the latter offers flexible ways to design and optimize the next generation of broad-spectrum antibiotics.

Anti-Bacterial Agents↗

Automated structure extraction and XML conversion of life science database flat files.

In the light of the increasing number of biological databases, their integration is a fundamental prerequisite for answering complex biological questions. Database integration, therefore, is an important area of research in bioinformatics. Since most of the publicly available life science databases are still exclusively exchanged by means of proprietary flat files, database integration requires parsers for very different flat file formats. Unfortunately, the development and maintenance of database specific flat file parsers is a nontrivial and time-consuming task, which takes considerable effort in large-scale integration scenarios. This paper introduces heuristically based concepts for automatic structure extraction from life science database flat files. On the basis of these concepts the FlatEx prototype is developed for the automatic conversion of flat files into XML representations.

Algorithms↗

Characterization of 43 non-protein-coding mRNA genes in Arabidopsis, including the MIR162a-derived transcripts.

Messenger RNAs that do not contain a long open reading frame (ORF) or non-protein-coding RNAs (npcRNAs) are an emerging novel class of transcripts. Their functions may involve the RNA molecule itself and/or short ORF-encoded peptides. npcRNA genes are difficult to identify using standard gene prediction programs that rely on the presence of relatively long ORFs. Here, we used detailed bioinformatic analyses of expressed sequence tag/cDNA databases to detect a restricted set of npcRNAs in the Arabidopsis (Arabidopsis thaliana) genome and further characterized these transcripts using a combination of bioinformatic and molecular approaches. Compositional analyses revealed strong nucleotide strand asymmetries in the npcRNAs, as well as a biased GC content, suggesting the existence of functional constraints on these RNAs. Thirteen of these transcripts display tissue-specific expression patterns, and three are regulated in conditions affecting root architecture. The npcRNA 78 gene contains the miR162 sequence in an alternative intron and corresponds to the MIR162a locus. Although DICER-LIKE 1 (DCL1) mRNA is known to be regulated by miR162-guided cleavage, its level does not change in a mir162a mutant. Alternative splicing of npcRNA 78 leads to several transcript isoforms, which all accumulate in a dcl1 mutant. This suggests that npcRNA 78 is a genuine substrate of DCL1 and that splicing of this microRNA primary transcript and miR162 processing are competitive nuclear events. Our results provide new insights into Arabidopsis npcRNA biology and the potential roles of these genes.

Alternative Splicing↗

[Methods of statistical genetics and use of database for genome information].

Knowledge and technology of bioinformatics have become inevitable for gene and genome research. Education and research in this field of science are not sufficient in Japan. There are two different approaches to trait mapping, the way by which traits are mapped on the genome. Thus, the knowledge-based approach uses functions of molecules while the statistics-based approach uses polymorphisms. Statistics-based approach uses two different methods, linkage analysis and analysis based on linkage disequilibrium. Various phenotypes are efficiently mapped on the genome using such methods. Recently, bioinformatic data base search is mostly performed using internet. Anyone can perform sequence-search, homology-search and SNP-search. Since such data bases change quickly, readers should access the databases themselves and be used to the procedures for them.

Computational Biology↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

Bioinformatics.

Explore the source record for details and available documents.

Biotechnology↗

Virtual resource development in the glycosciences.

The development of Internet-based virtual resources is a relatively new area of scientific and technical activity that is currently undergoing rapid expansion. Major factors fuelling recent growth include the emergence of multimedia capabilities through the rapid evolution of the World Wide Web, the reduction in cost of high quality personal computers and graphics workstations and the provision of mass-marketed provider services. Prior to 1995 the presence of Internet resources in the glycosciences was virtually non-existent. Existing scientific knowledge was primarily made available on the Net through the provision of databases from gopher and ftp sites. A particular example in the glycosciences is the Carbbank database of biological carbohydrate sequences. We will describe here our efforts in 1994-95 in establishing The Glycoscience Network (TGN, http:@ bellatrix.pcl.ox.ac.uk/TGN/). These activities included the establishment of a newsgroup, mailing lists, Web resources and the running of the First Electronic Glycoscience Conference (EGC-1, http:@bellatrix.pcl.ox.ac.uk/egc/). EGC-1 included many novel initiatives in the glycosciences including electronic posters and papers, a Virtual Conference Centre, a Web-based hyperglossary, Virtual Trade and Employment Centres, refereed electronic publishing, and the creation of a Virtual Reality Gallery. We would like to look towards the near future and discuss several initiatives in virtual resource creation that we believe will have significant scientific impact on the glycosciences including the development of bioinformatics-based servers, sophisticated interactive databases, and videoconferencing. Furthermore, we cherish the belief that these resources will foster international scientific collaboration and progress of an extent never previously possible. Finally, we indulge in speculation and make some suggestions on the form and long-term impact of Glycoscience Virtual Resources. We predict that their development may completely reconstruct the scientific environment that we work in as scientists and we reflect on the probable benefits and pitfalls to be encountered.

Carbohydrates↗

Implications for molecular mechanisms of glycoprotein hormone receptors using a new sequence-structure-function analysis resource.

Comparison between wild-type and mutated glycoprotein hormone receptors (GPHRs), TSH receptor, FSH receptor, and LH-chorionic gonadotropin receptor is established to identify determinants involved in molecular activation mechanism. The basic aims of the current work are 1) the discrimination of receptor phenotypes according to the differences between activity states they represent, 2) the assignment of classified phenotypes to three-dimensional structural positions to reveal 3) functional-structural hot spots and 4) interrelations between determinants that are responsible for corresponding activity states. Because it is hard to survey the vast amount of pathogenic and site-directed mutations at GPHRs and to improve an almost isolated consideration of individual point mutations, we present a system for systematic and diversified sequence-structure-function analysis (http://www.fmp-berlin.de/ssfa). To combine all mutagenesis data into one set, we converted the functional data into unified scaled values. This at least enables their comparison in a rough classification manner. In this study we describe the compiled data set and a wide spectrum of functions for user-driven searches and classification of receptor functionalities such as cell surface expression, maximum of hormone binding capability, and basal as well as hormone-induced Galphas/Galphaq mediated cAMP/inositol phosphate accumulation. Complementary to known databases, our data set and bioinformatics tools allow functional and biochemical specificities to be linked with spatial features to reveal concealed structure-function relationships by a semiquantitative analysis. A comprehensive discrimination of specificities of pathogenic mutations and in vitro mutant phenotypes and their relation to signaling mechanisms of GPHRs demonstrates the utility of sequence-structure-function analysis. Moreover, new interrelations of determinants important for selective G protein-mediated activation of GPHRs are resumed.

Animals↗

Molecular modeling of phosphorylation sites in proteins using a database of local structure segments.

A new bioinformatics tool for molecular modeling of the local structure around phosphorylation sites in proteins has been developed. Our method is based on a library of short sequence and structure motifs. The basic structural elements to be predicted are local structure segments (LSSs). This enables us to avoid the problem of non-exact local description of structures, caused by either diversity in the structural context, or uncertainties in prediction methods. We have developed a library of LSSs and a profile--profile-matching algorithm that predicts local structures of proteins from their sequence information. Our fragment library prediction method is publicly available on a server (FRAGlib), at http://ffas.ljcrf.edu/Servers/frag.html . The algorithm has been applied successfully to the characterization of local structure around phosphorylation sites in proteins. Our computational predictions of sequence and structure preferences around phosphorylated residues have been confirmed by phosphorylation experiments for PKA and PKC kinases. The quality of predictions has been evaluated with several independent statistical tests. We have observed a significant improvement in the accuracy of predictions by incorporating structural information into the description of the neighborhood of the phosphorylated site. Our results strongly suggest that sequence information ought to be supplemented with additional structural context information (predicted with our segment similarity method) for more successful predictions of phosphorylation sites in proteins.

Amino Acid Sequence↗

ODB: a database of operons accumulating known operons across multiple genomes.

Operon structures play an important role in co-regulation in prokaryotes. Although over 200 complete genome sequences are now available, databases providing genome-wide operon information have been limited to certain specific genomes. Thus, we have developed an ODB (Operon DataBase), which provides a data retrieval system of known operons among the many complete genomes. Additionally, putative operons that are conserved in terms of known operons are also provided. The current version of our database contains about 2000 known operon information in more than 50 genomes and about 13 000 putative operons in more than 200 genomes. This system integrates four types of associations: genome context, gene co-expression obtained from microarray data, functional links in biological pathways and the conservation of gene order across the genomes. These associations are indicators of the genes that organize an operon, and the combination of these indicators allows us to predict more reliable operons. Furthermore, our system validates these predictions using known operon information obtained from the literature. This database integrates known literature-based information and genomic data. In addition, it provides an operon prediction tool, which make the system useful for both bioinformatics researchers and experimental biologists. Our database is accessible at http://odb.kuicr.kyoto-u.ac.jp/.

Databases, Nucleic Acid↗

Detection of hypothetical proteins in human fetal perireticular nucleus.

There is a legion of hypothetical proteins (HP) in prokaryotic and eukaryotic proteomes and the aim of this study was to describe HP in the perireticular nucleus (PN), a key structure in human brain development. Tissue from four PNs was homogenized and extracted proteins were run on two-dimensional gel electrophoresis followed by in-gel digestion and mass spectrometrical identification of proteins. Several databases were used for obtaining bioinformatic information and searching for functional and structural domains. Five spots represented HP: KIAA0423 protein (Q9Y4F4), hypothetical protein KIAA0153 (Q14166), hypothetical protein DKFZp564A2416 (Q9NTW4), hypothetical protein DKFZp564H1122 (Q9H0W9), and hypothetical protein DKFZp564D1378 (Q9H0R4). These structures were predicted to serve in cell cycle, DNA-condensation, neurogenesis, or apoptosis. The existence of formerly HP proteins in the PN of human fetal brain is shown, thus extending knowledge of the brain proteome and proposing the method used as a suitable analytical tool for searching HP.

Apoptosis↗

LASS6, an additional member of the longevity assurance gene family.

Longevity assurance genes (LAGs) represent a subgroup of the homeobox gene family. Five mammalian homologs have been reported, and the corresponding proteins have previously been investigated with respect to their key role in ceramide synthesis. However, members of the LAG family have been shown to be involved in cell growth regulation and cancer differentiation. In an effort to characterize additional members of the LAG family, we have screened the latest releases of genomic databases and report on the bioinformatic characterization of yet another member, LAG1 longevity assurance homolog 6 (LASS6). Like other LAG family members, the LASS6 protein contained a homeodomain and LAG1 domain. In phylogenetic analyses, it displayed highest homology to LASS5. The corresponding gene was localized to human chromosome 2q24.3, spanning a rather large genomic region of 318 kb. Orthologous sequences in mouse and zebrafish suggested a conservation of LASS6 in vertebrates as the protein and corresponding genomic sequences were highly conserved. LASS6 expression was analyzed in silico, and the gene was shown to be broadly expressed in a wide range of tissues. Furthermore, available microarray data suggested a role in cancer differentiation and early embryonic development.

Amino Acid Sequence↗

FEMME database: topologic and geometric information of macromolecules.

FEMME (Feature Extraction in a Multi-resolution Macromolecular Environment: http://www.biocomp.cnb.uam.es/FEMME/) database version 1.0 is a new bioinformatics data resource that collects topologic and geometric information obtained from macromolecular structures solved by three-dimensional electron microscopy (3D-EM). Although the FEMME database is focused on medium resolution data, the methodology employed (based on the so-called alpha-shape theory) is applicable to atomic resolution data as well. The alpha-shape representation allows the automatic extraction of structural features from 3D-EM volumes and their subsequent characterisation. FEMME is being populated with 3D-EM data stored in the electron microscopy database EMD-DB (http://www.ebi.ac.uk/msd/). However, and since the number of entries in EMD-DB is still relatively small, FEMME is also being populated in this initial phase with structural data from PDB and PQS databases (http://www.rcsb.org/pdb/ and pqs.ebi.ac.uk/, respectively) whose resolution has been lowered accordingly. Each FEMME entry contains macromolecular geometry and topology information with a detailed description of its structural features. Moreover, FEMME data have facilitated the study and development of a method to retrieve macromolecular structures by their structural content based on the combined use of spin images and neural networks with encouraging results. Therefore, the FEMME database constitutes a powerful tool that provides a uniform and automatic way of analysing volumes coming from 3D-EM that will hopefully help the scientific community to perform wide structural comparisons.

Computational Biology↗

Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories.

With the human Plasma Proteome Project (PPP) pilot phase completed, the largest and most ambitious proteomics experiment to date has reached its first milestone. The correspondingly impressive amount of data that came from this pilot project emphasized the need for a centralized dissemination mechanism and led to the development of a detailed, PPP specific data gathering infrastructure at the University of Michigan, Ann Arbor as well as the protein identifications database project at the European Bioinformatics Institute as a general proteomics data repository. One issue that crept up while discussing which data to store for the PPP concerns whether the raw, binary data coming from the mass spectrometers should be stored, or rather the more compact and already significantly processed peak lists. As this debate is not restricted to the PPP but relates to the proteomics community in general, we will attempt to detail the relative merits and caveats associated with centralized storage and dissemination of raw data and/or peak lists, building on the extensive experience gained during the PPP pilot phase. Finally, some suggestions are made for both immediate and future storage of MS data in public repositories.

Computational Biology↗

Toward computer-based cleavage site prediction of cysteine endopeptidases.

Identification of relevant substrates is essential for elucidation of in vivo functions of peptidases. The recent availability of the complete genome sequences of many eukaryotic organisms holds the promise of identifying specific peptidase substrates by systematic proteome analyses in combination with computer-based screening of genome databases. Currently available proteomics and bioinformatics tools are not sufficient for reliable endopeptidase substrate predictions. To address these shortcomings the bioinformatics tool 'PEPS' (Prediction of Endopeptidase Substrates) has been developed and is presented here. PEPS uses individual rule-based endopeptidase cleavage site scoring matrices (CSSM). The efficiency of PEPS in predicting putative caspase 3, cathepsin B and cathepsin L cleavage sites is demonstrated in comparison to established algorithms. Mortalin, a member of the heat shock protein family HSP70, was identified by PEPS as a putative cathepsin L substrate. Comparative proteome analyses of cathepsin L-deficient and wild-type mouse fibroblasts showed that mortalin is enriched in the absence of cathepsin L. These results indicate that CSSM/PEPS can correctly predict relevant peptidase substrates.

Animals↗

Recent developments in proteomics: implications for the study of cardiac hypertrophy and failure.

The key components to the molecular understanding of the pathophysiology of various forms of heart failure involve global and/or large-scale identifications of proteins, their patterns of expression, posttranslational modifications, and functional characterization. Particularly, proteins involved in the induction of cardiac (mal)adaptive hypertrophic growth, interstitial fibrosis, and contractile dysfunction are of interest. In general, with the accumulation of vast amounts of DNA sequences in databases, researchers have become aware that merely having complete sequences of genomes and transcriptional changes for thousands of genes simultaneously will not be sufficient to elucidate, in molecular terms, the etiology and pathophysiology of cardiovascular disease. In the last decade, a new technology called proteomics has become available that allows biological and (patho)physiological questions to be approached exclusively from the protein perspective. Proteomics may enable us to map the entire complement of proteins expressed by the heart at any time and condition. This approach creates the unique possibility to identify, by differential analysis, protein alterations associated with the etiology of heart disease and its progression, outcome, and response to therapy. To illustrate the true power of proteomics, most of the currently available methodologies are first reviewed, including their limitations. This review also deals with the current status and the perspectives of proteomics applications in research on heart failure in general. Furthermore, examples of our recent data on global protein profiling of the pressure-overloaded rat right ventricle and of endothelin-1-stimulated cultures of neonatal rat cardiac myocytes are provided. The last section is devoted to the continuous advances in proteomic technologies, including protein separation methods, mass spectrometric instrumentation, computational analysis, and bioinformatic tools, together with integrative databases.

Cardiomegaly↗