Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Molecular evolution of serine/arginine splicing factors family (SR) by positive selection.

The serine-rich (SR) protein family is involved in the pre-mRNA splicing process and the DNA sequences of the corresponding genes are highly conserved in the metazoan organisms. The mammalian SR proteins consist of one or two characteristic RNA binding domains (RBD), containing the signature sequences RDAEDA and SWQDLKD and a RS (arginine/serine-rich) domain. We used the amino acid and nucleotide sequences deposited in GenBank and Swiss-Prot databases to perform a phylogenetic analysis using bioinformatics tools. The results of the phylogenetic trees suggest that this family has evolved by several gene duplication events as a result of a positive selection mechanism.

Amino Acid Sequence↗

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics↗

SSEP: Secondary structural elements of proteins.

SSEP is a comprehensive resource for accessing information related to the secondary structural elements present in the 25 and 90% non-redundant protein chains. The database contains 1771 protein chains from 1670 protein structures and 6182 protein chains from 5425 protein structures in 25 and 90% non-redundant protein chains, respectively. The current version provides information about the alpha-helical segments and beta-strand fragments of varying lengths. In addition, it also contains the information about 3(10)-helix, beta- and nu-turns and hairpin loops. The free graphics program RASMOL has been interfaced with the search engine to visualize the three-dimensional structures of the user queried secondary structural fragment. The database is updated regularly and is available through Bioinformatics web server at http://cluster.physics.iisc.ernet.in/ssep/ or http://144.16.71.148/ssep/.

Databases, Protein↗

SDAP: database and computational tools for allergenic proteins.

SDAP (Structural Database of Allergenic Proteins) is a web server that provides rapid, cross-referenced access to the sequences, structures and IgE epitopes of allergenic proteins. The SDAP core is a series of CGI scripts that process the user queries, interrogate the database, perform various computations related to protein allergenic determinants and prepare the output HTML pages. The database component of SDAP contains information about the allergen name, source, sequence, structure, IgE epitopes and literature references and easy links to the major protein (PDB, SWISS-PROT/TrEMBL, PIR-ALN, NCBI Taxonomy Browser) and literature (PubMed, MEDLINE) on-line servers. The computational component in SDAP uses an original algorithm based on conserved properties of amino acid side chains to identify regions of known allergens similar to user-supplied peptides or selected from the SDAP database of IgE epitopes. This and other bioinformatics tools can be used to rapidly determine potential cross-reactivities between allergens and to screen novel proteins for the presence of IgE epitopes they may share with known allergens. SDAP is available via the World Wide Web at http://fermi.utmb.edu/SDAP/.

Allergens↗

MitBASE pilot: a database on nuclear genes involved in mitochondrial biogenesis and its regulation in Saccharomyces cerevisiae.

In the framework of the EU BIOTECH PROGRAM and within the 'MITBASE: a comprehensive and integrated database on mtDNA' project, we have prepared a pilot database (MitBASE Pilot) on nuclear genes involved in mitochondrial biogenesis and its regulation in Saccharomyces cerevisiae. MitBASE Pilot includes nuclear genes encoding mitochondrial proteins as well as nuclear genes encoding products which are localised in other sub-cellular compartments but nevertheless interact with mitochondrial functions. Genes have been classified on the basis of the mitochondrial process in which they participate and the mitochondrial phenotype of the gene knockout. The structure of the MitBASE Pilot database has been conceived for a flexible organisation of the information. An intuitive visual query system has been developed which allows users to select information in different combinations, both in the query and the output format, according to their needs. MitBASE Pilot is a relational database, is maintained at the EMBL-European Bioinformatics Institute (EBI) and is available at the World Wide Web site http://www3.ebi.ac. uk/Research/Mitbase/mitbiog.pl

Cell Nucleus↗

MetaCyc: a multiorganism database of metabolic pathways and enzymes.

MetaCyc is a database of metabolic pathways and enzymes located at http://MetaCyc.org/. Its goal is to serve as a metabolic encyclopedia, containing a collection of non-redundant pathways central to small molecule metabolism, which have been reported in the experimental literature. Most of the pathways in MetaCyc occur in microorganisms and plants, although animal pathways are also represented. MetaCyc contains metabolic pathways, enzymatic reactions, enzymes, chemical compounds, genes and review-level comments. Enzyme information includes substrate specificity, kinetic properties, activators, inhibitors, cofactor requirements and links to sequence and structure databases. Data are curated from the primary literature by curators with expertise in biochemistry and molecular biology. MetaCyc serves as a readily accessible comprehensive resource on microbial and plant pathways for genome analysis, basic research, education, metabolic engineering and systems biology. Querying, visualization and curation of the database is supported by SRI's Pathway Tools software. The PathoLogic component of Pathway Tools is used in conjunction with MetaCyc to predict the metabolic network of an organism from its annotated genome. SRI and the European Bioinformatics Institute employed this tool to create pathway/genome databases (PGDBs) for 165 organisms, available at the BioCyc.org website. These PGDBs also include predicted operons and pathway hole fillers.

Animals↗

Design and implementation of a library-based information service in molecular biology and genetics at the University of Pittsburgh.

SETTING: In summer 2002, the Health Sciences Library System (HSLS) at the University of Pittsburgh initiated an information service in molecular biology and genetics to assist researchers with identifying and utilizing bioinformatics tools. PROGRAM COMPONENTS: This novel information service comprises hands-on training workshops and consultation on the use of bioinformatics tools. The HSLS also provides an electronic portal and networked access to public and commercial molecular biology databases and software packages. EVALUATION MECHANISMS: Researcher feedback gathered during the first three years of workshops and individual consultation indicate that the information service is meeting user needs. NEXT STEPS/FUTURE DIRECTIONS: The service's workshop offerings will expand to include emerging bioinformatics topics. A frequently asked questions database is also being developed to reuse advice on complex bioinformatics questions.

Computational Biology↗

Identification of novel genes preferentially expressed in the retina using a custom human retina cDNA microarray.

PURPOSE: To construct a custom cDNA microarray for comprehensive human retinal gene expression profiling and apply it to the identification of genes that are preferentially expressed in the retina. METHODS: A cDNA microarray was constructed based on the predicted human retina gene expression profile according to expressed sequence tag (EST) databases. Gene expression profiles were obtained from five human retinas, two livers, and the cerebral cortical regions of two brains. Each sample was studied in duplicate, using a reference sample experimental design. Retina-enriched genes were identified by using the significance analysis for microarray (SAM) algorithm. Quantitative real time PCR was used to confirm microarray results. Bioinformatic analysis was performed to compare the array results with expression data available from public databases. RESULTS: The cDNA microarray contains 10,034 sequences: 67% represent known genes and 33% represent ESTs. Differential hybridization with the array identified, in addition to known retinal genes, 186 retina-enriched genes that do not have known retinal function. Of these, 96 represent novel genes. Quantitative real-time PCR of 11 of the identified genes and ESTs confirmed their retina-enriched expression pattern. Bioinformatic analysis of EST databases suggests that of the 186 genes, approximately 40% are predominantly expressed in the retina, whereas the remainder show significant expression in other tissues. Comparison of this study's microarray-based retina-enriched gene set with three published similar sets identified using complementary high-throughput approaches demonstrated only limited overlap of the identified genes. CONCLUSIONS: Because previous studies have demonstrated that many retina-enriched genes are crucial for maintaining normal retinal function, the genes identified here are likely to include ones that have important roles in the retina and ones that when mutated can cause or modulate retinal disease. In addition, the retina custom array should provide a useful resource for comparing expression profiles between normal and diseased human retinas.

Adult↗

Bioinformatics approaches to cancer gene discovery.

The Cancer Gene Anatomy Project (CGAP) database of the National Cancer Institute has thousands of known and novel expressed sequence tags (ESTs). These ESTs, derived from diverse normal and tumor cDNA libraries, offer an attractive starting point for cancer gene discovery. Data-mining the CGAP database led to the identification of ESTs that were predicted to be specific to select solid tumors. Two genes from these efforts were taken to proof of concept for diagnostic and therapeutics indications of cancer. Microarray technology was used in conjunction with bioinformatics to understand the mechanism of one of the targets discovered. These efforts provide an example of gene discovery by using bioinformatics approaches. The strengths and weaknesses of this approach are discussed in this review.

Basic Helix-Loop-Helix Proteins↗

The EMBL nucleotide sequence database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/) is maintained at the European Bioinformatics Institute (EBI) in an international collaboration with the DNA Data Bank of Japan (DDBJ) and GenBank at the NCBI (USA). Data is exchanged amongst the collaborating databases on a daily basis. The major contributors to the EMBL database are individual authors and genome project groups. Webin is the preferred web-based submission system for individual submitters, whilst automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO). Database releases are produced quarterly. Network services allow free access to the most up-to-date data collection via ftp, email and World Wide Web interfaces. EBI's Sequence Retrieval System (SRS), a network browser for databanks in molecular biology, integrates and links the main nucleotide and protein databases plus many specialized databases. For sequence similarity searching a variety of tools (e.g. Blitz, Fasta, BLAST) are available which allow external users to compare their own sequences against the latest data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Computational Biology↗

Hickam 2000: the maturation of, and linkages between, medical informatics and bioinformatics.

I have always been infatuated with computers and convinced of their potential for solving problems in biologic research and clinical care. In the 1960s I thought we could use the computer to predict the shape of macromolecules from their chemical formulas and fundamental physical chemical principles. However, with the computers of the 1960s that was a fantasy. So I focused on the use of computers to manage medical record content and to assist with clinical care. The Electronic Medical Record (EMR) we began developing in 1972 with 33 diabetes patients now carries nearly 300 million separate results for more than 3 million patients. The data include lab and other diagnostic studies, dictated notes, orders, encounter records, radiology images, electrocardiograph tracings, and motion cardiac echoes, and the care provider at Indiana University and Wishard Hospital is accessed 10 million times per year. We have also agitated for standards to make the collection of these data easier. This work has become part of a field called medical informatics. In the meantime, the application of computers to biology has rapidly matured into a field called bioinformatics, and researchers in this field now provide annotated databases for many categories of molecules, programs for "matching" newly discovered genomic sequences with previously studied sequences, and systems for storing and processing massive amounts of genomic and molemic data. They have developed sophisticated methods for predicting the shape of biologic macromolecules and other important insights about biology and evolution. Medical informatics and bioinformatics intersect at many points. The most important intersection is between electronic medical records and the human specimen databases that can link genotype to the phenotype, as needed, to unravel polygenetic disease causality. The National Cancer Institute is embarking on an intriguing effort to use EMRs (phenotype) to link to paraffin blocks (genotype) in pathology laboratories where opportunities for cancer genomic discovery are open. We will participate in this effort and look forward to bending the EMR we developed for clinical use to bioinformatics uses as well.

Clinical Medicine↗

RADARS, a bioinformatics solution that automates proteome mass spectral analysis, optimises protein identification, and archives data in a relational database.

RADARS, a rapid, automated, data archiving and retrieval software system for high-throughput proteomic mass spectral data processing and storage, is described. The majority of mass spectrometer data files are compatible with RADARS, for consistent processing. The system automatically takes unprocessed data files, identifies proteins via in silico database searching, then stores the processed data and search results in a relational database suitable for customized reporting. The system is robust, used in 24/7 operation, accessible to multiple users of an intranet through a web browser, may be monitored by Virtual Private Network, and is secure. RADARS is scalable for use on one or many computers, and is suited to multiple processor systems. It can incorporate any local database in FASTA format, and can search protein and DNA databases online. A key feature is a suite of visualisation tools (many available gratis), allowing facile manipulation of spectra, by hand annotation, reanalysis, and access to all procedures. We also described the use of Sonar MS/MS, a novel, rapid search engine requiring 40 MB RAM per process for searches against a genomic or EST database translated in all six reading frames. RADARS reduces the cost of analysis by its efficient algorithms: Sonar MS/MS can identifiy proteins without accurate knowledge of the parent ion mass and without protein tags. Statistical scoring methods provide close-to-expert accuracy and brings robust data analysis to the non-expert user.

Amino Acid Sequence↗

EMBL Nucleotide Sequence Database: developments in 2005.

The EMBL Nucleotide Sequence Database (www.ebi.ac.uk/embl) at the EMBL European Bioinformatics Institute, UK, offers a comprehensive set of publicly available nucleotide sequence and annotation, freely accessible to all. Maintained in collaboration with partners DDBJ and GenBank, coverage includes whole genome sequencing project data, directly submitted sequence, sequence recorded in support of patent applications and much more. The database continues to offer submission tools, data retrieval facilities and user support. In 2005, the volume of data offered has continued to grow exponentially. In addition to the newly presented data, the database encompasses a range of new data types generated by novel technologies, offers enhanced presentation and searchability of the data and has greater integration with other data resources offered at the EBI and elsewhere. In stride with these developing data types, the database has continued to develop submission and retrieval tools to maximise the information content of submitted data and to offer the simplest possible submission routes for data producers. New developments, the submission process, data retrieval and access to support are presented in this paper, along with links to sources of further information.

Animals↗

Sustainable databases.

Although the flood of cell biological knowledge rises relentlessly, many databases face an uncertain future. Unless funding for essential bioinformatic resources is set in stone, the next storm may wash away the foundation of future cell biology research.

Databases as Topic↗

[Construction of rice dwarf virus genome database].

Secondary database construction is an important subject in the field of bioinformatics. As the full genomic sequences of some organisms are being completed and followed by structural and functional studies, construction of secondary database becomes essential on the agenda. The rice dwarf virus (RDV) is a pathogen infecting rice in China, Japan and the Southeastern Asia region and leading to considerable economic loss. Based on the data generated from recent genomic research and earlier biochemical studies scattered in various primary databases and scientific journals, we have constructed a compact, user-friendly and non-redundant job-oriented secondary database. This work will provide compiled useful information for plant molecular biologists as well as in achieving preliminary experiences in secondary database construction.

Databases, Factual↗

Predicting the nuclear localization signals of 107 types of HPV L1 proteins by bioinformatic analysis.

In this study, 107 types of human papillomavirus (HPV) L1 protein sequences were obtained from available databases, and the nuclear localization signals (NLSs) of these HPV L1 proteins were analyzed and predicted by bioinformatic analysis. Out of the 107 types, the NLSs of 39 types were predicted by PredictNLS software (35 types of bipartite NLSs and 4 types of monopartite NLSs). The NLSs of the remaining HPV types were predicted according to the characteristics and the homology of the already predicted NLSs as well as the general rule of NLSs. According to the result, the NLSs of 107 types of HPV L1 proteins were classified into 15 categories. The different types of HPV L1 proteins in the same NLS category could share the similar or the same nucleocytoplasmic transport pathway. They might be used as the same target to prevent and treat different types of HPV infection. The results also showed that bioinformatic technology could be used to analyze and predict NLSs of proteins.

Amino Acid Sequence↗

In silico tools for signal transduction research.

Signal transduction is a fundamental process that takes place in all living organisms and understanding how this event occurs at the cellular level is of vital importance to virtually all fields of biomedicine. There are several major steps involved in deciphering the signalling pathways: (a) Which molecules are involved in signalling? (b) Who talks to whom?, ie making sense of the molecular interactions in a context-dependent way. (c) Where are the signalling events taking place?, eg when a resting cell becomes activated. The challenge lies in reconstructing signalling modules and networks evoked in a particular response to a single input as well as correlating the signalling response to different cellular inputs. There is also the need for interpretation of cross-talk between signalling modules in response to single and multiple inputs. To follow up these questions there are many good databases that provide an information system on regulatory networks. This review aims to find some of the bioinformatics tools and websites available to conduct signal transduction research and to discuss the representation of databases available for the processes of signalling. The databases considered here can provide a well-structured overview on the subject and a basis for advanced bioinformatics analysis to interpret the function of genomic sequences or to analyse signalling networks within a cell. However, the knowledge of most signalling pathways is incomplete and for this reason the existing databases will provide insight, but very rarely a more complete picture.

Computational Biology↗

Multiexon skipping leading to an artificial DMD protein lacking amino acids from exons 45 through 55 could rescue up to 63% of patients with Duchenne muscular dystrophy.

Approximately two-thirds of Duchenne muscular dystrophy (DMD) patients show intragenic deletions ranging from one to several exons of the DMD gene and leading to a premature stop codon. Other deletions that maintain the translational reading frame of the gene result in the milder Becker muscular dystrophy (BMD) form of the disease. Thus the opportunity to transform a DMD phenotype into a BMD phenotype appeared as a new treatment strategy with the development of antisense oligonucleotides technology, which is able to induce an exon skipping at the pre-mRNA level in order to restore an open reading frame. Because the DMD gene contains 79 exons, thousands of potential transcripts could be produced by exon skipping and should be investigated. The conventional approach considers skipping of a single exon. Here we report the comparison of single- and multiple-exon skipping strategies based on bioinformatic analysis. By using the Universal Mutation Database (UMD)-DMD, we predict that an optimal multiexon skipping leading to the del45-55 artificial dystrophin (c.6439_8217del) could transform the DMD phenotype into the asymptomatic or mild BMD phenotype. This multiple-exon skipping could theoretically rescue up to 63% of DMD patients with a deletion, while the optimal monoskipping of exon 51 would rescue only 16% of patients.

Adolescent↗