Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

A graphical query generator for clinical research databases.

Clinical research involves recording, storage and retrieval of disease-related patient data, typically using a database system. In order to facilitate ad hoc queries to clinical databases we have developed a query generator with a graphical interface. The query generator uses an object-oriented data model which is visualized by directed graphs. The main focus of our work was the definition of object-oriented user views to the partly complex data structures of a relational database. Furthermore, we tried to define graphical abstractions for all common types of queries. Thus, even for non-expert database users such as clinicians, it is easy to assemble highly complex queries for a thorough examination of the content of large research databases.

Computer Graphics↗

The AAQA/ACHS National QA database: a 12 month update. Australasian Association for Quality in Health Care. Australian Council on Healthcare Standards.

In its first 12 months of operation, the AAQA/ACHS National QA Database has succeeded in its aim of providing a networking channel for quality coordinators and managers. Thanks to the cooperation and interest of many health care professionals, the database now holds some 1800 entries and in its first year was accessed by more than 160 individuals. Contributors rated 86% of quality activities submitted as effective or very effective. The majority of entries were submitted by facilities with less than 150 beds and almost half of the entries came from NSW. Over half the entries were submitted by nursing and allied health professionals. The article describes the most common QA methods contributed to the database. Ongoing input to the database is essential. Individuals actively involved in quality activities are urged to contribute their experience, particularly in areas currently under-represented in the database.

Australia↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

Improving interoperability between microbial information and sequence databases.

BACKGROUND: Biological resources are essential tools for biomedical research. Their availability is promoted through on-line catalogues. Common Access to Biological Resources and Information (CABRI) is a service for distribution of biological resources and related data collected by 28 European culture collections. Linking this information to bioinformatics databanks can make the collections' holdings more visible after a search in molecular biology databanks and vice-versa. Identification of links to sequence databases can be useful, but annotation and indexing problems, together with compilation errors, immediately arise. In this paper, we present our efforts for the identification of cross-references between CABRI catalogues and the EMBL Data Library and related results. RESULTS: An SRS site with both EMBL and CABRI catalogues has been set up. Ad-hoc changes in indexing scripts allowed to achieve homogeneous index keys and SRS link features have been used to identify links between databases. After manual checking and comparison with an alternative procedure, about 67,500 valid cross-references were identified, added to the EMBL Data Library and are now distributed with it. HTML links can be established from EMBL to CABRI network service. Procedures can be executed whenever needed. CONCLUSION: Links between EMBL and CABRI catalogues constitute an improved access to micro-organisms of certified quality and can produce positive effects on biomedical research. Further links between CABRI catalogues and other bioinformatics databases can now easily be defined by using these cross-references. Linking genetic information onto natural resources information may stand model for the integration of other databases containing empirical data on these materials.

Base Sequence↗

Rhinoplasty perioperative database using a personal digital assistant.

OBJECTIVE: To construct a reliable, accurate, and easy-to-use handheld computer database that facilitates the point-of-care acquisition of perioperative text and image data specific to rhinoplasty. METHODS: A user-modified database (Pendragon Forms [v.3.2]; Pendragon Software Corporation, Libertyville, Ill) and graphic image program (Tealpaint [v.4.87]; Tealpaint Software, San Rafael, Calif) were used to capture text and image data, respectively, on a Palm OS (v.4.11) handheld operating with 8 megabytes of memory. The handheld and desktop databases were maintained secure using PDASecure (v.2.0) and GoldSecure (v.3.0) (Trust Digital LLC, Fairfax, Va). The handheld data were then uploaded to a desktop database of either FileMaker Pro 5.0 (v.1) (FileMaker Inc, Santa Clara, Calif) or Microsoft Access 2000 (Microsoft Corp, Redmond, Wash). DESIGN: Patient data were collected from 15 patients undergoing rhinoplasty in a private practice outpatient ambulatory setting. Data integrity was assessed after 6 months' disk and hard drive storage. RESULTS: The handheld database was able to facilitate data collection and accurately record, transfer, and reliably maintain perioperative rhinoplasty data. Query capability allowed rapid search using a multitude of keyword search terms specific to the operative maneuvers performed in rhinoplasty. CONCLUSIONS: Handheld computer technology provides a method of reliably recording and storing perioperative rhinoplasty information. The handheld computer facilitates the reliable and accurate storage and query of perioperative data, assisting the retrospective review of one's own results and enhancement of surgical skills.

Computers, Handheld↗

Accessing the Kabat antibody sequence database by computer.

The Kabat antibody sequence database has for many years been the primary site for depositing sequence information on antibodies and other proteins of immunological interest. The chief drawback of this database has been that it has only been available in the form of a printed book (Kabat et al., Sequences of Proteins of Immunological Interest, 1991). These data have recently become available on the global computer Internet, but no method of searching the data has, as yet, been provided. Here, the development of a specialized database program for accessing the antibody data is described. This database software has been made accessible over the World Wide Web, together with a program which allows a novel antibody sequence to be tested against the Kabat sequence database, to identify unusual features of an antibody sequence which may represent cloning artifacts or sequencing errors.

Amino Acid Sequence↗

Influence of protein structure databases on the predictive power of statistical pair potentials.

A long standing goal in protein structure studies is the development of reliable energy functions that can be used both to verify protein models derived from experimental constraints as well as for theoretical protein folding and inverse folding computer experiments. In that respect, knowledge-based statistical pair potentials have attracted considerable interests recently mainly because they include the essential features of protein structures as well as solvent effects at a low computing cost. However, the basis on which statistical potentials are derived have been questioned. In this paper, we investigate statistical pair potentials derived from protein three-dimensional structures, addressing in particular questions related to the form of these potentials, as well as to the content of the database from which they are derived. We have shown that statistical pair potentials depend on the size of the proteins included in the database, and that this dependence can be reduced by considering only pairs of residue close in space (i.e., with a cutoff of 8 A). We have shown also that statistical potentials carry a memory of the quality of the database in terms of the amount and diversity of secondary structure it contains. We find, for example, that potentials derived from a database containing alpha-proteins will only perform best on alpha-proteins in fold recognition computer experiments. We believe that this is an overall weakness of these potentials, which must be kept in mind when constructing a database.

Chemical Phenomena↗

7th International HUGO Mutation Database Meeting, October 19, 1999, San Francisco, USA.

The 7th International HUGO Mutation Database Meeting was held on October 19, 1999 in conjunction with the annual meeting of the American Society of Genetics in San Fransisco, California, U.S.A. Meeting highlights are described, including discussions of topics such as the ethical aspects of variation databases, ethical guidelines which should be established immediately, data protection laws which may affect access to data, and plans to make variation databases financially self-sustaining. A resolution was passed which encourages HUGO and Mutation database Initiative (MDI) collaboration (under the name HUGO-MDI) to provide an integrated, properly funded system of variation databases.

Alleles↗

A two-dimensional electrophoresis database of rat heart proteins.

More than 3000 myocardial protein species of Wistar Kyoto rat, an important animal model, were separated by high-resolution two-dimensional gel electrophoresis (2-DE) and characterized in terms of isoelectric point (pI) and molecular mass (Mr). Currently, the 2-DE database contains 64 identified proteins; forty-three were identified by peptide mass fingerprinting (PMF) using matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS), nine by exclusive comparison with other 2-DE heart protein databases, and in only 12 cases of 60 attempts N-terminal sequencing was successful. We used the Make2ddb software package downloaded from the ExPASy server for the construction of a rat myocardial 2-DE database. The Make2ddb package simplifies the creation of a new 2-DE database if the Melanie II software and a Sun workstation under Solaris are available. Our 2-DE database of rat heart proteins can be accessed at URL http://gelmatching.inf.fu-berlin.de/pleiss/2d.

Animals↗

8th International HUGO-Mutation Database Initiative Meeting, April 9, 2000, Vancouver, Canada.

The 8th International HUGO-Mutation Database Initiative Meeting was held on April 9, 2000, in Vancouver, Canada. Meeting highlights are here described. The discussion predominantly revolved around the concept of a central mutation database, which would serve as a repository of gene sequence variants for the community. Specifications for such a central database were prepared and presented by a working group including bioinformatics experts, LSDB operators, central database operators, and an industry representative. Refinement of the specifications and consideration of the implementation of such a database was conducted through a consortium of public and private collaborators and funding. Members were urged to generate broad community support for the project.

Canada↗

A proteome database of human primary T helper cells.

We have established the first public database of human primary T helper cell proteome using two-dimensional electrophoresis (2-DE) and matrix assisted laser desorption/ionization-time of flight-mass spectrometry. For the database, CD4+ human T cells were activated with anti-CD3+anti-CD28 antibodies and metabolically labeled with [35S]methionine for 24 h. Cells were lysed and proteins were separated by 2-DE. About 1500 protein spots were detected in the resulting 2-DE gels with silver staining, and 2000 spots with autoradiography. We have identified 91 proteins from the 2-DE gels using peptide mass fingerprinting, and annotated them to our database. The identified proteins are also linked to SWISS-PROTand NCBI protein databases. Our database is available via the Internet at http://www3.btk.utu.fi:8080/Genomics/Proteomics/Database.

Databases, Protein↗

ProbID: a probabilistic algorithm to identify peptides through sequence database searching using tandem mass spectral data.

With the recent quick expansion of DNA and protein sequence databases, intensive efforts are underway to interpret the linear genetic information of DNA in terms of function, structure, and control of biological processes. The systematic identification and quantification of expressed proteins has proven particularly powerful in this regard. Large-scale protein identification is usually achieved by automated liquid chromatography-tandem mass spectrometry of complex peptide mixtures and sequence database searching of the resulting spectra [Aebersold and Goodlett, Chem. Rev. 2001, 101, 269-295]. As generating large numbers of sequence-specific mass spectra (collision-induced dissociation/CID) spectra has become a routine operation, research has shifted from the generation of sequence database search results to their validation. Here we describe in detail a novel probabilistic model and score function that ranks the quality of the match between tandem mass spectral data and a peptide sequence in a database. We document the performance of the algorithm on a reference data set and in comparison with another sequence database search tool. The software is publicly available for use and evaluation at http://www.systemsbiology.org/research/software/proteomics/ProbID.

Algorithms↗

An iterative calibration method with prediction of post-translational modifications for the construction of a two-dimensional electrophoresis database of mouse mammary gland proteins.

Protein databases serve as general reference resources providing an orientation on two-dimensional electrophoresis (2-DE) patterns of interest. The intention behind constructing a 2-DE database of the water soluble proteins from wild-type mouse mammary gland tissue was to create a reference before going on to investigate cancer-associated protein variations. This database shall be deemed to be a model system for mouse tissue, which is open for transgenic or knockout experiments. Proteins were separated and characterized in terms of their molecular weight (M(r)) and isoelectric point (pI) by high resolution 2-DE. The proteins were identified using prevalent proteomics methods. One method was peptide mass fingerprinting by matrix-assisted laser desorption/ionization-mass spectrometry. Another method was N-terminal sequencing by Edman degradation. By N-terminal sequencing M(r) and pI values were specified more accurately and so the calibration of the master gel was obtained more systematically and exactly. This permits the prediction of possible post-translational modifications of some proteins. The mouse mammary gland 2-DE protein database created presently contains 66 identified protein spots, which are clickable on the gel pattern. This relational database is accessible on the WWW under the URL: http://www.mpiib-berlin.mpg.de/2D-PAGE.

Animals↗

ALFRED: An allele frequency database for anthropology.

The deluge of data from the human genome project (HGP) presents new opportunities for molecular anthropologists to study human variation through the promise of vast numbers of new polymorphisms (e.g., single nucleotide polymorphisms or SNPs). Collecting the resulting data into a single, easily accessible resource will be important to facilitate this research. We created a prototype Web-accessible database named ALFRED (ALelle FREquency Database, http://alfred.med.yale.edu/alfred/) to store and make publicly available allele frequency data on diverse polymorphic sites for many populations. In constructing this database, we considered many different concerns relating to the types of information needed for anthropology, population genetics, molecular genetics, and statistics, as well as issues of data integrity and ease of access to data. We also developed links to other Web-based databases as well as procedures for others to make links to the data in ALFRED. Here we present an overview of the issues considered and provisional solutions, as well as an example of data already available. It is our hope that this database will be useful for research and teaching in a wide range of fields, and that colleagues from various fields will contribute to making ALFRED an important resource for many studies as yet unforeseen.

Anthropology, Physical↗

An integrated database of flavonoids.

Flavonoids are polyphenolic compounds that occur ubiquitously in foods of plant origin. Some of these molecules exhibit various physiological activities. Among existing drugs, there are a huge number of compounds bearing a flavonoid-related skeleton. Because of the relevance for pharmaceutical research, it would be beneficial to collect these compounds into a database. Recently, various databases of chemicals were compiled to help biological and/or chemical research, but no comprehensive database of flavonoids with chemical structures and physicochemical parameters, supposedly related to their activity, is available yet. The aim of this research was to merge the information about flavonoids of plant origin and flavonoids used as medicines into a database. Moreover, predictions of activities against various targets were performed using a virtual screening procedure to demonstrate a possible application of the database for pharmaceutical research.

Chemistry, Physical↗

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding↗

Microsequences of 145 proteins recorded in the two-dimensional gel protein database of normal human epidermal keratinocytes.

Microsequencing of proteins recovered from two-dimensional (2-D) gels is being used systematically to identify proteins in the master human keratinocyte 2-D gel database. To date, about 250 protein spots recorded in human 2-D gel databases have been microsequenced and, of these, 145 are recorded in the keratinocyte database under the entry partial amino acid sequence. Coomassie Brilliant Blue-stained protein spots cut from several (up to 40) dry gels were concentrated by elution-concentration gel electrophoresis, electroblotted onto PVDF membranes and digested in situ with trypsin. Eluting peptides were separated by reversed-phase HPLC, collected individually and sequenced. Computer search using the FASTA and TFASTA programs from Genetics Computer Group indicated that 110 of the microsequenced polypeptides shared significant similarity with proteins contained in the PIR, Mipsx or GenEMBL databases. Only 35 polypeptides corresponded to hitherto unknown proteins. Peptide sequences of all 145 proteins are listed together with their coordinates (apparent molecular weight and pI) in the keratinocyte database.

Amino Acid Sequence↗