Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

PIR-ALN: a database of protein sequence alignments.

MOTIVATION: The Protein Information Resource (PIR) maintains a database of annotated and curated alignments in order to visually represent interrelationships among sequences in the PIR-International Protein Sequence Database, to spread and standardize protein names, features and keywords among members of a family or superfamily, and to aid us in classifying sequences, in identifying conserved regions, and in defining new homology domains. RESULTS: Release 22.0, (December 1998), of the PIR-ALN database contains a total of 3806 alignments, including 1303 superfamily, 2131 family and 372 homology domain alignments. This is an appropriate dataset to develop and extract patterns, test profiles, train neural networks or build Hidden Markov Models (HMMs). These alignments can be used to standardize and spread annotation to newer members by homology, as well as to understand the modular architecture of multidomain proteins. PIR-ALN includes 529 alignments that can be used to develop patterns not represented in PROSITE, Blocks, PRINTS and Pfam databases. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. AVAILABILITY: PIR-ALN is currently being distributed as a single ASCII text file along with the title, member, species, superfamily and keyword indexes. The quarterly and weekly updates can be accessed via the WWW at pir.georgetown.edu. The quarterly updates can also be obtained by anonymous FTP from the PIR FTP site at NBRF.Georgetown.edu, directory [ANONYMOUS.PIR.ALIGNMENT].

Amino Acid Sequence↗

Object-oriented parsing of biological databases with Python.

MOTIVATION: While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way. RESULTS: We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.

Databases, Factual↗

Thermodynamic database for protein-nucleic acid interactions (ProNIT).

MOTIVATION: Protein-nucleic acid interactions are fundamental to the regulation of gene expression. In order to elucidate the molecular mechanism of protein-nucleic acid recognition and analyze the gene regulation network, not only structural data but also quantitative binding data are necessary. Although there are structural databases for proteins and nucleic acids, there exists no database for their experimental binding data. Thus, we have developed a Thermodynamic Database for Protein-Nucleic Acid Interactions (ProNIT). RESULTS: We have collected experimentally observed binding data from the literature. ProNIT contains several important thermodynamic data for protein-nucleic acid binding, such as dissociation constant (K(d)), association constant (K(a)), Gibbs free energy change (DeltaG), enthalpy change (DeltaH), heat capacity change (DeltaC(p)), experimental conditions, structural information of proteins, nucleic acids and the complex, and literature information. These data are integrated into a relational database system together with structural and functional information to provide flexible searching facilities by using combinations of various terms and parameters. A www interface allows users to search for data based on various conditions, with different display and sorting options, and to visualize molecular structures and their interactions. AVAILABILITY: ProNIT is freely accessible at the URL http://www.rtc.riken.go.jp/jouhou/pronit/pronit.html.

Amino Acid Sequence↗

Clustering of highly homologous sequences to reduce the size of large protein databases.

We present a fast and flexible program for clustering large protein databases at different sequence identity levels. It takes less than 2 h for the all-against-all sequence comparison and clustering of the non-redundant protein database of over 560,000 sequences on a high-end PC. The output database, including only the representative sequences, can be used for more efficient and sensitive database searches.

Algorithms↗

LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures.

A database comprising all ligand-binding sites of known structure aligned with all related protein sequences and structures is described. Currently, the database contains approximately 50000 ligand-binding sites for small molecules found in the Protein Data Bank (PDB). The structure-structure alignments are obtained by the Combinatorial Extension (CE) program (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998) and sequence-structure alignments are extracted from the ModBase database of comparative protein structure models for all known protein sequences (Sanchez et al., Nucleic Acids Res., 28, 250-253, 2000). It is possible to search for binding sites in LigBase by a variety of criteria. LigBase reports summarize ligand data including relevant structural information from the PDB file, such as ligand type and size, and contain links to all related protein sequences in the TrEMBL database. Residues in the binding sites are graphically depicted for comparison with other structurally defined family members. LigBase provides a resource for the analysis of families of related binding sites.

Binding Sites↗

SST: an algorithm for finding near-exact sequence matches in time proportional to the logarithm of the database size.

MOTIVATION: Searches for near exact sequence matches are performed frequently in large-scale sequencing projects and in comparative genomics. The time and cost of performing these large-scale sequence-similarity searches is prohibitive using even the fastest of the extant algorithms. Faster algorithms are desired. RESULTS: We have developed an algorithm, called SST (Sequence Search Tree), that searches a database of DNA sequences for near-exact matches, in time proportional to the logarithm of the database size n. In SST, we partition each sequence into fragments of fixed length called 'windows' using multiple offsets. Each window is mapped into a vector of dimension 4(k) which contains the frequency of occurrence of its component k-tuples, with k a parameter typically in the range 4-6. Then we create a tree-structured index of the windows in vector space, with tree-structured vector quantization (TSVQ). We identify the nearest neighbors of a query sequence by partitioning the query into windows and searching the tree-structured index for nearest-neighbor windows in the database. When the tree is balanced this yields an O(logn) complexity for the search. This complexity was observed in our computations. SST is most effective for applications in which the target sequences show a high degree of similarity to the query sequence, such as assembling shotgun sequences or matching ESTs to genomic sequence. The algorithm is also an effective filtration method. Specifically, it can be used as a preprocessing step for other search methods to reduce the complexity of searching one large database against another. For the problem of identifying overlapping fragments in the assembly of 120 000 fragments from a 1.5 megabase genomic sequence, SST is 15 times faster than BLAST when we consider both building and searching the tree. For searching alone (i.e. after building the tree index), SST 27 times faster than BLAST. AVAILABILITY: Request from the authors.

Algorithms↗

siRecords: an extensive database of mammalian siRNAs with efficacy ratings.

UNLABELLED: Short interfering RNAs (siRNAs) have been gaining popularity as the gene knock-down tool of choice by many researchers because of the clean nature of their workings as well as the technical simplicity and cost efficiency in their applications. We have constructed siRecords, a database of siRNAs experimentally tested by researchers with consistent efficacy ratings. This database will help siRNA researchers develop more reliable siRNA design rules; in the mean time, siRecords will benefit experimental researchers directly by providing them with information about the siRNAs that have been experimentally tested against the genes of their interest. Currently, more than 4100 carefully annotated siRNA sequences obtained from more than 1200 published siRNA studies are hosted in siRecords. This database will continue to expand as more experimentally tested siRNAs are published. AVAILABILITY: The siRecords database can be accessed at http://siRecords.umn.edu/siRecords/

Abstracting and Indexing↗

ZooDDD: a cross-species database for digital differential display analysis.

UNLABELLED: In this article, we combined EST information from the UniGene database and orthologous relationships from the Ensembl database to construct a ZooDDD database. The primary function of ZooDDD is to mine evolutionary conserved, highly expressed, tissue-specific orthologues in model animals. The candidate genes of interest derived from the ZooDDD database will provide biologists with a good step for comparing the expression, functions and evolution of animal genomes. AVAILABILITY: http://bio301.iis.sinica.edu.tw/~ZooDDDNew/main.php.

Base Sequence↗

Differential profiles of genes expressed in neonatal brain of 129X1/SvJ and C57BL/6J mice: A database to aid in analyzing DNA microarrays using nonisogenic gene-targeted mice.

Strain-specific differences in gene expression have been observed among various inbred mouse strains. Two strains that are commonly used in gene-targeting research today are the 129 substrains, which are used to produce ES cell lines, and C57BL/6J, which is used for the extensive backcrosses required to produce isogenic knockout mice. When F2 nonisogenic littermates are assessed using DNA microarrays, one must determine whether the expression profiles obtained resulted either from specific alteration(s) induced by the targeted gene mutation or from gene expression differences related to the genetic background of the parent mouse strains. In the present study, we report the differential expression profile of genes expressed in neonatal brains and adult spleen and liver of 129X1/SvJ and C57BL/6J strains of mice. These comprehensive profiles were assessed using two types of Agilent Mouse Oligo Microarrays (development and standard) and were compiled into a publicly available database. Researchers can use this database to determine whether their microarray findings represent strain-specific differences in gene expression by comparing their data with those cataloged in our database. This database is useful for effectively analyzing DNA microarray data from nonisogenic littermates, and would help researchers avoid time-consuming backcrosses and confirmatory experiments requiring the use of many mice.

Animals↗

Characteristics of the SAGE database: a new resource for research on outcomes in long-term care. SAGE (Systematic Assessment of Geriatric drug use via Epidemiology) Study Group.

BACKGROUND: Because there is a lack of databases specific to long-term care, standardized assessments of nursing home residents are seen as a potential new resource for studying an important but neglected population. We describe the design and principal population characteristics of the first integrated database combining detailed clinical information and administrative claims data. METHODS: We studied nearly 300,000 residents admitted between 1992 and 1994 to all Medicare/Medicaid certified nursing homes of five U.S. states (Kansas, Maine, Mississippi, New York, and South Dakota). The database crosslinks: (a) Resident Data: over 350 items (demographic, diagnostic, clinical, and treatments) collected with the Minimum Data Set; (b) Drug Data: brand name, dosage route, and frequency of administration for all drugs consumed by each resident; (c) Medicare Data: eligibility and inpatient hospital claims; (d) Facilities Data: structural and staffing information on nursing homes; and (e) Country Data: information on population, health professions and facility data, and economic parameters. RESULTS: Ninety-two percent of the residents were aged 65 years and older. Residents were predominantly white (85%) and female (72%). The average number of medical diagnoses was above three, and residents were receiving an average of six medications. Sixty-five percent of residents had at least one hospital claim following the initial assessment, most commonly related to cardiovascular diseases and metabolic disorders. Fifty-five percent of the facilities were for-profit and 33% were of small size. Quality indicators and staffing level varied significantly by state. CONCLUSIONS: The SAGE (Systematic Assessment of Geriatric drug use via Epidemiology) database provides a unique resource to study the relation between treatments received and outcomes experienced, particularly functional and health services outcomes, that have not been possible before in very old, frail people.

Activities of Daily Living↗

Archive of mass spectral data files on recordable CD-ROMs and creation and maintenance of a searchable computerized database.

A database containing names of mass spectral data files generated in a forensic toxicology laboratory and two Microsoft Visual Basic programs to maintain and search this database is described. The data files (approximately 0.5 KB/each) were collected from six mass spectrometers during routine casework. Data files were archived on 650 MB (74 min) recordable CD-ROMs. Each recordable CD-ROM was given a unique name, and its list of data file names was placed into the database. The present manuscript describes the use of search and maintenance programs for searching and routine upkeep of the database and creation of CD-ROMs for archiving of data files.

CD-ROM↗

DNA Methylation Database "MethDB": a user guide.

The DNA Methylation Database (MethDB; http://www.methdb.net) is a public database dedicated to DNA methylation. It attempts to store all data about DNA methylation in a common source. MethDB can be searched in different ways, ranging from a simple browse mode to detailed queries for sequence-specific DNA methylation profiles and patterns, total methylation content data or environmental conditions that can influence the methylation state of DNA. Currently, the database contains 2570 values for methylation content and 5278 methylation patterns and profiles for a total of 83 genes or loci. MethDB is an annotated database, and the scientific community is invited to submit their own data (submit@methdb.de).

Computer Graphics↗

Obtaining maximal concatenated phylogenetic data sets from large sequence databases.

To improve the accuracy of tree reconstruction, phylogeneticists are extracting increasingly large multigene data sets from sequence databases. Determining whether a database contains at least k genes sampled from at least m species is an NP-complete problem. However, the skewed distribution of sequences in these databases permits all such data sets to be obtained in reasonable computing times even for large numbers of sequences. We developed an exact algorithm for obtaining the largest multigene data sets from a collection of sequences. The algorithm was then tested on a set of 100,000 protein sequences of green plants and used to identify the largest multigene ortholog data sets having at least 3 genes and 6 species. The distribution of sizes of these data sets forms a hollow curve, and the largest are surprisingly small, ranging from 62 genes by 6 species, to 3 genes by 65 species, with more symmetrical data sets of around 15 taxa by 15 genes. These upper bounds to sequence concatenation have important implications for building the tree of life from large sequence databases.

Algorithms↗

Status of the transcription factors database (TFD).

The Transcription Factors Database is a specialized database focusing on transcription factors and their properties. This report describes the present status of this database and developments during the past year. Within this time, the size of this database has increased by a 2799 total records, and has become accessible through a number of new mechanisms.

Databases, Factual↗

ECD--a totally integrated database of Escherichia coli K12.

We have compiled the DNA sequence data for E. coli available from the GENBANK and EMBL data libraries and independently from the literature. Starting with this update of our Escherichia coli database (ECD release 20) we provide major changes compared to previous issues. This update not only represents another substantial increase in sequence information, it also allows now to find the exact physical location of each individual gene or regulatory region, even regarding discrepancies in nomenclature. In order to save space this printed version does not contain the database itself anymore, but we provide several examples. The complete database is publically available in electronic form together with a self explaining application program or as a flat file. The complete compilation including a full set of genetic map data and the E. coli protein index can be obtained in machine readable form from the EMBL data library as a part of the CD-ROM issue of the EMBL sequence database, released and updated every three months. After deletion of all detected overlaps a total of 2,878,364 individual bp is found to be determined till the end of June 1994. This corresponds to a total of 60.98% of the entire E. coli chromosome consisting of about 4,720 kbp. This number may actually be higher by 9161 bp derived from other strains of E. coli.

Base Sequence↗

The translational termination signal database (TransTerm) now also includes initiation contexts.

The TransTerm database of termination codon contexts has been extended to include sense codon usage, and initiation codon contexts. The database was constructed from 23,721 coding sequences from 93 organisms. The database contains: a) the sequence around the termination codon (-10, +10); b) the sequence around the initiation codon (-20, +10); c) the length, 'G+C%' of the third position of codons (GC3), the 'codon adaptation index' (CAI) and the 'effective number of codons' statistic (Nc); d) summary tables for each organism including total codon usage, stop codon and tetranucleotide stop-signal usage, and matrices tallying base frequencies at each position around the initiation and termination codons. The data are arranged to facilitate investigation of the relationships between the three phases of protein synthesis. The database is available electronically from EMBL.

Animals↗

The haemophilia A mutation search test and resource site, home page of the factor VIII mutation database: HAMSTeRS.

In order to facilitate easy access to and aid understanding of the causes of haemophilia A at the molecular level we have constructed HAMSTeRS, the third release of the factor VIII mutation database and the first release of this database that may be accessed and interrogated over the internet through a World Wide Web browser. The database also presents a review of the structure and function of factor VIII and the molecular genetics of haemophilia A, a real time update of the biostatistics of each parameter in the database, a molecular model of the A1, A2 and A3 domains of the factor VIII protein (based on the crystal structure of caeruloplasmin) and a bulletin board for discussion of issues in the molecular biology of factor VIII.

Databases, Factual↗

The PIR-International Protein Sequence Database.

From its origin the Protein Sequence Database has been designed to support research and has focused on comprehensive coverage, quality control and organization of the data in accordance with biological principles. Since 1988 the database has been maintained collaboratively within the framework of PIR-International, an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. The database is widely distributed and is available on the World Wide Web, via ftp, email server, on CD-ROM and magnetic media. It is widely redistributed and incorporated into many other protein sequence data compilations, including SWISS-PROT and the Entrez system of the NCBI.

Amino Acid Sequence↗