Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

COPASAAR--a database for proteomic analysis of single amino acid repeats.

BACKGROUND: Single amino acid repeats make up a significant proportion in all of the proteomes that have currently been determined. They have been shown to be functionally and medically significant, and are associated with cancers and neuro-degenerative diseases such as Huntington's Chorea, where a poly-glutamine repeat is responsible for causing the disease. The COPASAAR database is a new tool to facilitate the rapid analysis of single amino acid repeats at a proteome level. The database aims to simplify the comparison of repeat distributions between proteomes in order to provide a better understanding of their function and evolution. RESULTS: A comparative analysis of all proteomes in the database (currently 244) shows that single amino acid repeats account for about 12-14% of the proteome of any given species. They are more common in eukaryotes (14%) than in either archaea or bacteria (both 13%). Individual analyses of proteomes show that long single amino acid repeats (6+ residues) are much more common in the Eukaryotes and that longer repeats are usually made up of hydrophilic amino acids such as glutamine, glutamic acid, asparagine, aspartic acid and serine. CONCLUSION: COPASAAR is a useful tool for comparative proteomics that provides rapid access to amino acid repeat data that can be readily data-mined. The COPASAAR database can be queried at the kingdom, proteome or individual protein level. As the amount of available proteome data increases this will be increasingly important in order to automate proteome comparison. The insights gained from these studies will give a better insight into the evolution of protein sequence and function.

Algorithms↗

PeanutMap: an online genome database for comparative molecular maps of peanut.

BACKGROUND: Molecular maps have been developed for many species, and are of particular importance for varietal development and comparative genomics. However, despite the existence of multiple sets of linkage maps, databases of these data are lacking for many species, including peanut. DESCRIPTION: PeanutMap http://peanutgenetics.tamu.edu/cmap provides a web-based interface for viewing specific linkage groups of a map set. PeanutMap can display and compare multiple maps of a set based upon marker or trait correspondences, which is particularly important as cultivated peanut is a disomic tetraploid. The database can also compare linkage groups among multiple map sets, allowing identification of corresponding linkage groups from results of different research projects. Data from the two published peanut genome map sets, and also from three maps sets of phenotypic traits are present in the database. Data from PeanutMap have been incorporated into the Legume Information System website http://www.comparative-legumes.org to allow peanut map data to be used for cross-species comparisons. CONCLUSION: The utility of the database is expected to increase as several SSR-based maps are being developed currently, and expanded efforts for comparative mapping of legumes are underway. Optimal use of these data will benefit from the development of tools to facilitate comparative analysis.

Arachis↗

RiboaptDB: a comprehensive database of ribozymes and aptamers.

BACKGROUND: Catalytic RNA molecules are called ribozymes. The aptamers are DNA or RNA molecules that have been selected from vast populations of random sequences, through a combinatorial approach known as SELEX. The selected oligo-nucleotide sequences (~200 bp in length) have the ability to recognize broad range of specific ligands by forming binding pockets. These novel aptamer sequences can bind to nucleic acids, proteins or small organic and inorganic chemical compounds and have many potential uses in medicine and technology. RESULTS: The comprehensive sequence information on aptamers and ribozymes that have been generated by in vitro selection methods are included in this RiboaptDB database. Such types of unnatural data generated by in vitro methods are not available in the public 'natural' sequence databases such as GenBank and EMBL. The amount of sequence data generated by in vitro selection experiments has been accumulating exponentially. There are 370 artificial ribozyme sequences and 3842 aptamer sequences in the total 4212 sequences from 423 citations in this RiboaptDB. We included general search feature, and individual feature wise search, user submission form for new data through online and also local BLAST search. CONCLUSION: This database, besides serving as a storehouse of sequences that may have diagnostic or therapeutic utility in medicine, provides valuable information for computational and theoretical biologists. The RiboaptDB is extremely useful for garnering information about in vitro selection experiments as a whole and for better understanding the distribution of functional nucleic acids in sequence space. The database is updated regularly and is publicly available at http://mfgn.usm.edu/ebl/riboapt/.

Aptamers, Nucleotide↗

i-Genome: a database to summarize oligonucleotide data in genomes.

BACKGROUND: Information on the occurrence of sequence features in genomes is crucial to comparative genomics, evolutionary analysis, the analyses of regulatory sequences and the quantitative evaluation of sequences. Computing the frequencies and the occurrences of a pattern in complete genomes is time-consuming. RESULTS: The proposed database provides information about sequence features generated by exhaustively computing the sequences of the complete genome. The repetitive elements in the eukaryotic genomes, such as LINEs, SINEs, Alu and LTR, are obtained from Repbase. The database supports various complete genomes including human, yeast, worm, and 128 microbial genomes. CONCLUSIONS: This investigation presents and implements an efficiently computational approach to accumulate the occurrences of the oligonucleotides or patterns in complete genomes. A database is established to maintain the information of the sequence features, including the distributions of oligonucleotide, the gene distribution, the distribution of repetitive elements in genomes and the occurrences of the oligonucleotides. The database can provide more effective and efficient way to access the repetitive features in genomes.

Alu Elements↗

GOLD.db: genomics of lipid-associated disorders database.

BACKGROUND: The GOLD.db (Genomics of Lipid-Associated Disorders Database) was developed to address the need for integrating disparate information on the function and properties of genes and their products that are particularly relevant to the biology, diagnosis management, treatment, and prevention of lipid-associated disorders. DESCRIPTION: The GOLD.db http://gold.tugraz.at provides a reference for pathways and information about the relevant genes and proteins in an efficiently organized way. The main focus was to provide biological pathways with image maps and visual pathway information for lipid metabolism and obesity-related research. This database provides also the possibility to map gene expression data individually to each pathway. Gene expression at different experimental conditions can be viewed sequentially in context of the pathway. Related large scale gene expression data sets were provided and can be searched for specific genes to integrate information regarding their expression levels in different studies and conditions. Analytic and data mining tools, reagents, protocols, references, and links to relevant genomic resources were included in the database. Finally, the usability of the database was demonstrated using an example about the regulation of Pten mRNA during adipocyte differentiation in the context of relevant pathways. CONCLUSIONS: The GOLD.db will be a valuable tool that allow researchers to efficiently analyze patterns of gene expression and to display them in a variety of useful and informative ways, allowing outside researchers to perform queries pertaining to gene expression results in the context of biological processes and pathways.

Adipocytes↗

angaGEDUCI: Anopheles gambiae gene expression database with integrated comparative algorithms for identifying conserved DNA motifs in promoter sequences.

BACKGROUND: The completed sequence of the Anopheles gambiae genome has enabled genome-wide analyses of gene expression and regulation in this principal vector of human malaria. These investigations have created a demand for efficient methods of cataloguing and analyzing the large quantities of data that have been produced. The organization of genome-wide data into one unified database makes possible the efficient identification of spatial and temporal patterns of gene expression, and by pairing these findings with comparative algorithms, may offer a tool to gain insight into the molecular mechanisms that regulate these expression patterns. DESCRIPTION: We provide a publicly-accessible database and integrated data-mining tool, angaGEDUCI, that unifies 1) stage- and tissue-specific microarray analyses of gene expression in An. gambiae at different developmental stages and temporal separations following a bloodmeal, 2) functional gene annotation, 3) genomic sequence data, and 4) promoter sequence comparison algorithms. The database can be used to study genes expressed in particular stages, tissues, and patterns of interest, and to identify conserved promoter sequence motifs that may play a role in the regulation of such expression. The database is accessible from the address http://www.angaged.bio.uci.edu. CONCLUSION: By combining gene expression, function, and sequence data with integrated sequence comparison algorithms, angaGEDUCI streamlines spatial and temporal pattern-finding and produces a straightforward means of developing predictions and designing experiments to assess how gene expression may be controlled at the molecular level.

Algorithms↗

LINE FUSION GENES: a database of LINE expression in human genes.

BACKGROUND: Long Interspersed Nuclear Elements (LINEs) are the most abundant retrotransposons in humans. About 79% of human genes are estimated to contain at least one segment of LINE per transcription unit. Recent studies have shown that LINE elements can affect protein sequences, splicing patterns and expression of human genes. DESCRIPTION: We have developed a database, LINE FUSION GENES, for elucidating LINE expression throughout the human gene database. We searched the 28,171 genes listed in the NCBI database for LINE elements and analyzed their structures and expression patterns. The results show that the mRNA sequences of 1,329 genes were affected by LINE expression. The LINE expression types were classified on the basis of LINEs in the 5' UTR, exon or 3' UTR sequences of the mRNAs. Our database provides further information, such as the tissue distribution and chromosomal location of the genes, and the domain structure that is changed by LINE integration. We have linked all the accession numbers to the NCBI data bank to provide mRNA sequences for subsequent users. CONCLUSION: We believe that our work will interest genome scientists and might help them to gain insight into the implications of LINE expression for human evolution and disease. AVAILABILITY: http://www.primate.or.kr/line.

Chromosome Mapping↗

RINGdb: an integrated database for G protein-coupled receptors and regulators of G protein signaling.

BACKGROUND: Many marketed therapeutic agents have been developed to modulate the function of G protein-coupled receptors (GPCRs). The regulators of G-protein signaling (RGS proteins) are also being examined as potential drug targets. To facilitate clinical and pharmacological research, we have developed a novel integrated biological database called RINGdb to provide comprehensive and organized RGS protein and GPCR information. RESULTS: RINGdb contains information on mutations, tissue distributions, protein-protein interactions, diseases/disorders and other features, which has been automatically collected from the Internet and manually extracted from the literature. In addition, RINGdb offers various user-friendly query functions to answer different questions about RGS proteins and GPCRs such as their possible contribution to disease processes, the putative direct or indirect relationship between RGS proteins and GPCRs. RINGdb also integrates organized database cross-references to allow users direct access to detailed information. The database is now available at http://ringdb.csie.ncu.edu.tw/ringdb/. CONCLUSION: RINGdb is the only integrated database on the Internet to provide comprehensive RGS protein and GPCR information. This knowledge base will be useful for clinical research, drug discovery and GPCR signaling pathway research.

Amino Acid Sequence↗

Mycobacterium tuberculosis complex genetic diversity: mining the fourth international spoligotyping database (SpolDB4) for classification, population genetics and epidemiology.

BACKGROUND: The Direct Repeat locus of the Mycobacterium tuberculosis complex (MTC) is a member of the CRISPR (Clustered regularly interspaced short palindromic repeats) sequences family. Spoligotyping is the widely used PCR-based reverse-hybridization blotting technique that assays the genetic diversity of this locus and is useful both for clinical laboratory, molecular epidemiology, evolutionary and population genetics. It is easy, robust, cheap, and produces highly diverse portable numerical results, as the result of the combination of (1) Unique Events Polymorphism (UEP) (2) Insertion-Sequence-mediated genetic recombination. Genetic convergence, although rare, was also previously demonstrated. Three previous international spoligotype databases had partly revealed the global and local geographical structures of MTC bacilli populations, however, there was a need for the release of a new, more representative and extended, international spoligotyping database. RESULTS: The fourth international spoligotyping database, SpolDB4, describes 1939 shared-types (STs) representative of a total of 39,295 strains from 122 countries, which are tentatively classified into 62 clades/lineages using a mixed expert-based and bioinformatical approach. The SpolDB4 update adds 26 new potentially phylogeographically-specific MTC genotype families. It provides a clearer picture of the current MTC genomes diversity as well as on the relationships between the genetic attributes investigated (spoligotypes) and the infra-species classification and evolutionary history of the species. Indeed, an independent Naïve-Bayes mixture-model analysis has validated main of the previous supervised SpolDB3 classification results, confirming the usefulness of both supervised and unsupervised models as an approach to understand MTC population structure. Updated results on the epidemiological status of spoligotypes, as well as genetic prevalence maps on six main lineages are also shown. Our results suggests the existence of fine geographical genetic clines within MTC populations, that could mirror the passed and present Homo sapiens sapiens demographical and mycobacterial co-evolutionary history whose structure could be further reconstructed and modelled, thereby providing a large-scale conceptual framework of the global TB Epidemiologic Network. CONCLUSION: Our results broaden the knowledge of the global phylogeography of the MTC complex. SpolDB4 should be a very useful tool to better define the identity of a given MTC clinical isolate, and to better analyze the links between its current spreading and previous evolutionary history. The building and mining of extended MTC polymorphic genetic databases is in progress.

Computational Biology↗

The Subviral RNA Database: a toolbox for viroids, the hepatitis delta virus and satellite RNAs research.

BACKGROUND: Viroids, satellite RNAs, satellites viruses and the human hepatitis delta virus form the 'brotherhood' of the smallest known infectious RNA agents, known as the subviral RNAs. For most of these species, it is generally accepted that characteristics such as cell movement, replication, host specificity and pathogenicity are encoded in their RNA sequences and their resulting RNA structures. Although many sequences are indexed in publicly available databases, these sequence annotation databases do not provide the advanced searches and data manipulation capability for identifying and characterizing subviral RNA motifs. DESCRIPTION: The Subviral RNA database is a web-based environment that facilitates the research and analysis of viroids, satellite RNAs, satellites viruses, the human hepatitis delta virus, and related RNA sequences. It integrates a large number of Subviral RNA sequences, their respective RNA motifs, analysis tools, related publication links and additional pertinent information (ex. links, conferences, announcements), allowing users to efficiently retrieve and analyze relevant information about these small RNA agents. CONCLUSION: With its design, the Subviral RNA Database could be considered as a fundamental building block for the study of these related RNAs. It is freely available via a web browser at the URL: http://subviral.med.uottawa.ca.

Base Sequence↗

Incomplete evidence: the inadequacy of databases in tracing published adverse drug reactions in clinical trials.

BACKGROUND: We would expect information on adverse drug reactions in randomised clinical trials to be easily retrievable from specific searches of electronic databases. However, complete retrieval of such information may not be straightforward, for two reasons. First, not all clinical drug trials provide data on the frequency of adverse effects. Secondly, not all electronic records of trials include terms in the abstract or indexing fields that enable us to select those with adverse effects data. We have determined how often automated search methods, using indexing terms and/or textwords in the title or abstract, would fail to retrieve trials with adverse effects data. METHODS: We used a sample set of 107 trials known to report frequencies of adverse drug effects, and measured the proportion that (i) were not assigned the appropriate adverse effects indexing terms in the electronic databases, and (ii) did not contain identifiable adverse effects textwords in the title or abstract. RESULTS: Of the 81 trials with records on both MEDLINE and EMBASE, 25 were not indexed for adverse effects in either database. Twenty-six trials were indexed in one database but not the other. Only 66 of the 107 trials reporting adverse effects data mentioned this in the abstract or title of the paper. Simultaneous use of textword and indexing terms retrieved only 82/107 (77%) papers. CONCLUSIONS: Specific search strategies based on adverse effects textwords and indexing terms will fail to identify nearly a quarter of trials that report on the rate of drug adverse effects.

Abstracting and Indexing↗

The Latin American Social Medicine database.

BACKGROUND: Public health practitioners and researchers for many years have been attempting to understand more clearly the links between social conditions and the health of populations. Until recently, most public health professionals in English-speaking countries were unaware that their colleagues in Latin America had developed an entire field of inquiry and practice devoted to making these links more clearly understood. The Latin American Social Medicine (LASM) database finally bridges this previous gap. DESCRIPTION: This public health informatics case study describes the key features of a unique information resource intended to improve access to LASM literature and to augment understanding about the social determinants of health. This case study includes both quantitative and qualitative evaluation data. Currently the LASM database at The University of New Mexico http://hsc.unm.edu/lasm brings important information, originally known mostly within professional networks located in Latin American countries to public health professionals worldwide via the Internet. The LASM database uses Spanish, Portuguese, and English language trilingual, structured abstracts to summarize classic and contemporary works. CONCLUSION: This database provides helpful information for public health professionals on the social determinants of health and expands access to LASM.

Databases, Bibliographic↗

Food composition database development for between country comparisons.

Nutritional assessment by diet analysis is a two-stepped process consisting of evaluation of food consumption, and conversion of food into nutrient intake by using a food composition database, which lists the mean nutritional values for a given food portion. Most reports in the literature focus on minimizing errors in estimation of food consumption but the selection of a specific food composition table used in nutrient estimation is also a source of errors. We are conducting a large prospective study internationally and need to compare diet, assessed by food frequency questionnaires, in a comparable manner between different countries. We have prepared a multi-country food composition database for nutrient estimation in all the countries participating in our study. The nutrient database is primarily based on the USDA food composition database, modified appropriately with reference to local food composition tables, and supplemented with recipes of locally eaten mixed dishes. By doing so we have ensured that the units of measurement, method of selection of foods for testing, and assays used for nutrient estimation are consistent and as current as possible, and yet have taken into account some local variations. Using this common metric for nutrient assessment will reduce differential errors in nutrient estimation and improve the validity of between-country comparisons.

Agriculture↗

Patient-reported outcome and quality of life instruments database (PROQOLID): frequently asked questions.

The exponential development of Patient-Reported Outcomes (PRO) measures in clinical research has led to the creation of the Patient-Reported Outcome and Quality of Life Instruments Database (PROQOLID) to facilitate the selection process of PRO measures in clinical research. The project was initiated by Mapi Research Trust in Lyon, France. Initially called QOLID (Quality of Life Instruments Database), the project's purpose was to provide all those involved in health care evaluation with a comprehensive and unique source of information on PRO and HRQOL measures available through the Internet.PROQOLID currently describes more than 470 PRO instruments in a structured format. It is available in two levels, non-subscribers and subscribers, at http://www.proqolid.org. The first level is free of charge and contains 14 categories of basic useful information on the instruments (e.g. author, objective, original language, list of existing translations, etc.). The second level provides significantly more information about the instruments. It includes review copies of over 350 original instruments, 120 user manuals and 350 translations. Most are available in PDF format. This level is only accessible to annual subscribers. PROQOLID is updated in close collaboration with the instruments' authors on a regular basis. Fifty or more new instruments are added to the database annually.Today, all of the major pharmaceutical companies, prestigious institutions (such as the FDA, the NIH's National Cancer Institute, the U.S. Veterans Administration), dozens of universities, public institutions and researchers subscribe to PROQOLID on a yearly basis. More than 800 users per day routinely visit the database.

Adult↗

A UK general practice database study of prevalence and mortality of people with neural tube defects.

OBJECTIVE: To investigate the prevalence of neural tube defects (NTDs) in the UK and to compare the mortality rate with that of the general population. METHODS: A cross-sectional study. The General Practice Research Database (GPRD) contains the prescribing and diagnostic records since 1990 of over 4 million people from throughout the UK. All patients aged 10-69 and registered on the database in the years 1994-1997 were included in the study. Patients with a diagnosis of NTD were identified from the database and prevalence and standardized mortality ratios in each year were calculated. RESULTS: The size of the GPRD reduced during the study period - there were 2116452 patients aged 10-69 years on the database in 1994, of whom 1751 had a prior record of NTD. In 1997 there were 998368 patients, of whom 842 had an NTD. The age standardized prevalence between 1994 and 1997 for NTDs ranged between 7.8 and 8.4 per 10000 for males and 9.0 and 9.4 per 10000 for females aged 10-69 years. There were 27 deaths in patients with a record of NTD over the four-year study period. The standardized mortality ratio for the years 1994 to 1997 for NTDs ranged between 1.9 and 2.9. CONCLUSIONS: These data give an estimate of the prevalence of NTDs in the general population. They also show that those who have survived to age 10 years still have double the mortality of the general population.

Adolescent↗

Worldwide Innovative Network Consortium: Building a Common Global Cancer Database.

This review shares the ongoing work of the global Worldwide Innovative Network (WIN) Consortium for Precision Medicine to synthesize emerging cancer treatment data and to define the requirements for a common global cancer database that can truly support precision oncology. We performed a narrative review of emerging cancer treatment data, molecular profiling technologies, and existing clinicogenomic databases, focusing on how tumors are characterized, how subgroups are defined, and how demographic, lifestyle, and environmental factors are captured. The growth in molecular profiling technologies and the development of new targeted therapies are transforming cancer care. Tumors, regardless of tissue origin, are increasingly defined as composites of multiple, often rare, subgroups, each with distinct biology and likely response to specific therapies, based on multidimensional profiling of the tumor and its microenvironment. The solution lies in building vast databases that capture racial and ethnic diversity, reflected in genomic data, as well as diet and lifestyle factors that may have epigenetic impact on gene expression and post-translational modifications. A truly inclusive and informative data set must reflect global diversity, and there are multiple examples of demography-dependent differences in genomic signals. With members caring for and studying patients with cancer across five continents, WIN is actively exploring pathways to create a global cancer database, rich in clinical and molecular detail, granular enough for precise analysis, and large enough to power artificial intelligence-driven insights, provided appropriate data quality, validation, and governance frameworks are in place. This review surveys the current landscape and outlines practical paths forward to achieve this goal.

Humans↗

Evolutionary relationships among G protein-coupled receptors using a clustered database approach.

Guanine nucleotide-binding protein-coupled receptors (GPCRs) comprise large and diverse gene families in fungi, plants, and the animal kingdom. GPCRs appear to share a common structure with 7 transmembrane segments, but sequence similarity is minimal among the most distant GPCRs. To reevaluate the question of evolutionary relationships among the disparate GPCR families, this study takes advantage of the dramatically increased number of cloned GPCRs. Sequences were selected from the National Center for Biotechnology Information (NCBI) nonredundant peptide database using iterative BLAST (Basic Local Alignment Search Tool) searches to yield a database of approximately 1700 GPCRs and unrelated membrane proteins as controls, divided into 34 distinct clusters. For each cluster, separate position-specific matrices were established to optimize sequence comparisons among GPCRs. This approach resulted in significant alignments between distant GPCR families, including receptors for the biogenic amine/peptide, VIP/secretin, cAMP, STE3/MAP3 fungal pheromones, latrophilin, developmental receptors frizzled and smoothened, as well as the more distant metabotrobic glutamate receptors, the STE2/MAM2 fungal pheromone receptors, and GPR1, a fungal glucose receptor. On the other hand, alignment scores between these recognized GPCR clades with p40 (putative GPCR) and pm1 (putative GPCR), as well as bacteriorhodopsins, failed to support a finding of homology. This study provides a refined view of GPCR ancestry and serves as a reference database with hyperlinks to other sources. Moreover, it may facilitate database annotation and the assignment of orphan receptors to GPCR families.

Amino Acid Sequence↗

The ALS patient care database: goals, design, and early results. ALS C.A.R.E. Study Group.

OBJECTIVE: The ALS Patient Care Database was created to improve the quality of care for patients with ALS by 1) providing neurologists with data to evaluate and improve their practices, 2) publishing data on temporal trends in the care of patients with ALS, and 3) developing hypotheses to be tested during formal clinical trials. BACKGROUND: Substantial variations exist in managing ALS, but there has been no North American database to measure outcomes in ALS until now. METHODS: This observational database is open to all neurologists practicing in North America, who are encouraged to enroll both incident and prevalent ALS patients. Longitudinal data are collected at intervals of 3 to 6 months by using standard data collection instruments. Forms are submitted to a central data coordinating center, which mails quarterly reports to participating neurologists. RESULTS: Beginning in September 1996 through November 30, 1998, 1,857 patients were enrolled at 83 clinical sites. On enrollment, patients had a mean age of 58.6 years +/-12.9 (SD) years (range, 20.1 to 95.1 years), 92% were white, and 61% were men. The mean interval between onset of symptoms and diagnosis was 1.2+/-1.6 years (range, 0 to 31.9 years). Riluzole was the most frequently used disease-specific therapy (48%). Physical therapy was the most common nonpharmacologic intervention (45%). The primary caregiver was generally the spouse (77%). Advance directives were in place at the time of death for 70% of 213 enrolled patients who were reported to have died. CONCLUSIONS: The ALS Patient Care Database appears to provide valuable data on physician practices and patient-focused outcomes in ALS.

Activities of Daily Living↗