Search PubMedSearch

SEARCH · Search PubMed

Results for “reference databases”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

The SWISS-PROT protein sequence data bank and its supplement TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, structure of its domains, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and the creation of TrEMBL, a computer annotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Academies and Institutes

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

Effect of growth hormone treatment in children with craniopharyngioma with reference to the KIGS (Kabi International Growth Study) database.

In children with craniopharyngioma, poor growth commonly precedes diagnosis, but is observed less frequently than neurological or visual symptoms. A deficiency of growth hormone (GH) is common before, and almost universal after, treatment of the tumour, and is usually treated with GH. However, a minority of these children with GH deficiency (GHD) grow well without GH replacement therapy but exhibit other metabolic effects of GHD that are correctable by GH treatment. This article provides a review of studies in 422 children with craniopharyngioma whose details have been entered into the database of KIGS, the Kabi International Growth Study. The response to GH during the first year of therapy was similar to that seen in children with idiopathic GHD (IGHD). Leg length was relatively greater than sitting height and this disproportion was maintained during treatment. Adiposity increased in some children receiving GH treatment. At the end of GH treatment in 82 patients, there was a median gain in height SD score of 1.51, with evidence of residual growth potential still remaining in the majority. Tumour recurrence occurred in 13.5% of the total group of patients with craniopharyngioma within KIGS, at a median of 3.9 years from diagnosis and 2.3 years from the start of GH therapy. Tumour recurrence was not associated with an impairment in height achieved, but there was a tendency towards greater adiposity in patients in whom recurrence occurred. Adverse events during GH treatment were more frequent in children with craniopharyngioma than in those with IGHD, and headache was commonly reported. The results of these studies suggest that GH treatment is recommended for the treatment of children with craniopharyngioma on the grounds of improved growth velocity, adult height and other GH-dependent metabolic functions, and of the good safety profile of GH in these patients.

Adipose Tissue

Standard method of diagnosis versus use of a computer database in the evaluation of skeletal dysplasias.

OBJECTIVE: The objective of this study was to compare reference textbooks and the computer database, OSSUM, for accuracy and ease of use in the diagnosis of skeletal dysplasias. Materials and methods. Twenty cases of clinically and and radiologically established skeletal dysplasias were evaluated as unknowns by four pediatric radiologists. Readers 1 and 2 evaluated group A (10 cases) using reference texts and group B (10 cases) using OSSUM. Readers 3 and 4 evaluated group B using reference texts. The radiologists independently listed their roentgenographic findings, the top three diagnoses, confidence level, difficulty level, and time spent on each case. RESULTS: The correct diagnosis was made in 68% of both the reference text cases and the OSSUM cases. Difficulty level was significantly higher (3.5 vs 2.9, P = 0.0013) and confidence significantly lower (3.3 vs. 2.3, P = 0.0001) when using OSSUM. Average time spent on cases was 25 min with references and 30 min with OSSUM (P > 0.05). However, there was a decrease in both the time (38 min vs 23 min, P = 0.05) and the difficulty (3.9 vs 3.1, P = 0.001) between the first five and the last five cases. The composite of four readers correctly identified 90% of the skeletal dysplasias when the results of both methods were combined. CONCLUSIONS: In the ability to reach a correct diagnosis, no difference was detected between the OSSUM and reference texts methods. The increased time necessary, greater difficulty and decreased confidence levels with OSSUM are expected to improve with increasing program familiarity. Use of both textbooks and the database was complementary.

Bone Diseases, Developmental

The genetics of tobacco-induced malignancy.

OBJECTIVE: Several areas of investigation contribute to an increasing understanding of the genetics of malignancies associated with tobacco use. While the strong influence of tobacco exposure on cancer development obscures genetic influences, there are indications that aspects of cancer susceptibility may have a heritable basis. In addition, specific sites within the genome appear to be commonly involved in these malignancies. This review describes research relevant to investigation of the genetics of tobacco-induced malignancy. DATA SOURCES: A review of the pertinent literature covered the past 20 years. References were gleaned from a variety of sources including manual review of the most recent journals, a computerized database (Mini-MEDLINE), references cited in previous works, and our own ongoing research in cancer genetics and molecular biology. STUDY SELECTION AND DATA EXTRACTION: Whenever possible, controlled studies from peer-reviewed journals were used. Where studies have shown conflicting results, the possible confounding factors are discussed. DATA SYNTHESIS: From a broad array of research areas, a view of the genetic aspects of tobacco carcinogenesis emerges. This includes syndromic and nonsyndromic susceptibility, genetic determinants of carcinogen metabolism, DNA adduct formation, and site-specific genetic alterations in tobacco-induced malignancy. In addition to specific gene alterations commonly seen in tobacco-induced malignancy, viral infections may contribute to cancer development through pathways related to tobacco carcinogenesis. CONCLUSION: Further research on many aspects of tobacco-induced carcinogenesis is warranted. Investigation of cancer susceptibility may contribute to understanding DNA surveillance and repair pathways. Carcinogen metabolism investigations have application in cancer detection and prevention schemes. Further understanding of tumor suppressor gene function and the role of gene amplification in carcinogenesis may allow design of gene-specific strategies in cancer treatment.

Carcinogens

Downloading from MEDLINE: a comparison of personal database software.

The compatibility of downloaded MEDLINE references with software packages designed for maintenance of personal bibliographic databases is considered. The nine packages reviewed enable one to import batches of records downloaded from one or more presentations of the MEDLINE database, online or CD-ROM; and enable the user to re-format records into the various styles required by journal editors. Some basic features of the packages are tabulated, and details of record-importing facilities are described. Details of suppliers and prices are given.

Catalogs, Commercial as Topic

The GASTER project: building a computer network in digestive endoscopy: the experience of the European Society for Gastrointestinal Endoscopy. Gastrointestinal Endoscopy Application for Standards in Telecommunication, Education and Research.

Digestive endoscopy is currently the main diagnostic procedure for investigation of the digestive tract whenever a digestive disease is suspected. From 1970 to 1985, digestive endoscopy was performed with endoscopes equipped with fiberoptic bundles, whereas the last decade was marked by the development of electronic endoscopes, characterized by the presence of a CCD (charge coupled device) at the tip of the endoscope. Thus the physician looks at a TV screen to control the procedure and examine in detail the gut wall. Endoscopes examine the foregut until the duodenum and the hindgut, up to the three last intestinal loops. When the endoscopic workstation comprises a computer, it is possible to acquire electronic images during the endoscopy and use these images as support of the information about the results of the procedure. These numeric images can then be stored in databases containing text attached to them. Starting with these images, one may expect many developments in the near future that will change the management of the patient with digestive diseases. Physicians will become able to exchange images and text related to one patient or one procedure, although they are equipped with different workstations. Therefore, it is obvious that the information exchanged must be written in a standard format that makes it understandable by all systems. The European Society of Gastrointestinal Endoscopy is a scientific society that groups most of the gastroenterologists in Europe. This society has initiated a research program to develop standards for the exchange of images and text. The Gastrointestinal Endoscopy Applications for Standards in Telecommunication, Education, and Research (GASTER) project intends to implement a multimedia database of endoscopic images based on a standard format of images and a standard terminology for descriptive terms. These standards must be validated by use in different endoscopy units. The database will collect images from these centers that will be linked to the coordinating center through a network based on an integrated services digital network (fast electronic connection). This database will then be used for the development of computer applications. The output of the GASTER project will bring advances at three levels: (1) The physicians will be able to exchange images about the procedures their patients have undergone and will thus obtain more complete information, improving quality of care. They will also benefit from help-to-decision applications based on validated reference images from the database. (2) At the patient level, the quality of care will be improved through a better dissemination of information between the physicians in charge of the patient, thus there is better follow-up of the patient and a decrease in redundant examinations. (3) At the level of national health care systems, the benefit will be a decrease in cost of care due to a better follow-up of the patients, a decrease in redundant examinations, and a faster decision made to treat the patient. The possibility of consulting a database of a scientifically validated images used as reference material will also improve quality control in digestive endoscopy.

Computer Communication Networks

Design of a diagnostic encyclopaedia using AIDA.

Diagnostic Encyclopaedia Workstation (DEW) is the name of a digital encyclopaedia constructed to contain reference knowledge with respect to the pathology of the ovary. Comparing DEW with the common sources of reference knowledge (i.e. books) leads to the following advantages of DEW: it contains more verbal knowledge, pictures and case histories, and it offers information adjusted to the needs of the user. Based on an analysis of the structure of this reference knowledge we have chosen AIDA to develop a relational database and we use a video-disc player to contain the pictorial part of the database. The system consists of a database input version and a read-only run version. The design of the database input version is discussed. Reference knowledge for ovary pathology requires 1-3 Mbytes of memory. At present 15% of this amount is available. The design of the run version is based on an analysis of which information must necessarily be specified to the system by the user to access a desired item of information. Finally, the use of AIDA in constructing DEW is evaluated.

Computer Systems

An evaluation of the TransFER model for sharing clinical decision-support applications.

TransFER is a formal model designed to facilitate the sharing of decision-support applications across institutions with heterogeneous clinical databases. The TransFER model provides a mechanism to automatically customize database queries based on a reference schema of clinical data and an encoded set of database mappings. In this paper, we describe the elements of the TransFER model and we present the results of a formal evaluation we conducted to assess the utility and generality of the model. The results suggest that the TransFER has significant potential for automating query translation and facilitating application sharing, but that further work on the representation of temporal semantics, on the modeling of missing data, and on the optimization of complex queries is required.

Decision Making, Computer-Assisted

Development of computerized storage facilities for twin data: a relational database system for a twin register.

Many twin registers hold information on flat file systems such as those provided by statistical packages or spreadsheets. Demographic details may be maintained separately from data collected in multiple different studies, leading to considerable problems with data consistency, redundancy, and integration. Ad hoc requests may be difficult. Implementation of a relational database system permits storage and maintenance of all records, simple data entry and validation procedures, linking of information from different projects with security of access, and the flexibility to provide rapid answers to ad hoc enquiries using standard Structured Query Language (SQL). Twin data provide a challenge for relational database design which rests on the technique of normalization and the use of unique identifiers to access associated groups of variables; for twins, "uniqueness" must preserve identification of both the pair and the individual twin subjects in the data structure to enable flexible access to and analysis of the data. An application on the Institute of Psychiatry Volunteer Twin Register (IOPVTR) database is described, through reference to one study of a sample of the twins, with simulated data. We show how a balance of adherence to database design principles and attention to ongoing clerical and research procedures has been used to produce an integrated, flexible, and open-ended system.

Data Collection

ENB--resource and careers department.

The English National Board for Nursing, Midwifery and Health Visiting is aware that, in the current climate of change, a range of information is required by nurses, midwives and health visitors to enable them to update their professional knowledge and plan their careers. In response to this need the Board offers a comprehensive range of services through the Resource and Careers Department based in Sheffield. I will describe in detail the services provided by the Resource Section and then give a brief overview of the role and functions of the other sections, i.e. Publications, Careers Service and Projects. The Resource Section comprises three separate but interactive parts. These are the Health Care Database, Open Learning Resource and Reference Room and ENB Campus.

Databases, Factual

Estimation of reference change limits using patient data.

Two approaches for deriving reference change limits from patient data are described. In the direct method, hospital database information is used for the selection of appropriate reference groups. If database information is not sufficient or reliable enough, but still most of the source data can be considered as health-related, an indirect method can be applied in the calculation of rough estimates for reference change limits. A computer program developed by us, GraphROC for Windows, includes both methods for the estimation of change limits from patient data. Time between specimen collections should be included as one classifying factor in the selection of source data. When only one previous result is available for comparison, change limits based on the reference sample group form the only available guide for clinical interpretation. However, when several previous results are available and the within-subject variances for the considered analyte are known to be heterogeneous between individuals, the clinical interpretation should rather be based on application of time series analysis.

Data Interpretation, Statistical

A multi-user networked database for analysis of clinical and temperature data from patients treated with simultaneous radiation and ultrasound hyperthermia.

A database was developed using commercially available development software that allows the entry of clinical data and automatically analyses temperature and power data from a commercial ultrasound hyperthermia system. The database can be accessed via network connections by more than one authorized user, thus facilitating the entry, management, and analysis of clinical data. The software automatically estimates ultrasound induced temperature artifacts and calculates thermal dose parameters such as T90s, equivalent minutes at 43 degrees, and time at or above index temperatures using the corrected temperatures. These parameters also become part of the database. Digital photographs of treatment setup, probe placement, and tumour or normal tissue response can be included in the database for documentation and reference. Ultrasound diagnostic images that document the depth and reproducibility of probe placement can be scanned into the PC and included in the database as well. This short communication documents experiences developing this tool that may be useful to other investigators.

Combined Modality Therapy

ERIC: a resource for researchers in nursing education.

This chapter provides information on the ERIC system of bibliographic information covering the field of education. Information on topics related to nursing and specifically to research in nursing education is presented. The number of references in these areas and in the ERIC database and the content of these references is described.

Bibliographies as Topic

Efficient literature searching in diffuse topics: lessons from a systematic review of research on communicating risk to patients in primary care.

Using the example of communication about risk in a primary care setting, this paper puts forward a method of developing and evaluating a detailed search strategy for locating the literature for a systematic review of a 'diffuse' subject. The aim of this paper is to show how to develop a search strategy that maximizes both recall and precision while keeping search outputs manageable. Six different databases were used, namely Medline, Embase, PsychLIT, CancerLIT, Cinahl and Social Science Citation Index (SSCI). The searches were augmented by hand-searching, contacting authors, citation searching and reference lists from included papers. Other databases were searched but yielded no extra references for this subject matter. Of the 99 papers included, 80 were indexed on Medline. The Medline search strategy identified 54 of them and the remaining 26 were located on other databases. The 19 further unique references were found using the other databases and methods of retrieval. A combination of several databases must be used to maximize recall and to increase the precision of searches on individual databases, thus improving the overall efficiency of the search.

Databases, Bibliographic

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV