Search PubMedSearch

SEARCH · Search PubMed

Results for “Reference database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Comparison of the effect of different reference data on Lunar DPX and Hologic QDR-1000 dual-energy X-ray absorptiometers.

We have investigated whether the Lunar DPX (software 3.4) and Hologic QDR-1000 dual-energy X-ray absorptiometers have comparable normal reference databases for the spine and femur of white UK and USA subjects. After conversion for systematic differences in absolute bone density values between the two systems, the reference databases were very similar for the spine in young subjects, but there were clear differences in the femur databases of young females and males of all ages. These differences were confirmed by comparing the percent age-matched and young values determined by the two systems for subjects scanned on both systems. Thus the diagnosis and management of a patient could differ, depending on the system used for the bone density measurements.

Absorptiometry, Photon

On dividing reference data into subgroups to produce separate reference ranges.

We consider statistical criteria for partitioning a reference database to obtain separate reference ranges for different subpopulations. Using general formulas relating population variances, sample sizes, and the normal deviate test for the significance of the difference between two subgroup means, we show that partitioning into separate ranges produces little reduction in between-person variability, even when the differences between means are highly significant statistically. However, when there is a clear physiological basis for distinguishing between certain subgroups, simulation studies show that partitioning may be necessary to obtain reference limits that cut off the desired proportions of low and high values in each subgroup. Guidelines based on these results are provided to help decide whether separate ranges should be obtained for a given analyte.

Analysis of Variance

Approaching the millennium: perinatal problems and software solutions.

Strategic planning for rational development of perinatal computing capabilities for the year 2000 should be driven by anticipated trends in (1) the health care business, (2) computer technology and (3) medicine, as well as (4) the needs of perinatal practitioners. In the USA, health care is the fastest growing segment of the economy. This will produce increasing attention from hardware and software developers, and vendors, and will lead to a proliferation of computing platforms, operating systems and specific medical application software. Desktop computers, already capable of 20 million instructions per second (MIPS) with massive storage capacities, will continue to evolve and fall in price. Increasingly, perinatologists will develop software packages to facilitate patient care in their own environments. All of these trends will lead to severe fragmentation in medical computing. Simultaneously, however, the need for integrated institutional computer-based data access for quality assurance and fiscal and operations management will increase. Perinatal care will be more regionalized, complex and rigorous with new clinical trial- and effectiveness research-based interventions, as well as molecular diagnosis and therapy. To practice appropriately, clinicians will need to be familiar with computer capabilities. Having been exposed to computer-aided instruction (CAI) at the undergraduate and postgraduate levels, they will except on-line access to detailed and accurate patient information with linkage to laboratory, radiology and other medical databases, as well as to reference databases, such as Medlines and the Oxford Database of Perinatal Trials. Artificial intelligence (AI) software may support perinatal decision making; computerized professional and facility billing will be available.

Forecasting

A profile for molecular biology databases and information resources.

This paper examines the requirements for building database management systems and multi-database information resources to support molecular biology research. The paper profiles the most important features of 16 integrated resources and 102 databases related to molecular biology research. The aspects surveyed in this paper include the nature of information in these databases, their sizes, update properties, cross-references, database management system heterogeneity, geographical distribution, data quality, use of temporal information and level of interpretation. The paper also comments on the access patterns to these databases. Since not all these aspects were available for all databases, specific comparisons sometimes compare fewer than the full 102 databases. Consequently, the same set of databases is not necessarily always being compared with respect to every aspect. The paper is organized primarily according to these comparison aspects and ends with some concluding remarks.

Databases, Bibliographic

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic

[Electroconvulsive therapy and benzodiazepine: antagonism or indifference? Review of the literature].

BACKGROUND: It is usually considered that the efficacy of electroconvulsive therapy is due to the induction of a seizure (20). Benzodiazepines are well known antiepileptics (11) and it is suggested that they should be withdrawn during electroconvulsive therapy (3). The purpose of our study was to determine the evidence that exists for this hypothesis in the literature. METHOD: We have reviewed the international literature on electroconvulsive therapy and benzodiazepines through the Medline and Pascal reference database, and performed a manual search of the major international journals. We have retained the references dealing with both benzodiazepine and electroconvulsive therapy. RESULTS: Up to June 1990, 16 references have been found (table I). Of these 16 references, 5 (6, 7, 12, 19, 25) are letters to the editor with no data reported. Six papers (1, 2, 8, 9, 14, 22) examine benzodiazepines as anesthetics and only focus on anesthetic variables, ignoring electroconvulsive therapy-treatment outcome. Only four studies (5, 17, 23, 24) actually examine the incidence of benzodiazepine use on electroconvulsive therapy. Of these, three studies (17, 23, 24) report the incidence of benzodiazepines used as tranquilizers or hypnotics on duration of the electroconvulsive therapy induced seizure. One is a retrospective study (24) showing a decrease in seizure duration due to benzodiazepine use. Of the two prospective studies one shows (23) no difference in seizure duration and the other one (17) a decrease. These studies, however, fail to actually examine how the antidepressive efficacy of electroconvulsive therapy is modified by benzodiazepine use. One recent study (5) measures the antidepressive efficacy of electro-convulsive therapy through the decrease of a validated depression scale. This is a randomized double-blind study examining a short acting benzodiazepine used as an anesthetic in lieu of methohexital. Electroconvulsive therapy efficacy does not appear impaired. CONCLUSION: From our review of the literature it appears that the supposed negative effect of benzodiazepines on the antidepressive action of electroconvulsive therapy has not been demonstrated. Further studies are necessary warranted.

Anti-Anxiety Agents

Non-categorical problem lists in a primary-care information system.

An ambulatory-care patient-tracking system has been implemented that records non-categorical problem descriptions in the outpatient problem list. The system does not restrict physicians to the use of predefined diagnostic categories. Instead, the system stores patient problems in a database as free-text records. Subsequent diagnostic categorization and coding is accomplished through prompted free-text input and appropriate reference databases. This system design allows an outpatient problem-list summary to reflect non-categorical health-status information in addition to coded medical diagnoses.

Ambulatory Care Information Systems

Standardizing clinical laboratory data for the development of transferable computer-based diagnostic programs.

The existence of systematic differences between test results obtained at different laboratories can compromise the development of generally accessible reference databases for interpretive pathology. We review approaches to the elimination of inter-laboratory bias from pathology test results through the use of standard unit transformations. A general transform procedure is described that will permit laboratories serving a common population to make use of reference data, decision rules, and computer-based interpretive programs developed around a larger clinical database than each of these test centers could amass for themselves.

Clinical Laboratory Techniques

Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction.

SUMMARY: Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health. AVAILABILITY: Data is available at https://gdo-meta2db.llnl.gov/ and https://zenodo.org/records/17315984.

Metadata

resLens: genomic language models to enhance antibiotic resistance gene detection.

The rise of antibiotic resistance necessitates advanced tools to detect and analyze antibiotic resistance genes (ARGs). We present resLens, a family of genomic language models that leverage latent genomic representations to enhance ARG detection and analysis. Unlike alignment-based methods constrained by reference databases, resLens fine-tunes a pre-trained DNA language model on curated ARG datasets, achieving competitive or superior performance in classifying resistance genes across multiple evaluation scenarios, including when ARGs exhibit sequences and mechanisms of resistance dissimilar to those in reference datasets.

Journal Article

Automatic analysis of heart rate variation: I. Method and reference values in healthy controls.

Many patients referred to an electrophysiological laboratory may have autonomic dysfunction. Some parasympathetic tests are based on the assessment of heart rate variation induced by breathing, Valsalva maneuver, and standing. We have developed fast and practical computer-based methods to analyze heart rate variation using standard EMG equipment and a personal computer. For quantitative description we have evaluated different algorithms, both earlier described and new ones. Findings in patients with diabetes have been compared with those obtained from healthy subjects in order to determine the diagnostic utility of the various algorithms. The optimal algorithm has been chosen by this and other criteria, and a reference database from healthy subjects has been developed.

Adolescent

nf-core/magmap: Map metatranscriptomes to large collections of genomes.

SUMMARY: The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. AVAILABILITY AND IMPLEMENTATION: The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Software

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial

A further assessment of factors influencing measurements of thioguanine-resistant mutant frequency in circulating T-lymphocytes.

We have used the T-Lymphocyte cloning technique as a method of monitoring the human population for somatic cell mutant frequency. We present a statistical analysis of the experimental factors which may influence the observed mutant frequency. We have obtained consistently high plating efficiencies of T-cells from the mononuclear cell fraction from donor blood samples (mean of 56%, based on 123 observations from 70 individuals). Nevertheless, an inverse correlation of mutant frequency with plating efficiency was observed, and some experimental factors (serum and interleukin-2 batch, and worker) may have a significant effect on the observed mutant frequency. We discuss the difficulties that these possible effects present in establishment of a reference database and design of long-term studies. No significant effect of donor sex on mutant frequency was observed, but age (1.3% increase per year for normal adults) and smoking (56% increase over normal non-smokers) both significantly increased the mutant frequency. We discuss the utility of the assay for the monitoring of populations for heritable DNA damage, and we compare the results to those obtained with lymphocytes using other endpoints, e.g. chromosome aberrations, micronuclei and sister-chromatid exchange.

Age Factors

Metax enables accurate cross-domain taxonomic profiling of metagenomes.

Taxonomic profiling is fundamental to microbiome research, yet achieving high species-level accuracy remains challenging for complex communities that span bacteria, viruses, eukaryotes, and archaea, and these limitations are exacerbated in low-biomass, host-dominated samples. We introduce Metax, a cross-domain taxonomic profiler that integrates coverage-based probabilistic modeling with an expectation-maximization framework to distinguish true microbial signals from artifacts. Across >600 samples from host-associated, environmental, wastewater, and low-biomass clinical settings, including benchmarks with limited reference representation, Metax improved profiling accuracy, achieving on average 55% higher F1 scores and 45% lower Bray-Curtis dissimilarity than other methods. Moreover, this broad evaluation demonstrated that Metax resolved bacterial and viral signatures of peri-implantitis in oral microbiomes and revealed signals suggestive of reagent-borne contaminants and reference misassemblies in plasma-cell-free DNA. By leveraging genome-wide coverage evidence, Metax enables robust cross-domain profiling across diverse sample types and sequencing depths, including settings where reference databases are highly incomplete.

abundance estimation

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Equity in genome sequencing for rare disease diagnosis: a cross-sectional analysis of data from the UK 100,000 Genomes Project.

BACKGROUND: Genome sequencing has improved rare disease diagnosis and is now part of routine clinical care in the National Health Service in England. Automated prioritisation pipelines narrow millions of variants per patient to a small subset for clinical review, a process that relies on allele frequency resources that do not fully represent human genetic diversity. We assessed ancestry-related differences in variant prioritisation and diagnostic outcomes in patients from the UK 100,000 Genomes Project. METHODS: We analysed 29,405 rare disease probands with genome sequencing and linked clinical outcomes data. We used multivariable regression to assess ancestry-related differences in the number of variants prioritised for clinical review, the proportion of prioritised variants that were recorded as diagnostic, and diagnostic yield. We also evaluated the use of ancestry-stratified allele frequency filters derived from an independent, diverse UK cohort (n = 33,724). FINDINGS: Compared with the European ancestry group, the East African group had nearly three times more variants prioritised for clinical review (IRR 2.77, 95% CI 2.33-3.29). Other non-European groups also had significantly higher counts. Diagnostic yield was similar across ancestry groups after adjustment (LRT p = 0.1650). Prioritised variants were less likely to be recorded as diagnostic in East African (OR 0.32, 95% CI 0.22-0.46), West African (0.47, 0.39-0.57), South Asian (0.65, 0.58-0.73), and Middle Eastern (0.68, 0.54-0.86) groups. Applying ancestry-stratified allele-frequency filters removed 3.1% of prioritised variants overall-24.3% in the East African group-without loss of diagnostic sensitivity, including 29.5% of recorded VUS in this group. INTERPRETATION: Differences in the likelihood of prioritised variants being recorded as diagnostic partly reflect limitations of current allele frequency resources, which use broad population groupings that mask within-group diversity. Increased representation of diverse ancestries in reference databases and better estimation of ancestry-appropriate allele frequencies will help reduce inefficiencies and improve equity in variant prioritisation for rare disease diagnosis. FUNDING: The UK Department of Health and Social Care and the EU's Horizon 2020 Research and Innovation Programme.

Humans