Correction of misspellings and typographical errors in a free-text medical English information storage and retrieval system.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
An interactive program has been developed for entering, verifying, and storing chemical structures encoded in the Wiswesser Line Notation. The program calculates a molecular formula from the WLN for checking and then generates a bit string fragment code and connection table of the atoms in the structure. The encoding and entry of the WLN using a CRT has significantly improved the speed of the total compound registration process.
An improved interactive system for searching substructure and biological activity data has been developed. Features of the system include a two-level substructure search (fragment screen and atom by atom) and an expanded biological activity data base. The system operates on a file of about 150 000 compounds.
Morbidity and mortality data are necessary bases for the decision-making processes relevant to allocation of public funds for animal disease diagnoses and research. A system for information storage and retrieval capable of handling diagnostic data such as results of microbiology, parasitology, necropsy, and histopathology as well as demographic data such as owner, species, sex, breed, or geographic origin of the animal is described. This information is available to veterinarians, epidemiologists, herdsmen, and others involved in disease prevention or control efforts. The system described utilizes natural language, thus overcoming difficulties encountered in systems with numerical intermediates. Used and revised for the last 10 years, the system described has proved useful for annual administrative quantitation of services performed. In fact, the Concordance Index serves as the annual report of the University of Missouri Veterinary Medical Diagnostic Laboratory. Having accurate detailed information on individual cases, as well as a variety of composite data, has been extremely helpful in the documentation necessary for attracting funding for study of specific disease states.
The capabilities of a mini-computer-oriented permanent pacemaker information system are described. Extensive patient and pacer functional data are maintained in readily accessible files which may be displayed on a CRT terminal or printed out. Selective sorting of stored information may be accomplished according to any desired set of inclusive or exclusive criteria. Intelligent, comprehensive patient follow-up has been greatly facilitated through application of the system. In view of the rapidly expanding pacemaker population, it is suggested that cooperative regional networks operating with comparable information storage and retrieval structures will provide the only means for adequate patient surveillance and compilation of necessary pacemaker data.
Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.
BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.
MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.
MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.
The potential of the diverse chemistries present in natural products (NP) for biotechnology and medicine remains untapped because NP databases are not searchable with raw data and the NP community has no way to share data other than in published papers. Although mass spectrometry (MS) techniques are well-suited to high-throughput characterization of NP, there is a pressing need for an infrastructure to enable sharing and curation of data. We present Global Natural Products Social Molecular Networking (GNPS; http://gnps.ucsd.edu), an open-access knowledge base for community-wide organization and sharing of raw, processed or identified tandem mass (MS/MS) spectrometry data. In GNPS, crowdsourced curation of freely available community-wide reference MS libraries will underpin improved annotations. Data-driven social-networking should facilitate identification of spectra and foster collaborations. We also introduce the concept of 'living data' through continuous reanalysis of deposited data.
The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.
PURPOSE: Typically stored as unstructured notes, surgical pathology reports contain data elements valuable to cancer research that require labor-intensive manual extraction. Although studies have described natural language processing (NLP) of surgical pathology reports to automate information extraction, efforts have focused on specific cancer subtypes rather than across multiple oncologic domains. To address this gap, we developed and evaluated an NLP method to extract tumor staging and diagnosis information across multiple cancer subtypes. METHODS: The NLP pipeline was implemented on an open-source framework called Leo. We used a total of 555,681 surgical pathology reports of 329,076 patients to develop the pipeline and evaluated our approach on subsets of reports from patients with breast, prostate, colorectal, and randomly selected cancer subtypes. RESULTS: Averaged across all four cancer subtypes, the NLP pipeline achieved an accuracy of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.89 for T staging, 0.90 for N staging, and 0.97 for M staging. It achieved an F1 score of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.88 for T staging, 0.90 for N staging, and 0.24 for M staging. CONCLUSION: The NLP pipeline was developed to extract tumor staging and diagnosis information across multiple cancer subtypes to support the research enterprise in our institution. Although it was not possible to demonstrate generalizability of our NLP pipeline to other institutions, other institutions may find value in adopting a similar NLP approach-and reusing code available at GitHub-to support the oncology research enterprise with elements extracted from surgical pathology reports.
Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.
Methods have been developed for assessing the cognitive parameters contributing to a memory disorder. Our findings suggest that individuals with Huntington disease have impairments in the encoding of new information and the consistent retrieval from storage of learned material. Their difficulties lie particularly in the realm of episodic memory.
Nineteen normal male subjects received 1.0 milligram of physostigmine or 1.0 milligram of saline by a slow intravenous infusion on two nonconsecutive days. Physostigmine significantly enhanced storage of information into long-term memory. Retrieval of information from long-term memory was also improved. Short-term memory processes were not significantly altered by physostigmine.
Explore the source record for details and available documents.
A word recognition task was designed to determine the stage in memory affected by a single 10-mg intravenous injection of diazepam and the duration of the effect. Injection in three experimental subjects produced an anterograde amnesia for the 14 to 24-minute period immediately after injection. Memory loss resulted from impaired storage, the stage during which information is entered into memory. Retention and retrieval stages of memory were unaffected. This temporary amnesia may result from increased inhibition in the hippocampal system produced by diazepam, which shares many properties with the inhibitory neurotransmitter gamma-aminobutyric acid.
A graphic format is presented for the display and storage of data relating to gestational age. The graph permits rapid retrieval and synthesis of often confusion information and is thereby useful in the management of complicated pregnancies.