Search PubMedSearch

SEARCH · Search PubMed

Results for “Textual analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

BioMedGraphica: An All-in-One Platform for Joint Textual Biomedical Prior Knowledge and Numeric Graph Generation.

Multi-omic data analysis is essential for scientific discovery in precision medicine. However, translating statistical results of omic data analysis into novel scientific hypothesis remains a significant challenge. Human experts must manually review analysis results and generate new hypothesis based on extensive and inter-connected biomedical prior knowledge, which is subjective and not scalable. While large language models (LLMs) can accelerate the discovery, their reasoning improves when grounded in structured, auditable and comprehensive biomedical prior knowledge. Biomedical knowledge, however, is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of AI systems to fully leverage biomedical data for scientific discovery. To address these challenges, we developed BioMedGraphica , an all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified knowledge graph containing 2,306,921 entities and 27,232,091 relations. In addition, to the best of our knowledge, this is the first work to propose a novel Textual-Numeric Graph (TNG) data-structure for multi-omics data analysis. In TNG, textual information captures prior biological knowledge (e.g., transcription start sites, functions, mechanisms), while numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data-structure for the development of graph foundation models, with the potential to improve prediction performance and interpretability, while also augmenting LLMs by supplying graph-structured mechanistic context to strengthen reasoning. The details for BioMedGraphica code can be accessed by github link: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph data can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

biomedical knowledge graph

Knowledge representation for platform-independent structured reporting.

Structured reporting systems allow health care providers to record observations using predetermined data elements and formats. We present a generalized language, based on the Standard Generalized Markup Language (SGML), for platform-independent structured reporting. DRML (Data-entry and Report Markup Language) specifies hierarchically organized concepts to be included in data-entry forms and reports. DRML documents serve as the knowledge base for SPIDER, a reporting system that uses the World Wide Web as its data-entry medium. SPIDER generates platform-independent documents that incorporate familiar data-entry objects such as text windows, checkboxes, and radio buttons. From the data entered on these forms, SPIDER uses its knowledge base to generate outline-format textual reports, and creates datasets for analysis of aggregate results. DRML allows knowledge engineers to design a wide variety of clinical reports and survey instruments.

Computer Communication Networks

BioMedGraphica: an all-in-one platform for joint textual biomedical prior knowledge and numeric graph generation.

MOTIVATION: Multiomics data analysis is essential for scientific discovery in precision medicine. However, translating analysis results of omics data analysis into novel scientific hypotheses remains a significant challenge. Human experts must manually review analysis results and generate new hypotheses based on extensive and interconnected biomedical prior knowledge, which is subjective and not scalable. While large language models can accelerate the discovery, their reasoning improves when grounded in structured, auditable, and comprehensive biomedical prior knowledge. However, biomedical knowledge is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of artificial intelligence systems to fully leverage biomedical data for scientific discovery. RESULTS: We developed BioMedGraphica, a novel all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified textual prior knowledge graph containing 2 306 921 entities and 27 232 091 relations. In addition, we present a novel textual-numeric graph (TNG) data structure concept, where textual information captures prior biological knowledge (e.g. transcription start sites, functions, mechanisms), numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data structure for developing novel graph analysis models. AVAILABILITY AND IMPLEMENTATION: The code is available at: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph database can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

Humans

The use of organizational strategies to improve memory for prose passages.

Previous studies have shown that incidental memory for the material increases when older adults are forced to analyze material to the extent necessary to impose an organizational structure. The present experiment sought to extend this finding by examining the effects of enforced organizational strategies on the memory of older adults for textual material. Young and old adults were required to sort the scrambled sentences of a prose passage into the correct order. A subsequent incidental memory test showed that, when older adults were required to make an in-depth analysis to sort the material, their incidental memory for the textual information was approximately equal to that of their younger counterparts. Additional analysis revealed that, although older adults spent more time sorting the material than did younger adults, it was only when required to analyze the material to a sufficient degree that the older adults showed any improvement in memory.

Adolescent

Models and practice in medicine: menopause as syndrome or life transition?

Biomedical knowledge, like scientific knowledge in general, is a product of a social and culture milieu. Moreover, in the analysis of biomedicine a distinction must be maintained between textual and clinical knowledge. Through examination of medical texts and of social science literature, the susceptibility of menopause to numerous interpretations is demonstrated. The generation of clinical models from current information available to physicians is examined and it is suggested that these can be thought of as folk models. Several current clinical models of menopause are presented and the implications for sociological analysis of the subject discussed.

Adult

Exploratory analysis of the medical record.

Current patient information systems such as SCAMP have the capacity to store not only highly-structured information such as problem codes, drug lists and laboratory values, but also richer, clinically-descriptive information such as comprehensive natural language problem summaries and other textual data that retain the full clinical information used by physicians in managing their patients. The clinical importance, richness, and extensiveness of this information suggest that techniques which allow computers to process textual data may play a helpful role in clinical research. The analysis programs of the SCAMP system have been developed to explore the potential of this approach.

Ambulatory Care Information Systems

Discourse analysis: a new methodology for understanding the ideologies of health and illness.

Discourse analysis is an interdisciplinary field of inquiry which has been little employed by public health practitioners. The methodology involves a focus upon the sociocultural and political context in which text and talk occur. Discourse analysis is, above all, concerned with a critical analysis of the use of language and the reproduction of dominant ideologies (belief systems) in discourse (defined here as a group of ideas or patterned way of thinking which can both be identified in textual and verbal communications and located in wider social structures). Discourse analysis adds a linguistic approach to an understanding of the relationship between language and ideology, exploring the way in which theories of reality and relations of power are encoded in such aspects as the syntax, style and rhetorical devices used in texts. This paper argues that discourse analysis is pertinent to the concerns of public health, for it has the potential to lay bare the ideological dimension of such phenomena as lay health beliefs, the doctor-patient relationship, and the dissemination of health information in the entertainment mass media. This dimension is often neglected by public health research. The method of discourse analysis is explained, and examples of its use in the area of public health given.

Communication

The CODATA/IUIS Hybridoma Data Bank: development of a hybrid system to handle complex data relationships.

System design for the Hybridoma Data Bank, a database of comprehensive information on immunoreagents for use by scientists in diverse disciplines, is described. Unique problems include: use of nomenclature from diverse fields that is neither static nor standard; the need for two representations of the database--textual for readability and numeric for complex search capabilities, analysis and data compression; and a method of translating between the two representations of the database.

Animals

On studying the discourse of medical encounters. A critique of quantitative and qualitative methods and a proposal for reasonable compromise.

Studies of doctor-patient communication, although leading to diverse findings, have not lent themselves to replication and also have not captured important features of medical discourse. Quantitative methods alone do not deal with the complexities of medical encounters, usually are not helpful in analyzing the social context of discourse, do not clarify underlying themes and structures, and are costly and tedious to use. With qualitative methods, the selection of discourse for analysis is not straightforward, quality of interpretation is difficult to evaluate, and textual presentation is not clear-cut. Several criteria of an appropriate method offer reasonable compromises in dealing with medical discourse: 1) discourse should be selected through a sampling procedure, preferably a randomized technique; 2) recordings of sampled discourse should be available for review by other observers; 3) standardized rules of transcription should be used; 4) the reliability of transcription should be assessed by multiple observers; 5) procedures of interpretation should be decided in advance, should be validated in relation to theory, and should address both content and structure of texts; 6) the reliability of applying interpretive procedures should be assessed by multiple observers; 7) a summary and excerpts from transcripts should accompany the interpretation, but full transcripts should also be available for review; and 8) texts and interpretations should convey the variability of content and structure across sampled texts. An ongoing study applies these criteria to research on ideology and social control in medical encounters.

Communication

The Irma dream, self-analysis, and self-supervision.

The Irma dream has special historical significance. Erikson and others have placed it in historical, social, and cultural context. The manifest dream was elaborated in terms of analytic surface with analysis of form and content, patterns and movement in time and space, etc. There are, however, limits to textual reinterpretations. Further psychobiographic consideration of the Irma dream highlights issues of transference, countertransference and their sources in unconscious conflict and trauma. The Irma dream was initially a secret dream which represented the initiation of a self-analytic and supervisory process. Freud's revealing the dream and imagining the commemoration of the discovery of "the secret of the dream" marked the termination of formal self-analysis within analysis interminable.

Countertransference

Free text analysis.

In the context of hospital information systems (HIS) medical free text analysis is reviewed with respect to current automated approaches to literature retrieval, case retrieval and fact retrieval from textual data in the patient record. The Unified Medical Language System (UMLS) project has enormously stimulated current research. It is expected that UMLS knowledge sources and SNOMED III (which need a translation into other languages as soon as possible) as well as the conceptual graphs formalism, could become standards to utilize free text information contained in HIS databases.

Abstracting and Indexing

The telephone interview as a data collection method.

The interview is one of the most frequently used methods of collecting qualitative data. This paper offers a description of how the telephone interview may be used to collect data in both qualitative and quantitative studies. It offers a description of how to plan and use the telephone as a means of collecting data and justifies the use of this method. The paper also describes and illustrates how textual data that arises from the use of telephone interviews may be analysed by computer. Two approaches to analysis are outlined: (1) simple drawing together of responses to questions and (2) the searching for categories within the data and the organisation of text within those categories. The paper identifies limitations of the approach to data collection and points to further reading on the topic.

Humans

FOLD: integrated analysis and display of protein secondary structure.

FOLD, a computer program for the definition and analysis of protein secondary structure, is described. Algorithms implemented in the software are reviewed. These include methods for the identification of simple features such as hydrogen bonds, alpha helices, beta strands, beta bulges, and beta and psi turns. Techniques are also described for the definition and analysis of higher-order structures, such as beta hairpins, beta sheets and their topology, and beta barrels. In addition to considerable textual output the program supports visualization of protein secondary structure in either an atom-based display style or one reproducing the characteristics of a so-called ribbon drawing.

Computer Graphics

Textual prompts as an antecedent cue self-management strategy for persons with mild disabilities.

Providing learners written task analyses to be used as textual prompts was examined as a self-management strategy for persons with mild disabilities. Initially, modeling, corrective verbal feedback, and contingent descriptive praise were employed to train participants to use the written task analysis to perform one home maintenance task. Subsequently, participants were tested on their use of different task analyses combined with general feedback to perform two novel home maintenance tasks. No training was provided on how to use these new task analyses. Either a multiple baseline or a multiple probe across settings experimental design was used to control extraneous variables. Results indicated that the written task analyses served as self-administered textual prompts and, along with general feedback, provided stimulus control for the second and third tasks. When the self-management task analyses and general feedback were withdrawn, transfer of stimulus control occurred to the natural discriminative stimuli for the majority of tasks. The research suggests that written task analyses, as presented in the present study, may have utility for the self-management of instruction by persons with mild disabilities.

Adult

On-line information sources on chemical substances.

Information technology has brought about changes in the work patterns of researchers and scientists. After some hints on the on-line facilities needed to be connected to the international host computers, an analysis is made of some of the main automated sources available to retrieve information on chemical substances. Special emphasis is given to textual-numeric data banks, first reviewing the main chemical dictionaries, like Registry and Chemline, and then focusing on those sources that offer immediate information in case of emergency. Among the Toxnet files, produced and managed within the US National Library of Medicine Toxicology Information Program, play a very important role in offering publicly available data on toxicology and on hazardous chemicals. Therefore, the Hazardous Substances Data Bank (HSDB) and the Registry of Toxic Effects of Chemical Substances (RTECS) are described for their relevance thereon. Other data banks produced in Europe, like the Environmental Chemicals Data Information Network (ECDIN) and the very specialized Major Hazard Incident Data Service (MHIDAS) are also briefly outlined. To integrate this overview on online information, the attention is then shifted on sources having the characteristic of reference databases: prestigious files covering the international scientific literature, as CA/Chemabs, Toxline/Toxlit, Embase, Medline are introduced. Implications of on-line technology in enhancing information access in the next future are discussed, pointing out the new tools created to meet the information needs of end-users.

Databases, Bibliographic

Proposed standard for image cytometry data files.

A number of different types of computers running a variety of operating systems are presently used for the collection and analysis of image cytometry data. In order to facilitate the development of sharable data analysis programs, to allow for the transport of image cytometry data from one installation to another, and to provide a uniform and controlled means for including textual information in data files, this document describes a data storage format that is proposed as a standard for use in image cytometry. In this standard, data from an image measurement are stored in a minimum of two files. One file is written in ASCII to include information about the way the image data are written and optionally, information about the sample, experiment, equipment, etc. The image data are written separately into a binary file. This standard is proposed with the intention that it will be used internationally for the storage and handling of biomedical image cytometry data. The method of data storage described in this paper is similar to those methods published in American Association of Physicists in Medicine (AAPM) Report Number 10 and in ACR-NEMA Standards Publication Number 300-1985.

Electronic Data Processing

TongueTwister: an integrated program for analyzing lickometer data.

The analysis of lickometer data is often rendered prohibitively tedious by the large volume of data generated by the typical experiment. TongueTwister is an integrated program for the rapid and automatic analysis, presentation, and summary of long- and medium-access data collected by lickometers or of brief-access data collected by multi-bottle lickometers such as the DiLog Instruments MS80. The program was written in C+2 for Macintosh computers, and analyzes data collected by MS-DOS PCs. It takes advantage of the Macintosh user interface to provide quick and convenient output from all the files of a single experimental session, and to export the data to third-party statistical software or other documents. It can batch-process data files by automatically opening and analyzing all the files in a directory; thus, the user can employ directories as a simple database for organizing experimental groups. When a lickometer data file is opened, a textual summary, a raster plot of the lick pattern, the cumulative licks, the lick rate, a histogram of inter-lick intervals, and a breakdown of the session by fractions are automatically calculated and displayed. When an MS80 brief-access file is opened, the lick pattern for each tube presentation and a textual summary of the mean values derived for each tube are automatically displayed. If a directory of files is opened, the mean values derived across all the individual files are calculated and graphed. Analysis parameters can be tailored to the investigator's liking. Tables or graphs can be saved to disk, or copied and pasted into other Macintosh programs for additional analysis. The program may also be used for general-purpose analysis of periodic event records.

Animals