Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Potential impact of advanced clinical information technology on cancer care in 2015.

New clinical information technologies now sporadically available will soon be in routine clinical use, bringing many changes to all phases of the cancer care continuum. For example, new technologies such as: (1) The next generation Internet; (2) Real-time clinical decision support systems; (3) Off-line, population-based systems; (4) Large, integrated, individual patient-level phenotypic and genotypic databases with intelligent data mining capabilities; (5) Wireless, invasive and non-invasive physiologic monitoring devices; (6) Natural Language Processing (NLP) systems; and (7) Mathematical models of complex biological systems all have the potential to impact significantly the provision of cancer care throughout its continuum. While new information management and communication techniques and technologies will reduce many of the inefficiencies and inaccuracies of our present systems, there will be an equal, and potentially far more dangerous, set of unintended consequences. Informatics investigators, cancer specialists, and health system administrators must focus on the study of what is working and what is not, as well as, on development and testing of the new clinical information management and communication technologies, if we are to be ready for the future.

Cancer Care Facilities↗

Context-free evolutionary grammars and the structural language of nucleic acids.

This paper introduces and investigates a generative mechanism based on some operations inspired by the large-scale mutations in genomes (deletion, inversion, transposition, duplication). Basic questions regarding these devices and their generated languages are investigated: generative capacity, closure properties, decidability. We also briefly discuss a few problems concerning our model with respect to some structural features of the nucleic acids.

Evolution, Molecular↗

Potential impact of advanced clinical information technology on healthcare in 2015.

Clinical information technologies now sporadically available will soon be in routine clinical use, bringing many changes to healthcare. For example, 1) The next generation Internet; 2) Real-time clinical decision support systems; 3) Off-line, population-based systems; 4) Large, integrated, individual patient-level phenotypic and genotypic databases with intelligent data mining capabilities; 5) Wireless, invasive and non-invasive physiologic monitoring devices; 6) Natural Language Processing (NLP) systems; and 7) Mathematical models of complex biological systems have the potential to impact significantly the future healthcare delivery system. While new information management and communication techniques and technologies will reduce many of the inefficiencies and inaccuracies of our present systems, there will be an equal, and potentially far more dangerous, set of unintended consequences. Informatics investigators and health system administrators must focus on the study of what is working and what is not, as well as, on development and testing of the new clinical information management and communication technologies, if we are to be ready for the future.

Databases as Topic↗

Evidence supporting the role of GIGYF2 in synapse development and autism.

Autism spectrum disorder (ASD) is a heterogeneous condition in which genetically defined subtypes offered insights into underlying biological mechanisms and potential targeted treatments. Here, we investigate the clinical and pathogenic significance of GIGYF2 variants in ASD through an integrated approach combining clinical genetics, conditional knockout (cKO) mouse models, neurobiology, and molecular studies. Through targeted sequencing, large-scale genomic data analysis of neurodevelopmental disorder cohorts, and international collaborations, we identified ten affected individuals from eight families harboring de novo or dominantly inherited likely gene-disruptive (LGD) variants and 13 affected individuals from 13 families with de novo missense variants in GIGYF2. Clinical characterization of 16 probands with GIGYF2 variants revealed common features, including ASD, language problems, intellectual disability, and anxiety. In a Gigyf2 cKO mouse model, we observed pronounced autistic-like behaviors, cognitive deficits, and anxiety-like behaviors, mirroring phenotypes observed in affected individuals. Mechanistically, Gigyf2 deficiency disrupted synaptic homeostasis, as evidenced by altered spine density and miniature excitatory postsynaptic currents, and impaired IGF-1R/mTOR signaling, along with dysregulation of synapse-related genes such as Nrp2. Pharmacological inhibition of mTOR with rapamycin or Torin1, as well as Nrp2 knockdown rescued synaptic defects in Gigyf2 KO neurons. These findings define a novel ASD subtype associated with GIGYF2 variants and establish GIGYF2 as a key regulator of synaptic development and function, implicating GIGYF2 dysfunction in ASD pathogenesis and highlighting the IGF-1R/mTOR pathway as a potential therapeutic target for GIGYF2-related ASD subtype.

Journal Article↗

Structural basis of carbohydrate recognition by lectin II from Ulex europaeus, a protein with a promiscuous carbohydrate-binding site.

Protein-carbohydrate interactions are the language of choice for inter- cellular communication. The legume lectins form a large family of homologous proteins that exhibit a wide variety of carbohydrate specificities. The legume lectin family is therefore highly suitable as a model system to study the structural principles of protein-carbohydrate recognition. Until now, structural data are only available for two specificity families: Man/Glc and Gal/GalNAc. No structural data are available for any of the fucose or chitobiose specific lectins. The crystal structure of Ulex europaeus (UEA-II) is the first of a legume lectin belonging to the chitobiose specificity group. The complexes with N-acetylglucosamine, galactose and fucosylgalactose show a promiscuous primary binding site capable of accommodating both N-acetylglucos amine or galactose in the primary binding site. The hydrogen bonding network in these complexes can be considered suboptimal, in agreement with the low affinities of these sugars. In the complexes with chitobiose, lactose and fucosyllactose this suboptimal hydrogen bonding network is compensated by extensive hydrophobic interactions in a Glc/GlcNAc binding subsite. UEA-II thus forms the first example of a legume lectin with a promiscuous binding site and illustrates the importance of hydrophobic interactions in protein-carbohydrate complexes. Together with other known legume lectin crystal structures, it shows how different specificities can be grafted upon a conserved structural framework.

Amino Acid Sequence↗

Causal thinking and causal language in epidemiology: it's in the details.

Although epidemiology is necessarily involved with elucidating causal processes, we argue that there is little practical need, having described an epidemiological result, to then explicitly label it as causal (or not). Doing so is a convention which obscures the valuable core work of epidemiology as an important constituent of public health practice. We discuss another approach which emphasizes the public health "use value" of research findings in regard to prediction and intervention independent from explicit metaphysical causal claims. Examples are drawn from smoking and lung cancer, with particular focus on the original 1964 Surgeon General's report on smoking and the new version released in 2004. The intent is to help the epidemiologist focus on the pertinent implications of research, which, from a public health point of view, in large part entails the ability to predict and to intervene. Further discussion will center on the importance of differentiating between technical/practical uses of causal language, as might be used in structural equations or marginal structural modeling, and more foundational notions of cause. We show that statistical/epidemiological results, such as "smoking two packs a day increases risk of lung cancer by 10 times" are in themselves a kind of causal argument that are not in need of additional support from relatively ambiguous language such as "smoking causes lung cancer." We will show that the confusion stemming from the use of this latter statement is more than mere semantics. Our goal is to allow researchers to feel more confident in the power of their research to tell a convincing story without resorting to metaphysical/unsupportable notions of cause.

Journal Article↗

Modeling hospital information systems. Part 1: The revised three-layer graph-based meta model 3LGM2.

OBJECTIVES: Not only architects but also information managers need models and modeling tools for their subject of work. Especially for supporting strategic information management in hospitals, the meta model 3LGM2 is presented as an ontological basis for modeling the comprehensive information system of a hospital (HIS). METHODS: In a case study, requirements for modeling HIS have been deduced. Accordingly 3LGM2 has been designed to describe HIS by concepts on three layers. The domain layer consists of enterprise functions and entity types, the logical tool layer focuses on application components and the physical tool layer describes physical data processing components. In contrast to other approaches a lot of inter-layer-relationships exist. 3LGM2 is defined using the Unified Modeling Language (UML). RESULTS: Models of HIS can be created which comprise not only technical and semantic aspects but also computer-based and paper-based information processing. A software tool supporting the creation of 3LGM2 compliant models in a graphical way has been developed. The tool supports in detecting those shortcomings at the logical or the physical tool layers which make it impossible to satisfy the information needs at the domain layer. 3LGM2 can also be used as an ontology for describing HIS in natural language. CONCLUSIONS: Strategic information management even in large hospitals should be and can be supported by dedicated methods and tools. Although there have been good experiences with 3LGM2 concerning digital document archiving at the Leipzig University Hospital, which are presented in part 2, the benefit of the proposed method and tool has to be further evaluated.

Hospital Information Systems↗

Novel approaches and applications in identifying DNA methylation markers of cardio-kidney-metabolic disease.

Cardio-kidney-metabolic (CKM) diseases represent a major public health challenge, accounting for a large proportion of global burden of morbidity and mortality. These conditions share risk factors, including genetic predisposition, environmental exposures, and lifestyle influences, which collectively drive disease development and progression. Epigenetic modifications, particularly DNA methylation (DNAm), serve as key mediators and biomarkers between these risk factors and disease phenotypes by regulating gene expression without altering the DNA sequence. Epigenome-wide association studies have identified DNAm markers associated with CKM diseases and related phenotypes, highlighting both shared pathways and disease-specific epigenetic signatures in inflammation, metabolic dysfunction, and aging-related processes. Longitudinal studies further demonstrate the dynamic nature of DNAm changes over time, offering insights into disease trajectories. Additionally, methylation risk scores integrating multiple epigenetic markers show promise in improving disease prediction and risk stratification beyond traditional clinical factors. To synthesize the current evidence, we conducted a targeted literature search in PubMed for English-language, peer-reviewed articles published between 2014 and the present. Future research leveraging large, well-phenotyped cohorts, advanced statistical methods, and innovative study designs will be critical for uncovering novel biomarkers, refining risk prediction models, and developing targeted epigenetic therapies to mitigate the global burden.

Humans↗

Next generation simulation tools: the Systems Biology Workbench and BioSPICE integration.

Researchers in quantitative systems biology make use of a large number of different software packages for modelling, analysis, visualization, and general data manipulation. In this paper, we describe the Systems Biology Workbench (SBW), a software framework that allows heterogeneous application components--written in diverse programming languages and running on different platforms--to communicate and use each others' capabilities via a fast binary encoded-message system. Our goal was to create a simple, high performance, opensource software infrastructure which is easy to implement and understand. SBW enables applications (potentially running on separate, distributed computers) to communicate via a simple network protocol. The interfaces to the system are encapsulated in client-side libraries that we provide for different programming languages. We describe in this paper the SBW architecture, a selection of current modules, including Jarnac, JDesigner, and SBWMeta-tool, and the close integration of SBW into BioSPICE, which enables both frameworks to share tools and compliment and strengthen each others capabilities.

Biochemical Phenomena↗

MAVL and StickWRLD: visually exploring relationships in nucleic acid sequence alignments.

Many powerful tools have been created to detect and describe the similarities between nucleic acid or protein sequences. Frequently these take the form of a sequence consensus, expressing simple most popular positional identities, positional identities with allowances for varying positions or some type of statistical description of the positional frequency characteristics of the defining sequence family. Despite the fact that some provide intuitively interpretable descriptions of the consensuses themselves, they typically do not give the viewer any information about regions of the sequence that might have inter-positional dependencies, and that therefore do not obey a strict consensus behavior. Herein, we present MAVL (Multiple Alignment Variation Linker) and StickWRLD. MAVL is our web-based application for detecting and displaying both positive and negative inter-positional correlations in nucleic acid sequences. MAVL examines all positional pairs in each of a collection of pre-aligned sequences and determines any pairs that occur with either greater or lesser frequency than a positional frequency matrix would predict. These data are then composited into a StickWRLD representation and supplied back to the user as a VRML (virtual reality modeling language) file. MAVL and StickWRLD can be accessed at http://www.microbial-pathogenesis.org/stickwrld/. A tutorial that explains MAVL features and demonstrates typical user interactions with StickWRLD graphs is available at http://www.microbial-pathogenesis.org/stickwrld/tutorial/sticktut2.html. This tutorial is quite large; please be patient while it loads.

Algorithms↗

The development of language-like communication without a language model.

Deaf children who are unable to acquire oral language naturally and who are not exposed to a standard manual language can spontaneously develop a structured sign system that has many of the properties of natural spoken language. This communication system appears to be largely the invention of the child himself rather than of the caretakers.

Child, Preschool↗

How knowledge drives understanding--matching medical ontologies with the needs of medical language processing.

In this article, we introduce a knowledge-based approach to medical text understanding. From an in-depth consideration of deep sentence and text understanding we distill basic requirements for an adequate knowledge representation framework. These requirements are then matched with currently available medical ontologies (thesauri, terminologies, etc.). A fundamental trade-off is recognized between large-scale conceptual coverage on the one hand, and formal mechanisms for integrity preservation and conceptual expressiveness on the other hand. We discuss various shortcomings of the most wide-spread ontologies to capture medical knowledge in-the-large. As a result, we argue for the need of a formally sound and expressive model along the lines of KL-ONE-style terminological representation systems in the format of description logics. These provide an adequate methodology for designing more sophisticated, flexible medical ontologies serving the needs of 'deep' knowledge applications which are by no means restricted to medical language processing.

Artificial Intelligence↗

A system that facilitates the orientation within procedure nomenclatures through a semantic approach.

The representation of medical concepts should provide the flexibility required to support several purposes. We have implemented a model in which medical terms are represented in a standard format based on a semantic description of the terms. We have focused on the description of procedures. Underlying this project is the assumption that information about medical procedures is crucial in the healthcare system. A prototype has been developed for urology. Because of the large number of terms in the Unified Medical Language System (UMLS) and the abundance of links between them, we have experimented in the use of the UMLS as the foundation for our concept base. We assess the usefulness of this approach and discuss its improvements.

Algorithms↗

The ERATO Systems Biology Workbench: enabling interaction and exchange between software tools for computational biology.

Researchers in computational biology today make use of a large number of different software packages for modeling, analysis, and data manipulation and visualization. In this paper, we describe the ERATO Systems Biology Workbench (SBW), a software framework that allows these heterogeneous application components--written in diverse programming languages and running on different platforms--to communicate and use each others' data and algorithmic capabilities. Our goal is to create a simple, open-source software infrastructure which is effective, easy to implement and easy to understand. SBW uses a broker-based architecture and enables applications (potentially running on separate, distributed computers) to communicate via a simple network protocol. The interfaces to the system are encapsulated in client-side libraries that we provide for different programming languages. We describe the SBW architecture and the current set of modules, as well as alternative implementation technologies.

Computational Biology↗

[Use of the ACSL simulation language for physiologic toxicokinetic models].

For the description of the processes of absorption, excretion or elimination of chemicals, the open one- or two-compartment models have been used thus far. The latter consist mainly of the fast (central) and slow (peripheral) compartments. The toxicological studies were based on an assumption that the organic processes develop according to is the first order kinetic reaction. However, the absorption, elimination or excretion of toxic chemicals are in fact much more complicated processes that should be explained using, e.g. the physiologically-based toxicokinetic (PBTK) models, covering physiological, biochemical and metabolic parameters, as well as the allometric calibration of selected parameters for interspecies extrapolations, and in vitro/in vivo extrapolations of metabolic parameters. Simulation languages, e.g. ACSL (Advanced Continuous Simulation Language) are indispensable application tools to be operated with PBTK models. They have been developed for modelling systems described by time-dependent non-linear differential equations and/or transfer functions. ACSL with its interfaces (ACSL Builder, ACSL Graphic Modeller, ACSL Math) ensures data input and communication inside the model by the control, transfer and computed parameters. The physiologically-based toxicokinetic models employ a large number of different parameters, which enables, e.g. forecasting the dose/effect or dose/response relationship absorption rate, metabolic pathways, excretion or elimination according to the absorbed dose of xenobiotic; evaluation of risk assessment; extrapolation from high to low doses characteristic of environmental exposure or setting biological exposure limits.

Body Fluid Compartments↗

Knowledge representation of signal transduction pathways.

MOTIVATIONS: Signal transduction is the common term used to define a diverse topic that encompasses a large body of knowledge about the biochemical mechanisms. Since most of the knowledge of signal transduction resides in scientific articles and is represented by texts in natural language or by diagrams, there is the need of a knowledge representation model for signal transduction pathways that can be as readily processed by a computer as it is easily understood by humans. RESULTS: A signal transduction pathway representation model is presented. It is based on a compound graph structure and is designed to handle the diversity and hierarchical structure of pathways. A prototype knowledge base was implemented on a deductive database and a number of biological queries are demonstrated on it.

Amino Acid Motifs↗

Simulators in clinical surgery.

Simulators are no replacement for patients in surgical learning. Live patients are required for teaching clinical signs and skills. Large numbers of students, a relative lack of motivation, a decreasing number of common cases, unwilling patients, differences in language, etc., make clinical teaching in India a bitter problem. Because patient-related problems are important, surgical training using models can help students to gain effective control over surgical signs and skills.

Education, Medical, Undergraduate↗

Why is that? Structural prediction and ambiguity resolution in a very large corpus of English sentences.

Previous psycholinguistic research has shown that a variety of contextual factors can influence the interpretation of syntactically ambiguous structures, but psycholinguistic experimentation inherently does not allow for the investigation of the role that these factors play in natural (uncontrolled) language use. We use regression modeling in conjunction with data from the British National Corpus to measure the amount and specificity of the information available for disambiguation in natural language use. We examine the Direct Object/Sentential Complement ambiguity and the closely related issue of complementizer use in sentential complements, and find that both ambiguity resolution and complementizer use can be predicted from contextual information.

Cognition↗