Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Feature expressions: creating and manipulating sequence datasets.

Annotation of features, such as introns, exons and protein coding regions in GenBank/EMBL/DDBJ entries is now standardized through use of the Features Table (FT) language. The essence of the FT language is described by the relation 'expression-->sequence', meaning that each FT expression evaluates to a sequence. For example, the expression M74750:1..50 evaluates to the first 50 bases of the sequence with accession number M74750. Because FT is intrinsic to the database definition, it can serve as a software- and platform-independent lingua franca for sequence manipulation. The XYLEM package makes it possible to create and manipulate sequence datasets using FT expressions. FEATURES is a program that resolves FT expressions into their corresponding sequences. Annotated features can be retrieved either by feature key or by expression. Even unannotated portions of a sequence can be retrieved by user-generated FT expressions. Applications of the FT language include retrieval of subsequences from large sequence entries, generation of chromosome models or artificial DNA constructs, and representation of restriction maps or mutants.

Base Sequence↗

Constructing influence views from data to support dynamic decision making in medicine.

A dynamic decision model can facilitate the complicated decision-making process in medicine, in which both time and uncertainty are explicitly considered. In this paper, we address the problem of automatic construction of a dynamic decision model from a large medical database. Within the DynaMoL (a dynamic decision modeling language) framework, a model can be represented in influence view. Thus, our proposed approach first learns the structures of the influence view based on the minimal description length (MDL) principle, and then obtains the conditional probabilities of the model by Bayesian method. The experiment results demonstrate that our system can efficiently construct the influence views from data with high fidelity.

Algorithms↗

Pitch characteristics of infant-directed speech affect infants' ability to discriminate vowels.

"Baby talk" or speech directed to prelinguistic infants is high in pitch and has exaggerated pitch contours (up/down patterns of pitch change) across languages and cultures. Using an acoustic model, we predicted that the large pitch contours of infant-directed speech should improve infants' ability to discriminate vowels. On the other hand, the same model predicted that high pitch would not benefit, and might actually impair, infants' ability to discriminate vowels. We then confirmed these predictions experimentally. We conclude that the exaggerated pitch contours of infant-directed speech aid infants' acquisition of vowel categories but that the high pitch of infant-directed speech must serve another function, such as attracting infants' attention or aiding emotional communication.

Attention↗

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article↗

The performance of typically developing 2 1/2-year-olds on dynamic display AAC technologies with different system layouts and language organizations.

The current generation of augmentative and alternative communication (AAC) technologies is largely based on conceptual models of adults who are not disabled (J. Light & P. Lindsay, 1991). As a result, there is a large "cost of learning" placed on young children. This paper presents the results of a study designed to investigate the learning demands of dynamic display systems that differed in system layout and language organization for children approximately 2 1/2 years old (2 years 5 months to 2 years 11 months). Thirty typically developing children were asked to locate 12 vocabulary items within a play context of a birthday party. Ten children were randomly assigned to each of 3 system approaches: vocabulary in a grid format organized taxonomically, vocabulary in a grid format organized schematically, and vocabulary in an integrated scene organized schematically. The children participated in 4 learning and testing sessions and 1 generalization session. Results indicated that the children performed poorly in all conditions but were able to locate more vocabulary items in the schematic scene condition than the taxonomic grid or schematic grid conditions. There was evidence that the children failed to generalize their knowledge of the vocabulary to facilitate learning of novel vocabulary items. The current design of AAC dynamic display systems appears to be inappropriate for very young children. Rather than relying solely on technology for these young children, early intervention should target multiple modes of communication. AAC technologies should be redesigned to reduce learning demands. Results are discussed with implications for practice and suggestions for future research.

Analysis of Variance↗

Dynamic decision analysis in medicine: a data-driven approach.

Dynamic decision analysis concerns decision problems in which both time and uncertainty are explicitly considered. Two major challenges in dynamic decision analysis are on proper formulation of a model for the problem and effective elicitation of the numerous time-dependent conditional probabilities for the model. Based on a new, general dynamic decision modeling framework called DynaMoL (Dynamic decision Modeling Language), we propose a data-driven approach to addressing these issues. Our approach uses available problem data from large medical databases, guides the decision modeling at a proper level of abstraction and establishes a Bayesian learning method for automatic extraction of the probabilistic parameters. We demonstrate the theoretical implications and practical promises of this new approach to dynamic decision analysis in medicine through a comprehensive case study in the optimal follow-up of patients after curative colorectal cancer surgery.

Bayes Theorem↗

Postmodernism and immune selfhood.

Two research traditions in immunology, supposedly centered on the same issue of immune identification, have followed different theoretical goals and originated from competing philosophical foundations. These may be labelled modernist and postmodernist, respectively, thereby applying cultural and philosophical categories to immunology in order to articulate potential scientific resonances with the broader culture. To accept that exercise an important caveat is imposed, namely, this translation is most appropriately discussed at the level of metaphor. In other words, I will structure my treatment of these issues as expressed in the metaphorical language of the discipline, and thus the bulk of this discussion will focus on how the language and modeling of the science draws from the culture-at-large. Scientists seek images from their everyday lives to describe phenomena that may be poorly articulated in their technical discourse; such is the utility and importance of metaphors generally, and thus it is not surprising that we might discern echoes of a postmodernist sentiment in the metaphors borrowed from post-World War II culture. I will also discuss, to a more limited extent, how postmodernists have sought support for their own ideological arguments in immunology. This last topic serves only to illustrate the bidirectionality of scientific discourse with the society in which it is embedded.

Allergy and Immunology↗

Distributional typicality: a new approach to estimating noun and verb usage from large scale text corpora.

This paper reports a new approach to estimating the extent to which words have predominant noun and verb usages which do not require human judgments about parts of speech. The Hyperspace Analog to Language model (HAL, Lund & Burgess, 1996) was used to computationally estimate noun vs verb usage based on the statistical regularities present in a large-scale electronic text corpus. This measure can be used to estimate the extent to which a given word occurs in typical noun or verb sentence contexts (i.e., its distributional typicality) in informal contemporary discourse.

Cognition↗

OQAFMA Querying agent for the Foundational Model of Anatomy: a prototype for providing flexible and efficient access to large semantic networks.

The development of large semantic networks, such as the UMLS, which are intended to support a variety of applications, requires a flexible and efficient query interface for the extraction of information. Using one of the source vocabularies of UMLS as a test bed, we have developed such a prototype query interface. We first identify common classes of queries needed by applications that access these semantic networks. Next, we survey StruQL, an existing query language that we adopted, which supports all of these classes of queries. We then describe the OQAFMA Querying Agent for the Foundational Model of Anatomy (OQAFMA), which provides an efficient implementation of a subset of StruQL by pre-computing a variety of indices. We describe how OQAFMA leverages database optimization by converting StruQL queries to SQL. We evaluate the flexibility and efficiency of our implementation using English queries written by anatomists. This evaluation verifies that OQAFMA provides flexible, efficient access to one such large semantic network, the Foundational Model of Anatomy, and suggests that OQAFMA could be an efficient query interface to other large biomedical knowledge bases, such as the Unified Medical Language System.

Abstracting and Indexing↗

Models of natural language understanding.

This paper surveys some of the fundamental problems in natural language (NL) understanding (syntax, semantics, pragmatics, and discourse) and the current approaches to solving them. Some recent developments in NL processing include increased emphasis on corpus-based rather than example- or intuition-based work, attempts to measure the coverage and effectiveness of NL systems, dealing with discourse and dialogue phenomena, and attempts to use both analytic and stochastic knowledge. Critical areas for the future include grammars that are appropriate to processing large amounts of real language; automatic (or at least semi-automatic) methods for deriving models of syntax, semantics, and pragmatics; self-adapting systems; and integration with speech processing. Of particular importance are techniques that can be tuned to such requirements as full versus partial understanding and spoken language versus text. Portability (the ease with which one can configure an NL system for a particular application) is one of the largest barriers to application of this technology.

Cognition↗

Word frequency effects in high-dimensional co-occurrence models: A new approach.

The HAL (hyperspace analog to language) model of lexical semantics uses'global word co-occurrence from a large corpus of text to calculate the distance between words in co-occurrence space. We have implemented a system called HiDEx (High Dimensional Explorer) that extends HAL in two ways: It removes unwanted influence of orthographic frequency from the measures of distance, and it finds the number of words within a certain distance of the word of interest (NCount, the number of neighbors). These two changes to the HAL model produce measures of word neighborhood density that are reliably predictive of human lexical decision reaction times.

Humans↗

Literacy for health: an interdisciplinary model.

Traditionally, many literacy and health education programs have had difficulty in significantly affecting vulnerable priority populations. The materials used were largely generalized for one language, one level of literacy, and one culture. A multidiscipline review of literature discusses the relationship between literacy, health, and culture and provides rationale for the interdisciplinary literacy for health model. The model's synthesis of anthropology, linguistics, literacy, nursing, and community partnership guides development of culturally and linguistically appropriate materials for successful adoption and diffusion within a priority population. In Nepal, the model is being used in the Mugom first-language literacy project among a group of remote Tibetan Buddhist peoples.

Culture↗

Language models based on Hebbian cell assemblies.

This paper demonstrates how associative neural networks as standard models for Hebbian cell assemblies can be extended to implement language processes in large-scale brain simulations. To this end the classical auto- and hetero-associative paradigms of attractor nets and synfire chains (SFCs) are combined and complemented by conditioned associations as a third principle which allows for the implementation of complex graph-like transition structures between assemblies. We show example simulations of a multiple area network for object-naming, which categorises objects in a visual hierarchy and generates different specific syntactic motor sequences ("words") in response. The formation of cell assemblies due to ongoing plasticity in a multiple area network for word learning is studied afterwards. Simulations show how assemblies can form by means of percolating activity across auditory and motor-related language areas, a process supported by rhythmic, synchronized propagating waves through the network. Simulations further reproduce differences in own EEG&MEG experiments between responses to word- versus non-word stimuli in human subjects.

Animals↗

Answering the connectionist challenge: a symbolic model of learning the past tenses of English verbs.

Supporters of eliminative connectionism have argued for a pattern association-based explanation of language learning and language processing. They deny that explicit rules and symbolic representations play any role in language processing and cognition in general. Their argument is based to a large extent on two artificial neural network (ANN) models that are claimed to be able to learn the past tenses of English verbs (Rumelhart & McClelland, 1986, Parallel distributed processing, Vol. 2, Cambridge, MA: MIT Press; MacWhinney & Leinbach, 1991, Cognition, 40, 121-157). In this article we critically review Rumelhart and McClelland's as well as MacWhinney and Leinbach's ANN models and conclude that they do not succeed in the assigned task of learning the past tenses of English verbs. In order to answer their challenge to the symbolic processing approach, we present our symbolic pattern associator (SPA)-a general-purpose pattern associator that can learn to associate arbitrary discrete patterns. We carried out several experiments with the SPA using the same set of verbs that was used in MacWhinney and Leinbach's simulation with more realistic training and testing procedures. The SPA outperformed the connectionist models by a wide margin in the accuracy of learning, and successful inductive generalizations to unseen verbs. Our SPA has very natural and psychologically realistic explanations to many psychological effects such as U-shaped learning curve, and is much closer to human subjects in predicting past tense of the pseudo-verbs. In contrast to ANNs, whose internal representations are entirely opaque, the SPA can represent the acquired knowledge in the form of production rules that allow for further higher-level processing and integration, resulting in linguistically realistic associative templates for irregular verbs and production rules for regular verbs. In the light of these findings, we conclude that eliminative connectionists' vision of cognition as simple pattern association and pattern recognition without symbolic representation is inadequate. Pattern association as such does not imply rule-less or cue-based models of language acquisition or of human learning in general.

Cognition↗

A five-phase model for clinical-outcome research.

UNLABELLED: Through a variety of approaches, speech-language pathologists and audiologists have produced strong evidence that treatments are generally potent. However, we have largely ignored the accepted standards for clinical-outcome testing used throughout the broader research community (e.g., by other clinical disciplines, federal regulators, and third-party payers). Several clinical professions recognize a comprehensive model for organizing and scaffolding the many forms of clinical-outcome research. An adaptation of this five-phase model of clinical-outcome research is examined as a means for structuring forms of clinical research throughout audiology and speech-language pathology. Within the organizing structure, relationships become apparent between types and grades of scientific evidence and the processes underpinning evidence-based practice which ultimately lead to decisions on the status of intervention protocols. LEARNING OUTCOMES: Readers will be able to distinguish the phases of clinical-outcome research in a comprehensive model. Readers will be able to identify relationships between the structure of the model and broadly recognized concepts associated with the terms 'efficacy' and 'effectiveness.' Readers will be able to identify indicators of quality for controlled clinical trials.

Biomedical Research↗

Image library of biological macromolecules.

An Image Library of Biological Macromolecules is described, which contains image and text files related to structures of biological macromolecules. Currently, the Library has approximately 3000 image files of approximately 300 structures of biological macromolecules whose coordinates are available in the Protein Data Bank and in the Nucleic Acid Database. The entries include all RNA structures, approximately 70 DNA structures, 150 proteins and a few carbohydrates. The Library contains further images of amino acids, of standard and modified nucleotides and of nucleic acid model structures. Each entry consists of an annotation file with bibliographic and sequence information and possibly comments, of a color-coded distance plot and of structure images. Almost all of the images are available both in a mono and in a stereo representation. Standard procedures for generating these images were strictly avoided. Therefore, mixed rendering, coloring and labeling techniques were used extensively. Since May 1995 the Library has a growing division of images in the new Virtual Reality Modeling Language (VRML) format. The Image Library of Biological Macromolecules can be accessed via the World-Wide Web (http://www.imb-jena.de/IMAGE.html). There is a large number of structures determined by experimental and/or modeling techniques which are not intended to be included into the Protein Data Bank or Nucleic Acid Database for some reason. The Image Library could be a repository of these structures and of images of these and other structures of biological macromolecules including structures which are not known at atomic detail. Authors who are willing to make available images or coordinates to the scientific community via the Image Library of Biological Macromolecules are requested to contact the author.

Computer Communication Networks↗

Lexical training through modeling and elicitation procedures with late talkers who have specific language impairment and developmental delays.

Late talkers with specific language impairment and developmental delay make up a large portion of our early childhood caseloads; therefore, an understanding of best clinical practices for these populations is essential. Early lexical learning was examined in 2 interactive treatment approaches with 29 late-talking preschoolers with language and developmental disabilities. Children were randomly assigned to either a mand-elicited imitation (MEI) condition in which elicitations and imitative prompts were used or to a modeling with auditory bombardment (Mod-AB) condition in which auditory bombardment and play modeling were incorporated with no response demands on participants. Lexical production of target vocabulary words already comprehended was measured during a 10-session training period and then during two 50-min play interactions with a parent/caretaker in the home after treatment was completed. Results indicated that the MEI procedure was relatively more effective in facilitating frequency and rate of target word learning in the treatment setting, but no significant differences were found between conditions in the number or percentage of target words generalized to the home setting. Mod-AB children produced more target words that were limited to the home setting than did MEI children, whose productivity was more balanced across settings. Treatment by aptitude regression analyses indicated that none of the preintervention language, cognitive, or total development aptitude scores were predictive of child performance in 1 treatment condition or the other, although Battelle Developmental Inventory communication scores and sizes of preintervention lexicons were predictive of child performance across conditions. Empirical and clinical issues pertaining to the efficacy of modeling- and elicitation-based procedures for late-talking preschoolers are discussed.

Child, Preschool↗