Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Lexical training through modeling and elicitation procedures with late talkers who have specific language impairment and developmental delays.

Late talkers with specific language impairment and developmental delay make up a large portion of our early childhood caseloads; therefore, an understanding of best clinical practices for these populations is essential. Early lexical learning was examined in 2 interactive treatment approaches with 29 late-talking preschoolers with language and developmental disabilities. Children were randomly assigned to either a mand-elicited imitation (MEI) condition in which elicitations and imitative prompts were used or to a modeling with auditory bombardment (Mod-AB) condition in which auditory bombardment and play modeling were incorporated with no response demands on participants. Lexical production of target vocabulary words already comprehended was measured during a 10-session training period and then during two 50-min play interactions with a parent/caretaker in the home after treatment was completed. Results indicated that the MEI procedure was relatively more effective in facilitating frequency and rate of target word learning in the treatment setting, but no significant differences were found between conditions in the number or percentage of target words generalized to the home setting. Mod-AB children produced more target words that were limited to the home setting than did MEI children, whose productivity was more balanced across settings. Treatment by aptitude regression analyses indicated that none of the preintervention language, cognitive, or total development aptitude scores were predictive of child performance in 1 treatment condition or the other, although Battelle Developmental Inventory communication scores and sizes of preintervention lexicons were predictive of child performance across conditions. Empirical and clinical issues pertaining to the efficacy of modeling- and elicitation-based procedures for late-talking preschoolers are discussed.

Child, Preschool↗

Probabilistic models of language processing and acquisition.

Probabilistic methods are providing new explanatory approaches to fundamental cognitive science questions of how humans structure, process and acquire language. This review examines probabilistic models defined over traditional symbolic structures. Language comprehension and production involve probabilistic inference in such models; and acquisition involves choosing the best model, given innate constraints and linguistic and other input. Probabilistic models can account for the learning and processing of language, while maintaining the sophistication of symbolic models. A recent burgeoning of theoretical developments and online corpus creation has enabled large models to be tested, revealing probabilistic constraints in processing, undermining acquisition arguments based on a perceived poverty of the stimulus, and suggesting fruitful links with probabilistic theories of categorization and ambiguity resolution in perception.

Brain↗

Accelerating inference in genomic and proteomic foundation models via speculative decoding.

MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.

Genomics↗

Computational physiology and the Physiome Project.

Bioengineering analyses of physiological systems use the computational solution of physical conservation laws on anatomically detailed geometric models to understand the physiological function of intact organs in terms of the properties and behaviour of the cells and tissues within the organ. By linking behaviour in a quantitative, mathematically defined sense across multiple scales of biological organization--from proteins to cells, tissues, organs and organ systems--these methods have the potential to link patient-specific knowledge at the two ends of these spatial scales. A genetic profile linked to cardiac ion channel mutations, for example, can be interpreted in relation to body surface ECG measurements via a mathematical model of the heart and torso, which includes the spatial distribution of cardiac ion channels throughout the myocardium and the individual kinetics for each of the approximately 50 types of ion channel, exchanger or pump known to be present in the heart. Similarly, linking molecular defects such as mutations of chloride ion channels in lung epithelial cells to the integrated function of the intact lung requires models that include the detailed anatomy of the lungs, the physics of air flow, blood flow and gas exchange, together with the large deformation mechanics of breathing. Organizing this large body of knowledge into a coherent framework for modelling requires the development of ontologies, markup languages for encoding models, and web-accessible distributed databases. In this article we review the state of the field at all the relevant levels, and the tools that are being developed to tackle such complexity. Integrative physiology is central to the interpretation of genomic and proteomic data, and is becoming a highly quantitative, computer-intensive discipline.

Biophysical Phenomena↗

mGrid: a load-balanced distributed computing environment for the remote execution of the user-defined Matlab code.

BACKGROUND: Matlab, a powerful and productive language that allows for rapid prototyping, modeling and simulation, is widely used in computational biology. Modeling and simulation of large biological systems often require more computational resources then are available on a single computer. Existing distributed computing environments like the Distributed Computing Toolbox, MatlabMPI, Matlab*G and others allow for the remote (and possibly parallel) execution of Matlab commands with varying support for features like an easy-to-use application programming interface, load-balanced utilization of resources, extensibility over the wide area network, and minimal system administration skill requirements. However, all of these environments require some level of access to participating machines to manually distribute the user-defined libraries that the remote call may invoke. RESULTS: mGrid augments the usual process distribution seen in other similar distributed systems by adding facilities for user code distribution. mGrid's client-side interface is an easy-to-use native Matlab toolbox that transparently executes user-defined code on remote machines (i.e. the user is unaware that the code is executing somewhere else). Run-time variables are automatically packed and distributed with the user-defined code and automated load-balancing of remote resources enables smooth concurrent execution. mGrid is an open source environment. Apart from the programming language itself, all other components are also open source, freely available tools: light-weight PHP scripts and the Apache web server. CONCLUSION: Transparent, load-balanced distribution of user-defined Matlab toolboxes and rapid prototyping of many simple parallel applications can now be done with a single easy-to-use Matlab command. Because mGrid utilizes only Matlab, light-weight PHP scripts and the Apache web server, installation and configuration are very simple. Moreover, the web-based infrastructure of mGrid allows for it to be easily extensible over the Internet.

Computational Biology↗

Syntactic-semantic tagging as a mediator between linguistic representations and formal models: an exercise in linking SNOMED to GALEN.

Natural language understanding applications are good candidates to solve the knowledge acquisition bottleneck when designing large scale concept systems. However, a necessary condition is that systems are built that transform sentences into a meaning representation that is independent of the subtleties of linguistic structure that nevertheless underly the way language works. The Cassandra II syntactic-semantic tagging system fulfills this goal partially. Within the GALEN-IN-USE project, it is used to transform linguistic representations of surgical procedure expressions into conceptual representations. In this paper, the proctology chapter of the SNOMED V3.1 procedure axis was used as a testbed to evaluate the usefulness of this approach. A quantitative and qualitative analysis of the data obtained is presented, showing that the Cassandra system can indeed complement the manual modelling efforts being conducted in the GALEN-IN-USE project. The different requirements related to linguistic modelling versus conceptual modelling can partly be accounted for by using an interface ontology, of which the fine tuning will however remain an important effort.

Artificial Intelligence↗

State-dependent decisions cause apparent violations of rationality in animal choice.

Normative models of choice in economics and biology usually expect preferences to be consistent across contexts, or "rational" in economic language. Following a large body of literature reporting economically irrational behaviour in humans, breaches of rationality by animals have also been recently described. If proven systematic, these findings would challenge long-standing biological approaches to behavioural theorising, and suggest that cognitive processes similar to those claimed to cause irrationality in humans can also hinder optimality approaches to modelling animal preferences. Critical differences between human and animal experiments have not, however, been sufficiently acknowledged. While humans can be instructed conceptually about the choice problem, animals need to be trained by repeated exposure to all contingencies. This exposure often leads to differences in state between treatments, hence changing choices while preserving rationality. We report experiments with European starlings demonstrating that apparent breaches of rationality can result from state-dependence. We show that adding an inferior alternative to a choice set (a "decoy") affects choices, an effect previously interpreted as indicating irrationality. However, these effects appear and disappear depending on whether state differences between choice contexts are present or not. These results open the possibility that some expressions of maladaptive behaviour are due to oversights in the migration of ideas between economics and biology, and suggest that key differences between human and nonhuman research must be recognised if ideas are to safely travel between these fields.

Analysis of Variance↗

NeuronC: a computational language for investigating functional architecture of neural circuits.

A computational language was developed to simulate neural circuits. A model of a neural circuit with up to 50,000 compartments is constructed from predefined parts of neurons, called "neural elements". A 2-dimensional (2-D) light stimulus and a photoreceptor model allow simulating a visual physiology experiment. Circuit function is computed by integrating difference equations according to standard methods. Large-scale structure in the neural circuit, such as whole neurons, their synaptic connections, and arrays of neurons, are constructed with procedural rules. The language was evaluated with a simulation of the receptive field of a single cone in cat retina, which required a model of cone-horizontal cell network on the order of 1000 neurons. The model was calibrated by adjusting biophysical parameters to match known physiological data. Eliminating specific synaptic connections from the circuit suggested the influence of individual neuron types on the receptive field of a single cone. An advantage of using neural elements in such a model is to simplify the description of a neuron's structure. An advantage of using procedural rules to define connections between neurons is to simplify the network definition.

Computer Simulation↗

CodonMoE: DNA language models for codon-dependent mRNA prediction.

MOTIVATION: Genomic language models (gLMs) face a fundamental efficiency challenge: one must either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multimodal architectures. Both approaches impose significant computational burdens-modality-specific models require redundant infrastructure despite inherent biological connections, while multi-modal architectures demand increased parameter counts and extensive cross-modality pretraining. RESULTS: To address this limitation, we introduce CodonMoE (Adaptive Mixture of Codon Reformative Experts), a lightweight adapter that transforms DNA language models into effective RNA analyzers without RNA-specific pretraining. Our theoretical analysis establishes CodonMoE as a universal approximator at the codon level, capable of mapping arbitrary functions from codon sequences to codon-dependent RNA properties given sufficient expert capacity. Across four RNA prediction tasks spanning stability, expression, and regulation, DNA models augmented with CodonMoE significantly outperform their unmodified counterparts, with the HyenaDNA+CodonMoE series achieving state-of-the-art results using 80% fewer parameters than specialized RNA models. By maintaining sub-quadratic complexity while achieving superior performance, our approach provides a principled path toward unifying genomic language modeling, leveraging more abundant DNA data and reducing computational overhead while preserving modality-specific performance advantages. AVAILABILITY AND IMPLEMENTATION: Source code for the method and to reproduce the results is available at https://github.com/Kingsford-Group/CodonMoE.

Codon↗

CeLLTra: aligning cell names with gene expression via a pathway-informed transformer.

MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans↗

A Canadian perspective on learning disabilities.

Canadian practice and research with children and adults with learning disabilities are described and analyzed. After an examination of the historical basis for current practice, the societal and cultural factors affecting education of children with learning disabilities, services for adults, and research are discussed. It was found that policy and legislation regarding special education vary considerably from province to province, and identification practices and service delivery models vary even within provinces. The fact that Canada has two official languages (English and French), a large multicultural community, and a Native population with special needs often arising from poverty has an impact on the education of children with learning disabilities and on sample description in research. Although school-age children are relatively well served, services for preschool children and adults with learning disabilities are minimal. The positive features of Canadian service delivery are that most programs are publicly funded, decision making tends to be nonadversarial and collaborative, and the needs of the whole child are typically considered.

Adult↗

Disconnection of language and memory in semantic dementia: a comparative and theoretical analysis.

In this paper, we present an illustrative case of Semantic Dementia (SD) and we review the literature on this relatively rare progressive neurodegenerative disorder. After reviewing the clinical, neuroimaging, neuropathological, and genetic features of SD, we propose a theoretical framework that addresses features of SD and relates them to features of other well known neuropsychiatric syndromes. Our 'on-line / off-line disconnection' model seeks to conceptualize SD as a syndrome of disconnection between two large distributed cortical networks, namely, between those networks that subserve language function and those that subserve memory function.

Brain↗

Common aetiology for diverse language skills in 4 1/2-year-old twins.

Multivariate genetic analysis was used to examine the genetic and environmental aetiology of the interrelationships of diverse linguistic skills. This study used data from a large sample of 4 1/2-year-old twins who were tested on measures assessing articulation, phonology, grammar, vocabulary, and verbal memory. Phenotypic analysis suggested two latent factors: articulation (2 measures) and general language (the remaining 7), and a genetic model incorporating these factors provided a good fit to the data. Almost all genetic and shared environmental influences on the 9 measures acted through the two latent factors. There was also substantial aetiological overlap between the two latent factors, with a genetic correlation of 0.64 and shared environment correlation of 1.00. We conclude that to a large extent, the same genetic and environmental factors underlie the development of individual differences in a wide range of linguistic skills.

Chi-Square Distribution↗

Domain-specific language models and lexicons for tagging.

Accurate and reliable part-of-speech tagging is useful for many Natural Language Processing (NLP) tasks that form the foundation of NLP-based approaches to information retrieval and data mining. In general, large annotated corpora are necessary to achieve desired part-of-speech tagger accuracy. We show that a large annotated general-English corpus is not sufficient for building a part-of-speech tagger model adequate for tagging documents from the medical domain. However, adding a quite small domain-specific corpus to a large general-English one boosts performance to over 92% accuracy from 87% in our studies. We also suggest a number of characteristics to quantify the similarities between a training corpus and the test data. These results give guidance for creating an appropriate corpus for building a part-of-speech tagger model that gives satisfactory accuracy results on a new domain at a relatively small cost.

Humans↗

Logic-based remodeling of the Digital Anatomist Foundational Model.

This paper describes a development cycle for the engineering of large knowledge bases: A graphical tool is used for editing and the content is transformed into a logic-based representation language. This representation is used to check the consistency of the knowledge base as well as to facilitate the reviewing process. Showing the usefulness of this approach, aspects of the Digital Anatomist Foundational Model will be transformed into a Description Logics representation. We introduce a special modeling technique to account for the representation of the complex part/whole relationships in the biomedical domain.

Anatomy↗

A query language for biological networks.

MOTIVATION: Many areas of modern biology are concerned with the management, storage, visualization, comparison and analysis of networks, but no appropriate query language for such complex data structures yet exists. RESULTS: We have designed and implemented the pathway query language (PQL) for querying large protein interaction or pathway databases. PQL is based on a simple graph data model with extensions reflecting properties of biological objects. Queries match subgraphs in the database based on node properties and paths between nodes. The syntax is easy to learn for anybody familiar with SQL. As an important feature, a query may require a certain structure in the database to exist but return a different subgraph. We have tested PQL queries on networks of up to 16,000 nodes and found it to scale very well. AVAILABILITY: The code is available on request from the author.

Computational Biology↗

Contributions of memory circuits to language: the declarative/procedural model.

The structure of the brain and the nature of evolution suggest that, despite its uniqueness, language likely depends on brain systems that also subserve other functions. The declarative/procedural (DP) model claims that the mental lexicon of memorized word-specific knowledge depends on the largely temporal-lobe substrates of declarative memory, which underlies the storage and use of knowledge of facts and events. The mental grammar, which subserves the rule-governed combination of lexical items into complex representations, depends on a distinct neural system. This system, which is composed of a network of specific frontal, basal-ganglia, parietal and cerebellar structures, underlies procedural memory, which supports the learning and execution of motor and cognitive skills, especially those involving sequences. The functions of the two brain systems, together with their anatomical, physiological and biochemical substrates, lead to specific claims and predictions regarding their roles in language. These predictions are compared with those of other neurocognitive models of language. Empirical evidence is presented from neuroimaging studies of normal language processing, and from developmental and adult-onset disorders. It is argued that this evidence supports the DP model. It is additionally proposed that "language" disorders, such as specific language impairment and non-fluent and fluent aphasia, may be profitably viewed as impairments primarily affecting one or the other brain system. Overall, the data suggest a new neurocognitive framework for the study of lexicon and grammar.

Humans↗

Evaluation (not validation) of quantitative models.

The present regulatory climate has led to increasing demands for scientists to attest to the predictive reliability of numerical simulation models used to help set public policy, a process frequently referred to as model validation. But while model validation may reveal useful information, this paper argues that it is not possible to demonstrate the predictive reliability of any model of a complex natural system in advance of its actual use. All models embed uncertainties, and these uncertainties can and frequently do undermine predictive reliability. In the case of lead in the environment, we may categorize model uncertainties as theoretical, empirical, parametrical, and temporal. Theoretical uncertainties are aspects of the system that are not fully understood, such as the biokinetic pathways of lead metabolism. Empirical uncertainties are aspects of the system that are difficult (or impossible) to measure, such as actual lead ingestion by an individual child. Parametrical uncertainties arise when complexities in the system are simplified to provide manageable model input, such as representing longitudinal lead exposure by cross-sectional measurements. Temporal uncertainties arise from the assumption that systems are stable in time. A model may also be conceptually flawed. The Ptolemaic system of astronomy is a historical example of a model that was empirically adequate but based on a wrong conceptualization. Yet had it been computerized--and had the word then existed--its users would have had every right to call it validated. Thus, rather than talking about strategies for validation, we should be talking about means of evaluation. That is not to say that language alone will solve our problems or that the problems of model evaluation are primarily linguistic. The uncertainties inherent in large, complex models will not go away simply because we change the way we talk about them. But this is precisely the point: calling a model validated does not make it valid. Modelers and policymakers must continue to work toward finding effective ways to evaluate and judge the quality of their models, and to develop appropriate terminology to communicate these judgments to the public whose health and safety may be at stake.

Models, Biological↗