Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Age, cohort, and time development muddles: easy in practice, hard in theory.

The debate among developmental psychologists over how best to combine longitudinal and cross-sectional data sequences can be traced back at least four decades. During the 1970s, a variation of this theme received much attention: could the developmental influences of age, cohort, and time be unraveled by sufficiently ingenious application of combined data sequences? We believe discussion of this question has been needlessly parochial and confused. Substantive and methodological contributions from other disciplines have, until recently, been largely ignored by developmental psychologists. Moreover, the solutions debated by psychologists have generally been formulated in language that obscured, rather than explicated, the formal indeterminacy implicit in models of age, cohort, and time parameters. Methodologists have outlined formal solutions to problems of indeterminacy in the contexts of model identification and estimability theory. Sociologists have proposed solutions, based on explicitly theoretical assumptions, that permit model identification and unambiguous interpretation. We review these contributions, and suggest a hierarchy of solutions to the problem of age, cohort, and time indeterminacy.

Aging↗

A conceptual model for information retrieval with UMLS.

Information retrieval in large information databases is a non-deterministic process which needs a sequence of search steps generally. One of the main problems to which the end-users are faced is to parse efficiently their questions into the query language that the computer systems allow. Conceptual graphs were initially designed for natural language analysis and understanding. Due to their closeness to semantic networks, their expressiveness is powerful enough to be applied to knowledge representation and use by computer systems. This work demonstrates that conceptual graphs are a suitable means to model the end-users querieson the basis of the thesaurus and the semantic network of the UMLS project.

Information Storage and Retrieval↗

The size distribution of conspecific populations: the peoples of New Guinea.

The size distribution of the language populations in New Guinea, which represent over 15% of the world's languages, is analysed using models analogous to the resource division models of species abundance distribution in ecological communities. A model distribution of resource segments reflecting population size is created by repeated selection of an existing resource segment and its division into two. We found that any dependency of the selection probability on the size of the segment generated negatively skewed abundance distributions after log transformation. Asymmetric segment division further exacerbated the negative skewness. Size-independent selection produced lognormal abundance distributions, irrespective of the segment division method. Size-dependent selection and asymmetric division were deemed reasonable assumptions since large language populations are more likely to generate isolates, which develop into new populations, than small ones, and these isolates are likely to be small relative to the progenitor population. A negatively skewed distribution of the log-transformed population sizes was therefore expected. However, the observed distributions were lognormal, scale invariant for areas containing between 100 and over 1000 language populations. The dynamics of language differentiation, as reflected by the models, may therefore be unimportant relative to the effect of variable growth rates among populations. All lognormal distributions from resource division models had a higher variance than the observed one, where half of the 1053 populations had between 350 and 3000 individuals. The possible mechanisms maintaining such a low variance around a modal population size of 1000 are discussed.

Algorithms↗

Reengineering the pharmaceutical industry by crash-testing molecules.

The recent decline in drug approvals and the increase in late-stage failures indicate that the ability to generate and screen large numbers of molecules has not improved the drug pipeline. Perhaps the pharmaceutical industry should follow the example of the automotive industry and agree upon a shared modeling language with vendors and academics to enable integration of predictive computational tools across the industry. This will then enable the virtual 'crash-testing' of drugs before synthesis, biological testing and, most importantly, clinical trials. This represents an ambitiously progressive approach using the models for simulating every stage of the drug discovery and development process. Combining the relevant computational algorithms into a grand unified model would enable prioritization of the best ideas before pursuing a discovery program, selecting a target or synthesizing a molecule. The successful application of these virtual crash-testing principles by any of its current proponents could revitalize the pharmaceutical industry so that failure is avoided.

Chemistry, Pharmaceutical↗

Language deficits, localization, and grammar: evidence for a distributive model of language breakdown in aphasic patients and neurologically intact individuals.

Selective deficits in aphasic patients' grammatical production and comprehension are often cited as evidence that syntactic processing is modular and localizable in discrete areas of the brain (e.g., Y. Grodzinsky, 2000). The authors review a large body of experimental evidence suggesting that morpho-syntactic deficits can be observed in a number of aphasic and neurologically intact populations. They present new data showing that receptive agrammatism is found not only over a range of aphasic groups, but is also observed in neurologically intact individuals processing under stressful conditions. The authors suggest that these data are most compatible with a domain-general account of language, one that emphasizes the interaction of linguistic distributions with the properties of an associative processor working under normal or suboptimal conditions.

Adult↗

Morphological units in the Arabic mental lexicon: evidence from an individual with deep dyslexia.

An ongoing debate in Arabic morphology concerns the nature of the smallest unit governing lexical organization and representation in this language. A standard model maintains that Arabic words are typically analyzable into a three-consonantal root morpheme carrying the core meaning of words and a prosodic template responsible mostly for grammatical information. This view has been largely supported by research in both theoretical linguistics and psycholinguistics. An alternative theory holds that the meaning of words in Arabic is, rather, encoded in the 'etymon' comprising two unordered consonants of the root only. Results from a recent priming experiment have shown that the etymon induces strong morphological priming effects, supporting its morphological/lexical status. In this paper we present data from a patient with deep dyslexia questioning the role of the etymon as a psychologically real representational unit in Arabic and arguing, instead, for the central role of the root in both morphological and lexical representation in this language.

Adult↗

NLP techniques associated with the OpenGALEN ontology for semi-automatic textual extraction of medical knowledge: abstracting and mapping equivalent linguistic and logical constructs.

This research project presents methodological and theoretical issues related to the inter-relationship between linguistic and conceptual semantics, analysing the results obtained by the application of a NLP parser to a set of radiology reports. Our objective is to define a technique for associating linguistic methods with domain specific ontologies for semi-automatic extraction of intermediate representation (IR) information formats and medical ontological knowledge from clinical texts. We have applied the Edinburgh LTG natural language parser to 2810 clinical narratives describing radiology procedures. In a second step, we have used medical expertise and ontology formalism for identification of semantic structures and abstraction of IR schemas related to the processed texts. These IR schemas are an association of linguistic and conceptual knowledge, based on their semantic contents. This methodology aims to contribute to the elaboration of models relating linguistic and logical constructs based on empirical data analysis. Advance in this field might lead to the development of computational techniques for automatic enrichment of medical ontologies from real clinical environments, using descriptive knowledge implicit in large text corpora sources.

Electronic Data Processing↗

Design of highly functional genome editors by modelling CRISPR-Cas sequences.

Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology and human health. CRISPR-based gene editors derived from microorganisms, although powerful, often show notable functional tradeoffs when ported into non-native environments, such as human cells1. Artificial-intelligence-enabled design provides a powerful alternative with the potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models2 trained on biological diversity at scale, we demonstrate successful precision editing of the human genome with a programmable gene editor designed with artificial intelligence. To achieve this goal, we curated a dataset of more than 1 million CRISPR operons through systematic mining of 26 terabases of assembled genomes and metagenomes. We demonstrate the capacity of our models by generating 4.8× the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate that an artificial-intelligence-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 to facilitate broad, ethical use across research and commercial applications.

CRISPR-Cas Systems↗

When novel sentences spoken or heard for the first time in the history of the universe are not enough: toward a dual-process model of language.

Although interest in the language sciences was previously focused on newly created sentences, more recently much attention has turned to the importance of formulaic expressions in normal and disordered communication. Also referred to as formulaic expressions and made up of speech formulas, idioms, expletives, serial and memorized speech, slang, sayings, clichés, and conventional expressions, non-propositional language forms a large proportion of every speaker's competence, and may be differentially disturbed in neurological disorders. This review aims to examine non-propositional speech with respect to linguistic descriptions, psycholinguistic experiments, sociolinguistic studies, child language development, clinical language disorders, and neurological studies. Evidence from numerous sources reveals differentiated and specialized roles for novel and formulaic verbal functions, and suggests that generation of novel sentences and management of prefabricated expressions represent two legitimate and separable processes in language behaviour. A preliminary model of language behaviour that encompasses unitary and compositional properties and their integration in everyday language use is proposed. Integration and synchronizing of two disparate processes in language behaviour, formulaic and novel, characterizes normal communicative function and contributes to creativity in language. This dichotomy is supported by studies arising from other disciplines in neurology and psychology. Further studies are necessary to determine in what ways the various categories of formulaic expressions are related, and how these categories are processed by the brain. Better understanding of how non-propositional categories of speech are stored and processed in the brain can lead to better informed treatment strategies in language disorders.

Aphasia↗

Evolutionary game dynamics in finite populations with strong selection and weak mutation.

We study stochastic game dynamics in finite populations. To this end we extend the classical Moran process to incorporate frequency-dependent selection and mutation. For 2 x 2 games, we give a complete analysis of the long-run behavior when mutation rates are small. For 3 x 3 coordination games, we provide a simple rule to determine which strategy will be selected in large populations. The expected motion in our model resembles the standard replicator dynamics when the population is large, but is qualitatively different when the population is small. Our analysis shows that even in large finite populations the behavior of a replicator-like system can be different from that of the standard replicator dynamics. As an application, we consider selective language dynamics. We determine which language will be spoken in finite large populations. The results have an intuitive interpretation but would not be expected from an analysis of the replicator dynamics.

Biological Evolution↗

Pleiotropy and preadaptation in the evolution of human language capacity.

The capacity for spoken language in the human is a genetic trait, but the information communicated by this means is to a large extent culturally determined. Using a gene-culture coevolutionary approach, we model the hypothesis that speech evolved as a channel for the communication of adaptive cultural traits from parent to offspring. The motivation for this paper is a condition obtained previously that initial increase of communication would require at least a two-fold advantage for the transmitted trait. Here, we show that under reasonable assumptions the invasion condition becomes less stringent. In Model 1, we assume that two adaptive cultural traits can be transmitted. A gene which permits communication of the second adaptive trait. In Model 2, we assume that a related function such as greater memory capacity is a prerequisite for speech, and that this function confers an advantage independent of its association with speech. In both models we assume haploid sexual genetics and a simple scheme of vertical transmission. The stability properties of all corner and edge equilibria of the models are analyzed. The two models taken together suggest a possible scenario for the initial stages of the evolution of speech.

Adaptation, Biological↗

Brain network interactions in auditory, visual and linguistic processing.

In the paper, we discuss the importance of network interactions between brain regions in mediating performance of sensorimotor and cognitive tasks, including those associated with language processing. Functional neuroimaging, especially PET and fMRI, provide data that are obtained essentially simultaneously from much of the brain, and thus are ideal for enabling one to assess interregional functional interactions. Two ways to use these types of data to assess network interactions are presented. First, using PET, we demonstrate that anterior and posterior perisylvian language areas have stronger functional connectivity during spontaneous narrative production than during other less linguistically demanding production tasks. Second, we show how one can use large-scale neural network modeling to relate neural activity to the hemodynamically-based data generated by fMRI and PET. We review two versions of a model of object processing - one for visual and one for auditory objects. The regions comprising the models include primary and secondary sensory cortex, association cortex in the temporal lobe, and prefrontal cortex. Each model incorporates specific assumptions about how neurons in each of these areas function, and how neurons in the different areas are interconnected with each other. Each model is able to perform a delayed match-to-sample task for simple objects (simple shapes for the visual model; tonal contours for the auditory model). We find that the simulated electrical activities in each region are similar to those observed in nonhuman primates performing analogous tasks, and the absolute values of the simulated integrated synaptic activity in each brain region match human fMRI/PET data. Thus, this type of modeling provides a way to understand the neural bases for the sensorimotor and cognitive tasks of interest.

Animals↗

International development of the Quality of Life in Depression Scale (QLDS).

BACKGROUND: The Quality of Life in Depression Scale (QLDS) employs the needs-based model of quality of life (QoL) and was developed in the UK and The Netherlands as an outcome measure for clinical trials. This paper describes the production and psychometric assessment of nine new language versions for Canada (French and English), Denmark, France, Germany, Italy, Morocco, Spain and the US. METHODS: Three adaptation stages were employed; production of conceptually equivalent translations, field-test interviews and assessment of reliability and construct validity by survey of patients with major depression. RESULTS: Few problems were experienced with producing conceptually equivalent translations, except in Morocco. Patients in the field-test interviews found the instrument to have appropriate content and to be easy to complete. Internal consistency and test-retest reliability were excellent for all language versions and scores were found to relate appropriately to measures of depression severity and health status. LIMITATIONS: Further investigation is required of the ability of the measure to assess individuals at the extremes of the QoL continuum. Data collected with the Arabic QLDS should not be combined with those from other countries. CONCLUSIONS: The QLDS is the first instrument designed to assess QoL in depression based on a coherent model of the construct. Each language version has been shown to be well accepted by respondents and to have excellent psychometric properties. As the instrument is now available in a large number of languages, the QLDS is the QoL instrument of choice for inclusion in clinical trials of interventions for depression.

Adult↗

Simplest random K-satisfiability problem.

We study a simple and exactly solvable model for the generation of random satisfiability problems. These consist of gammaN random boolean constraints which are to be satisfied simultaneously by N logical variables. In statistical-mechanics language, the considered model can be seen as a diluted p-spin model at zero temperature. While such problems become extraordinarily hard to solve by local search methods in a large region of the parameter space, still at least one solution may be superimposed by construction. The statistical properties of the model can be studied exactly by the replica method and each single instance can be analyzed in polynomial time by a simple global solution method. The geometrical and topological structures responsible for dynamic and static phase transitions as well as for the onset of computational complexity in the local search method are thoroughly analyzed. Numerical analysis on very large samples allows for a precise characterization of the critical scaling behavior.

Journal Article↗

A short history of data banking in the United States from 1974 to 2003.

There have been 4 major longitudinal data banking efforts within the United States: ARAMIS, the Western Consortium, and the individual data banks of Drs. Ted Pincus and Fred Wolfe. ARAMIS began in the 1970s, and helped to develop the language and methodology of rheumatology data banks using biannual surveys. The National Data Bank for Rheumatic Diseases used the ARAMIS model beginning in the late 1990s to form a very large contemporary rheumatology data bank. Hybrid models using both survey data and clinical data were put into practice by Pincus and Wolfe, and by Paulus at the Western Consortium.

Arthritis, Rheumatoid↗

[Intercultural communication in general practice].

What is the reason for possible communication problems between physicians and immigrants from the third world? How is the interaction between the two groups regulated? To answer such questions, 15 doctors and 10 immigrants were interviewed about their experience of the doctor-patient interaction in unstructured open-ended interviews. Culturally based verbal and non-verbal expressions were particularly difficult to interpret, being based on different thought models and language. Owing to the world wide generalisation of the doctor-patient roles the gap between the two "partners" in the communication has been to some extent bridged. Largely independent of the doctors' will, the immigrants assigned considerable authority to the doctors. Thus the power of the doctor is based on the institutionalisation of the more universal doctor-patient role.

Communication↗

GOFCOX: a computer program for the goodness-of-fit analysis of the Cox proportional hazards model.

GOFCOX is a user-friendly FORTRAN program for assessing the adequacy of the Cox proportional hazards model. The underlying methodology is based on the comparison of the maximum partial likelihood estimator and a weighted parameter estimator. The latter is the root to an estimation equation that assigns varying weights to the individual contributions to the partial likelihood score function. The weighted and unweighted parameter estimators have the same expectation under the Cox model, but tend to differ when the model is inappropriate. The GOFCOX program computes a rich class of weighted parameter estimators and corresponding goodness-of-fit test statistics. The program runs on both mainframe computers and microcomputers. The running time is minimal even for large data sets. A simple example is provided to illustrate the features of the program.

Computers, Mainframe↗