Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Evaluation (not validation) of quantitative models.

The present regulatory climate has led to increasing demands for scientists to attest to the predictive reliability of numerical simulation models used to help set public policy, a process frequently referred to as model validation. But while model validation may reveal useful information, this paper argues that it is not possible to demonstrate the predictive reliability of any model of a complex natural system in advance of its actual use. All models embed uncertainties, and these uncertainties can and frequently do undermine predictive reliability. In the case of lead in the environment, we may categorize model uncertainties as theoretical, empirical, parametrical, and temporal. Theoretical uncertainties are aspects of the system that are not fully understood, such as the biokinetic pathways of lead metabolism. Empirical uncertainties are aspects of the system that are difficult (or impossible) to measure, such as actual lead ingestion by an individual child. Parametrical uncertainties arise when complexities in the system are simplified to provide manageable model input, such as representing longitudinal lead exposure by cross-sectional measurements. Temporal uncertainties arise from the assumption that systems are stable in time. A model may also be conceptually flawed. The Ptolemaic system of astronomy is a historical example of a model that was empirically adequate but based on a wrong conceptualization. Yet had it been computerized--and had the word then existed--its users would have had every right to call it validated. Thus, rather than talking about strategies for validation, we should be talking about means of evaluation. That is not to say that language alone will solve our problems or that the problems of model evaluation are primarily linguistic. The uncertainties inherent in large, complex models will not go away simply because we change the way we talk about them. But this is precisely the point: calling a model validated does not make it valid. Modelers and policymakers must continue to work toward finding effective ways to evaluate and judge the quality of their models, and to develop appropriate terminology to communicate these judgments to the public whose health and safety may be at stake.

Models, Biological↗

Observations on the recent history of drug user counseling.

The drug user counselor role is explored in terms of its changing nature over the course of the past 25 years. Initially, the drug user counselor could be characterized as a professional based on his or her experience, or as an ex-addict paraprofessional in the language of that time. Working very largely with a heroin-using clientele, the counselor was the advisor and role model who could not be conned and, thereby, the essential counterpart to the mental health professionals who were entering the drug misuse field. Over time, these latter professionals based on education have become increasingly evident and the "professionals of experience" have become less so in accord with changes in the demography, drug-using characteristics, and psychological functioning of drug user clients. Nonetheless, studies that support the particular efficacy of counselors of education for all but drug user clients with significant psychopathology are lacking. Moreover, aspects of therapeutic interaction that are more largely engaged in by "professionals of experience" are threatened by the diminution in that group's numbers and the credentialing out of nontraditional job functions. Over the past few years, awareness of the significance of the contributions of "professionals of experience" has been reawakened by the threat of AIDS and the recognition of counselors' contributions to outreach and AIDS prevention counseling.

Acquired Immunodeficiency Syndrome↗

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems↗

Language trees support the express-train sequence of Austronesian expansion.

Languages, like molecules, document evolutionary history. Darwin observed that evolutionary change in languages greatly resembled the processes of biological evolution: inheritance from a common ancestor and convergent evolution operate in both. Despite many suggestions, few attempts have been made to apply the phylogenetic methods used in biology to linguistic data. Here we report a parsimony analysis of a large language data set. We use this analysis to test competing hypotheses--the "express-train" and the "entangled-bank" models--for the colonization of the Pacific by Austronesian-speaking peoples. The parsimony analysis of a matrix of 77 Austronesian languages with 5,185 lexical items produced a single most-parsimonious tree. The express-train model was converted into an ordered geographical character and mapped onto the language tree. We found that the topology of the language tree was highly compatible with the express-train model.

Archaeology↗

Language disorder: a functional linguistic perspective.

This paper explores the issues involved in the linguistic characterisation of disordered discourse and the ways in which a Systemic Functional Linguistic framework addresses these issues. For many years, language disorders were described in terms of formal grammars, with "breakdown" discussed in terms of one or more of the traditional levels of language, i.e., phonology, syntax, and semantics. While it was acknowledged that an individual could have difficulty at one or more of these levels, each was viewed quite separately, with semantics viewed largely from a referential perspective. More recent approaches using functional grammar have broadened this view of language and have provided a model of language that re-conceptualizes the notion of meaning and embraces context as integral to its organisation. Such a model has introduced a different perspective on language into clinical fields, and has enabled researchers and clinicians to explore the skills of speakers with language disorders across a variety of situations and contextual variables, examining the linguistic resources still available to them. This paper introduces principles involved in a functional framework and provides an overview of how these principles have been applied to language disorders to date. In addition, the notion of "disorder" itself is discussed as it is situated in this alternative model.

Aphasia↗

A nonlinear electrical-thermal model of the skin.

This work presents a model for the skin which accounts for both the nonlinearities and the asymmetries in its voltage-current characteristic. This model consists of an electrical submodel and a heat transfer submodel. The electrical submodel uses nonlinear devices in which some parameters depend on skin temperature. The heat transfer submodel models the heat exchange between the skin, the surrounding tissues, and the ambient medium and calculates the temperature of the skin to update the necessary parameters of the electrical submodel. The model is based on experiments designed to determine: 1) the dry skin voltage-current characteristic; 2) the changes in the skin breakdown voltage with location; 3) the moist skin voltage-current characteristic; 4) the changes in the voltage-current characteristic of the skin with duration after the onset of stimulation; and 5) the effect of skin temperature on its voltage-current characteristic. During these experiments we used 84-mm2 square Ag-AgCl electrodes to apply sinusoidal voltage of 0.2 and 20 Hz. The simulations were performed using the Advanced Continuous Simulation Language (ACSL), capable of solving differential and integral equations with variable coefficients. The model predicted the skin behavior satisfactorily for a large range of amplitudes and frequencies. We found that the breakdown occurred when the energy delivered to the skin exceeded a threshold. Above this threshold the voltage-current characteristic of the skin became nonlinear and asymmetric and, in a real situation, the subject would experience an uncomfortable sensation which could rapidly develop into pain.

Electric Conductivity↗

Semantic similarity measures as tools for exploring the gene ontology.

Many bioinformatics resources hold data in the form of sequences. Often this sequence data is associated with a large amount of annotation. In many cases this data has been hard to model, and has been represented as scientific natural language, which is not readily computationally amenable. The development of the Gene Ontology provides us with a more accessible representation of some of this data. However it is not clear how this data can best be searched, or queried. Recently we have adapted information content based measures for use with the Gene Ontology (GO). In this paper we present detailed investigation of the properties of these measures, and examine various properties of GO, which may have implications for its future design.

Classification↗

The countering of overgeneralization.

Commenting on Goldberg's (1995) 'construction grammar', Tomasello (1998) proposes a model of language acquisition in which children move from highly specific utterance-event pairings to abstract, verb-general structures. Despite their many strengths, models of this kind predict considerably more overgeneralization of the argument structures of verbs than seems to occur. In recognition of this, the paper explains (and supports with data from a previously unpublished study of 44 children aged 2;0 to 4;4) how processes which are side effects of the emergence of the verb form class could counter the overgeneralizing tendencies. It is argued that these processes are consistent not just with the model proposed by Tomasello but also (in large part) with the grammatical theory developed by Goldberg.

Child Language↗

Readiness to change in a clinical sample of problem drinkers: relation to alcohol use, self-efficacy, and treatment outcome.

According to the transtheoretical model of behaviour change, individuals addicted to psychotropic drugs typically cycle through a sequence of five discrete stages (precontemplation, contemplation, preparation, action, and maintenance) before achieving sustained long-term abstinence and moderation, respectively. A number of English-language questionnaires have been developed to assess client motivation in accordance with the stages of change approach. The present study aimed to expand the research on the transtheoretical model by establishing the factor structure of a German-language version of the Stages of Change Readiness and Treatment Eagerness Scale (SOCRATES) in a large sample of alcohol-dependent inpatients (n = 350). Furthermore, the relation of client motivation to alcohol use, self-efficacy and treatment outcome at 3-month follow-up was examined. Exploratory factor analysis revealed three separate dimensions of readiness to change (Taking Steps, Recognition, and Ambivalence). The factorial structure of the German-language SOCRATES corresponded almost exactly to that of the original version. Readiness to change accounted for 9.4% of the variance in treatment outcome. Moreover, readiness to change was positively related to pretreatment self-efficacy.

Adult↗

Ask a silly question: two decades of troublesome trials.

BACKGROUND: Randomized control trials and the use of meta-analysis in systematic reviews are the basis of evidence-based practice. The paper reviews their use in the development of evidence-based practice in speech and language therapy. AIMS: It is accepted that clinical outcome research should develop in a sequence of phases. A model of this process is described. Examples of outcome research in speech and language therapy are used to illustrate the use of the model and the problems that result when it is not followed. MAIN CONTRIBUTION: Existing research has largely ignored the agreed procedures for outcome research. Particular problems have arisen when randomized control trials are used to examine therapy provision for a client group. Clients are often a heterogeneous group and receive different therapies. Consequently, it is unlikely that trials can obtain significant results, nor, if they do, can they provide clinicians with useful information about the choice of treatment. Systematic reviews are equally uninformative. Many of the studies on which they are based have methodological problems and their frequent failure adequately to describe the therapies used mean that reviews cannot evaluate or compare different types of therapy. CONCLUSIONS: Researchers in speech and language therapy have given too little attention to the basics of clinical outcome research. This requires that clinical and theoretical insights are used to identify specific therapies for well-defined groups of clients. These therapies must be tested first in efficacy, then in effectiveness studies, and their results should be disseminated to clinicians. Only then is it meaningful to carry out large-scale trials of the effectiveness of therapy provision for a client group or to conduct systematic reviews of existing research.

Child↗

The use of VRML in chemical education.

The Virtual Reality Modelling Language can be used for molecular modelling. The language provides some significant advantages in the area of chemical education; it can be used to communicate 3D concepts not normally covered by existing modelling packages, the data can be distributed to a large number of students over the web, and the viewers are free to students. The strengths and weakness of VRML in various aspects of molecular modelling are discussed.

Chemistry↗

Model validation software for classification models using repeated partitioning: MVREP.

The process of assessing the prediction ability of a computational model is called model validation. For models predicting a categorical response, the prediction ability is usually quantified by prediction measures such as sensitivity, specificity, and accuracy. This paper presents a software Model Validation using Repeated Partitioning (MVREP) that implements a computer-intensive, nonparametric approach to model validation, which we call the re-partitioning method. MVREP, developed using the SAS Macro language, repeats the process of randomly partitioning a dataset and subsequently performing standard model validation procedures, such as cross-validation, a large number of times and generates the empirical sampling distributions of prediction measures. The means of the sampling distributions serve as the point estimates of prediction measures of the model. The variances of the sampling distributions provide a direct assessment of variability for the point estimates of prediction measures. An example is presented using a mouse developmental toxicity chemical dataset to illustrate how the software can be used for the assessment of structure-activity relationships models.

Animals↗

Medical image databases: a content-based retrieval approach.

Information contained in medical images differs considerably from that residing in alphanumeric format. The difference can be attributed to four characteristics: (1) the semantics of medical knowledge extractable from images is imprecise; (2) image information contains form and spatial data, which are not expressible in conventional language; (3) a large part of image information is geometric; (4) diagnostic inferences derived from images rest on an incomplete, continuously evolving model of normality. This paper explores the differentiating characteristics of text versus images and their impact on design of a medical image database intended to allow content-based indexing and retrieval. One strategy for implementing medical image databases is presented, which employs object-oriented iconic queries, semantics by association with prototypes, and a generic schema.

Abstracting and Indexing↗

Automatic structuring of radiology free-text reports.

A natural language processor was developed that automatically structures the important medical information (eg, the existence, properties, location, and diagnostic interpretation of findings) contained in a radiology free-text document as a formal information model that can be interpreted by a computer program. The input to the system is a free-text report from a radiologic study. The system requires no reporting style changes on the part of the radiologist. Statistical and machine learning methods are used extensively throughout the system. A graphical user interface has been developed that allows the creation of hand-tagged training examples. Various aspects of the difficult problem of implementing an automated structured reporting system have been addressed, and the relevant technology is progressing well. Extensible Markup Language is emerging as the preferred syntactic standard for representing and distributing these structured reports within a clinical environment. Early successes hold out hope that similar statistically based models of language will allow deep understanding of textual reports. The success of these statistical methods will depend on the availability of large numbers of high-quality training examples for each radiologic subdomain. The acceptability of automated structured reporting systems will ultimately depend on the results of comprehensive evaluations.

Humans↗

A commercial large-vocabulary discrete speech recognition system: DragonDictate.

DragonDictate is currently the only commercially available general-purpose, large-vocabulary speech recognition system. It uses discrete speech and is speaker-dependent, adapting to the speaker's voice and language model with every word. Its acoustic adaptability is based in a three-level phonology and a stochastic model of production. The phonological levels are phonemes, augmented triphones (phonemes-in-context or PICs), and steady-state spectral slices that are concatenated to approximate the spectra of these PICs (phonetic elements or PELs) and thus of words. Production is treated as a hidden Markov process, which the recognizer has to identify from its output, the spoken word. Findings of practical value to speech recognition are presented from research on six European languages.

Female↗

Changing health inequalities in the Nordic countries?

The Nordic countries, referring here to Denmark, Finland, Norway, and Sweden, have often been viewed as a group of countries with many features in common, such as geographical location, history, culture, religion, language, and economic and political structures. It has also been habitual to refer to a "Nordic model" of welfare states comprising a large public sector, active labour market policies, high costs for social welfare as well as high taxes, and a general commitment to social equality. Recent research suggests that much of this "Nordicness" appears to remain despite the fact that the Nordic countries have experienced quite different changes during the 1980s and 1990s. How this relates to changes in health inequalities is in the focus of this supplement.

Finland↗

Generating patient-specific interactive natural language explanations.

Patient compliance is a significant problem and is strongly correlated with the patients' understanding of their condition and prescribed treatment. Since doctors typically do not have large amounts of time to educate patients, and impersonal, voluminous patient handouts are largely ineffective, we propose the use of a sophisticated computer-based information system to generate tailored, interactive handouts to communicate with patients. Our system uses text planning and user modeling techniques to generate natural language descriptions of migraine, its symptoms, triggering factors and prescriptions. The system is capable of handling follow-up questions requesting further information, and generating responses in the context of previously supplied information--a capability unavailable in previous patient information systems. The system tailors its interaction to: (i) the class of migraine patients, (ii) the individual patient, and (iii) the previous dialogue. Preliminary evaluation of the system indicates that patients find it useful and informative. More extensive evaluation is in progress.

Computer-Assisted Instruction↗

Incorporating constraint-based shape models into an interactive system for functional brain mapping.

Through intraoperative electrical stimulation mapping, it is possible to identify sites on the surface of the brain that are essential for language function. Interesting correlations have been found between the distribution of these sites and behavioral traits such as verbal IQ. In previous work, tools were developed for building a reconstruction of a patient's cortical surface and using it to recover coordinates of essential language sites. However, considerable expertise was required to produce good reconstructions. This paper describes an improved version of the mapping procedure, in which segmentation is driven by a 3-D shape model. The model-based approach provides more intuitive control over the system, allowing a trained user to complete a surface reconstruction and mapping in about two hours. This level of performance makes it feasible to gather language maps for a large number of patients, which hopefully will lead to significant new findings about language organization in the brain.

Anatomy, Cross-Sectional↗