Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Large language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Generating patient-specific interactive natural language explanations.

Patient compliance is a significant problem and is strongly correlated with the patients' understanding of their condition and prescribed treatment. Since doctors typically do not have large amounts of time to educate patients, and impersonal, voluminous patient handouts are largely ineffective, we propose the use of a sophisticated computer-based information system to generate tailored, interactive handouts to communicate with patients. Our system uses text planning and user modeling techniques to generate natural language descriptions of migraine, its symptoms, triggering factors and prescriptions. The system is capable of handling follow-up questions requesting further information, and generating responses in the context of previously supplied information--a capability unavailable in previous patient information systems. The system tailors its interaction to: (i) the class of migraine patients, (ii) the individual patient, and (iii) the previous dialogue. Preliminary evaluation of the system indicates that patients find it useful and informative. More extensive evaluation is in progress.

Computer-Assisted Instruction↗

Incorporating constraint-based shape models into an interactive system for functional brain mapping.

Through intraoperative electrical stimulation mapping, it is possible to identify sites on the surface of the brain that are essential for language function. Interesting correlations have been found between the distribution of these sites and behavioral traits such as verbal IQ. In previous work, tools were developed for building a reconstruction of a patient's cortical surface and using it to recover coordinates of essential language sites. However, considerable expertise was required to produce good reconstructions. This paper describes an improved version of the mapping procedure, in which segmentation is driven by a 3-D shape model. The model-based approach provides more intuitive control over the system, allowing a trained user to complete a surface reconstruction and mapping in about two hours. This level of performance makes it feasible to gather language maps for a large number of patients, which hopefully will lead to significant new findings about language organization in the brain.

Anatomy, Cross-Sectional↗

Modeling the UMLS using an OODB.

The Unified Medical Language System combines many well established authoritative medical informatics terminologies in one system. Such a resource is very valuable to the healthcare industry. However, the UMLS is very large and complex and poses serious comprehension problems for users and maintenance personnel. Furthermore, the sets of concepts of semantic types are not semantically uniform and thus are difficult to study. We describe a method to represent two components of the UMLS, the Metathesaurus (META) and the Semantic Network, as an OODB. The resulting UMLS OODB schema is deeper and more refined than the Semantic Network. It offers semantically uniform classes, which improves support for comprehension and navigation of META. The UMLS OODB also exposes problems in the semantic type classifications.

Classification↗

Indo-European origins: a computer-simulation test of five hypotheses.

Allele frequency distributions were generated by computer simulation of five models of microevolution in European populations. Genetic distances calculated from these distributions were compared with observed genetic distances among Indo-European speakers. The simulated models differ in complexity, but all incorporate random genetic drift and short-range gene flow (isolation by distance). The best correlations between observed and simulated data were obtained for two models where dispersal of Neolithic farmers from the Near East depends only on population growth. More complex models, where the timing of the farmers' expansion is constrained by archaeological time data, fail to account for a larger fraction of the observed genetic variation; this is also the case for a model including late Neolithic migrations from the Pontic steppes. The genetic structure of current populations speaking Indo-European languages seems therefore to largely reflect a Neolithic expansion. This is consistent with the hypothesis of a parallel spread of farming technologies and a proto-Indo-European language in the Neolithic. Allele-frequency gradients among Indo-European speakers may be due either to incomplete admixture between dispersing farmers, who presumably spoke proto-Indo-European, and pre-existing hunters and gatherers (as in the traditional demic diffusion hypothesis), or to founder effects during the farmers' dispersal. By contrast, successive migrational waves from the East, if any, do not seem to have had genetic consequences detectable by the present comparison of observed and simulated allele frequencies.

Alleles↗

Design of a data model for developing laboratory information management and analysis systems for protein production.

Data management has emerged as one of the central issues in the high-throughput processes of taking a protein target sequence through to a protein sample. To simplify this task, and following extensive consultation with the international structural genomics community, we describe here a model of the data related to protein production. The model is suitable for both large and small facilities for use in tracking samples, experiments, and results through the many procedures involved. The model is described in Unified Modeling Language (UML). In addition, we present relational database schemas derived from the UML. These relational schemas are already in use in a number of data management projects.

Algorithms↗

Creation of clinical research databases in the 21st century: a practical algorithm for HIPAA Compliance.

BACKGROUND: Enforcement of the Health Insurance Portability and Accountability Act (HIPAA) began in April, 2003. Designed as a law mandating health insurance availability when coverage was lost, HIPAA imposed sweeping and broad-reaching protections of patient privacy. These changes dramatically altered clinical research by placing sizeable regulatory burdens upon investigators with threat of severe and costly federal and civil penalties. This report describes development of an algorithmic approach to clinical research database design based upon a central key-shared data (CK-SD) model allowing researchers to easily analyze, distribute, and publish clinical research without disclosure of HIPAA Protected Health Information (PHI). METHODS: Three clinical database formats (small clinical trial, operating room performance, and genetic microchip array datasets) were modeled using standard structured query language (SQL)-compliant databases. The CK database was created to contain PHI data, whereas a shareable SD database was generated in real-time containing relevant clinical outcome information while protecting PHI items. Small (< 100 records), medium (< 50,000 records), and large (> 10(8) records) model databases were created, and the resultant data models were evaluated in consultation with an HIPAA compliance officer. RESULTS: The SD database models complied fully with HIPAA regulations, and resulting "shared" data could be distributed freely. Unique patient identifiers were not required for treatment or outcome analysis. Age data were resolved to single-integer years, grouping patients aged > 89 years. Admission, discharge, treatment, and follow-up dates were replaced with enrollment year, and follow-up/outcome intervals calculated eliminating original data. Two additional data fields identified as PHI (treating physician and facility) were replaced with integer values, and the original data corresponding to these values were stored in the CK database. Use of the algorithm at the time of database design did not increase cost or design effort. CONCLUSIONS: The CK-SD model for clinical database design provides an algorithm for investigators to create, maintain, and share clinical research data compliant with HIPAA regulations. This model is applicable to new projects and large institutional datasets, and should decrease regulatory efforts required for conduct of clinical research. Application of the design algorithm early in the clinical research enterprise does not increase cost or the effort of data collection.

Algorithms↗

Bidirectional incremental parsing for automatic pathway identification with combinatory categorial grammar.

As the importance of automatically extracting and analyzing various natural language assertions about protein-protein interactions in biomedical publications is recognized, many uses of natural language processing techniques are proposed in the literature. However, most proposals to date make rather simplifying assumptions about the syntactic aspects of natural language due to various reasons including efficiency. In this paper, we describe an implemented system that utilizes combinatory categorical grammar known to be competent in modeling natural language, with a controlled mechanism for the parser to operate bidirectionally and incrementally. We discuss the performance of the system on a large set of abstracts in Medline with quite encouraging results.

Data Interpretation, Statistical↗

Representing the UMLS as an object-oriented database: modeling issues and advantages.

OBJECTIVE: The Unified Medical Language System (UMLS) combines many well-established authoritative medical informatics terminologies in one knowledge representation system. Such a resource is very valuable to the health care community and industry. However, the UMLS is very large and complex and poses serious comprehension problems for users and maintenance personnel. The authors present a representation to support the user's comprehension and navigation of the UMLS. DESIGN: An object-oriented database (OODB) representation is used to represent the two major components of the UMLS-the Metathesaurus and the Semantic Network-as a unified system. The semantic types of the Semantic Network are modeled as semantic type classes. Intersection classes are defined to model concepts of multiple semantic types, which are removed from the semantic type classes. RESULTS: The authors provide examples of how the intersection classes help expose omissions of concepts, highlight errors of semantic type classification, and uncover ambiguities of concepts in the UMLS. The resulting UMLS OODB schema is deeper and more refined than the Semantic Network, since intersection classes are introduced. The Metathesaurus is classified into more mutually exclusive, uniform sets of concepts. The schema improves the user's comprehension and navigation of the Metathesaurus. CONCLUSIONS: The UMLS OODB schema supports the user's comprehension and navigation of the Metathesaurus. It also helps expose and resolve modeling problems in the UMLS.

Databases as Topic↗

Topology-induced coarsening in language games.

We investigate how very large populations are able to reach a global consensus, out of local "microscopic" interaction rules, in the framework of a recently introduced class of models of semiotic dynamics, the so-called naming game. We compare in particular the convergence mechanism for interacting agents embedded in a low-dimensional lattice with respect to the mean-field case. We highlight that in low dimensions consensus is reached through a coarsening process that requires less cognitive effort of the agents, with respect to the mean-field case, but takes longer to complete. In one dimension, the dynamics of the boundaries is mapped onto a truncated Markov process from which we analytically computed the diffusion coefficient. More generally we show that the convergence process requires a memory per agent scaling as N and lasts a time N1+2/d in dimension d < or = 4 (the upper critical dimension), while in mean field both memory and time scale as N3/2 , for a population of agents. We present analytical and numerical evidence supporting this picture.

Journal Article↗

ScrumPy: metabolic modelling with Python.

ScrumPy is a software package used for the definition and analysis of metabolic models. It is written using the Python programming language that is also used as a user interface. ScrumPy has features for both kinetic and structural modelling, but the emphasis is on structural modelling and those features of most relevance to analysis of large (genome-scale) models. The aim is at describing ScrumPy's functionality to readers with some knowledge of metabolic modelling, but implementation, programming and other computational details are omitted. ScrumPy is released under the Gnu Public Licence, and available for download from http://mudshark.brookes.ac.uk/ ScrumPy.

Adaptation, Physiological↗

Concern about environmental pollution: how much difference do race and ethnicity make? A New Jersey case study.

A survey conducted among 1,513 residents of New Jersey during March-May 2004 showed that non-Hispanic black, non-Hispanic white, and English-speaking Hispanic Americans were significantly more concerned about environmental pollution problems than were Asian Americans and Spanish-language Hispanic Americans. For example, an average of > 40% of the first three groups was very concerned about New Jersey's environmental problems, compared with 15% of the last two populations. There were also racial/ethnic differences among these groups in their desire for government action to protect the environment and in their personal support of the environmental movement. Regression analyses suggest that the 1970s and 1980s model of core support for environmental protection from white, female, young, educated, and politically liberal people has largely, but not completely, continued among non-Hispanic white, non-Hispanic black, and English-language Hispanic populations. But these demographic pointers do not hold for Asian and Spanish-language Hispanic Americans, except indicating more support among the more formally educated. The last two groups are the two fastest-growing subpopulations in the United States, and although acculturation may slowly increase their concern about environmental pollution, it is more prudent for proponents of environmental protection not to wait and instead to try to better understand the environmental perceptions of these groups.

Adult↗

The factor structure of Greek personality adjectives.

Personality descriptors--3,302 adjectives--were extracted from a dictionary of the modern Greek language. Those terms with the highest frequency were administered to large samples in Greece to test the universality of the Big-Five dimensions of personality in comparison to alternative models. One- and 2-factor structures were the most stable across variable selections and subsamples and replicated such structures found in previous studies. Among models with more moderate levels of replication, recently proposed 6- and 7-lexical-factor models were approximately as well replicated as the Big Five. An emic 6-factor structure showed relative stability; these factors were labeled Negative-Valence/Honesty, Agreeableness/Positive Affect, Prowess/Heroism, Introversion/Melancholia, Even Temper, and Conscientiousness.

Culture↗

A genome-wide search strategy for identifying quantitative trait loci involved in reading and spelling disability (developmental dyslexia).

Family and twin studies of developmental dyslexia have consistently shown that there is a significant heritable component for this disorder. However, any genetic basis for the trait is likely to be complex, involving reduced penetrance, phenocopy, heterogeneity and oligogenic inheritance. This complexity results in reduced power for traditional parametric linkage analysis, where specification of the correct genetic model is important. One strategy is to focus on large multigenerational pedigrees with severe phenotypes and/or apparent simple Mendelian inheritance, as has been successfully demonstrated for speech and language impairment. This approach is limited by the scarcity of such families. An alternative which has recently become feasible due to the development of high-throughput genotyping techniques is the analysis of large numbers of sib-pairs using allele-sharing methodology. This paper outlines our strategy for conducting a systematic genome-wide search for genes involved in dyslexia in a large number of affected sib-pair familites from the UK. We use a series of psychometric tests to obtain different quantitative measures of reading deficit, which should correlate with different components of the dyslexia phenotype, such as phonological awareness and orthographic coding ability. This enable us to use QTL (quantitative trait locus) mapping as a powerful tool for localising genes which may contribute to reading and spelling disability.

Alleles↗

Content-based indexing of images and video.

By representing image content using probabilistic models of an object's appearance we can obtain semantics-preserving compression of the image data. Such compact representations of an image's salient features allow rapid computer searches of even large image databases. Examples are shown for databases of face images, a video of American sign language (ASL), and a video of facial expressions.

Algorithms↗

Linking verbal and non-verbal representations: computer analysis of referential activity.

The objective of this study was to develop a computer assisted procedure to model the Referencial Activity scales as scored by raters. Referential Activity is defined as the function of connecting non-verbal experience with language. Using a large text corpus that had been rated by experienced and reliable judges, extreme samples from both ends of the Referential Activity Scales were selected. The Characteristic Vocabularies for each of these corpora, words that were significantly more frequent in each corpus as compared to the other, were then identified. A small set of 181 frequent words was derived that accounted for half of all words in the text corpora. These words were used as dictionaries for a Computerized Referential Activity measure based on computer assisted content analysis techniques. The new measure showed a correlation with judge-scored Referential Activity of around .50 across both the development and test corpora.

Artificial Intelligence↗

Academic emotions from a social-cognitive perspective: antecedents and domain specificity of students' affect in the context of Latin instruction.

This study concentrates on two assumptions of a social-cognitive model outlining the development of academic emotions (emotions directly linked to learning, classroom instruction, and achievement), namely on their antecedents and domain-specific organization. Our sample consisted of 200 students from Grades 7 to 10. Proposed relationships concerning the antecedents of academic emotions were tested in the context of Latin language instruction. Correlational analyses substantiated our assumptions concerning the relationships between academic emotions, students' cognitions, and aspects of the social environment. The mediating mechanisms proposed in the model were also confirmed using linear structural equation modelling. Subjective control- and value-related cognitions were found to mediate the relationship between aspects of the social environment and students' emotional experience. Our results further suggest that academic emotions are largely organized along domain-specific lines, with the degree of domain specificity varying according to the emotion in question. Implications for research and practice are discussed.

Achievement↗

Sequentiality of speech acts in conversational structure.

One aspect of the phenomenon of coherence in conversational discourse was addressed in the present study: sequentiality of speech acts. Several models of discourse structure have postulated sequencing rules between speech acts in conversations, but these efforts have been hampered by the lack of an efficient empirical method that can characterize a large body of language data. The lag sequential technique is proposed here as a tool that can be used to abstract a "grammar" of speech act contingency from spoken discourse. Derived patterns of discourse between female adults and preschool children confirmed expectations that most discourse is based upon three fundamental speech act pairings: question--answer, statement--reply, and directive--acknowledgement. It was also found that interlocutor differences in status, knowledge, and conversational ability affected the structure of the discourse in predictable ways.

Child↗

Age, cohort, and time development muddles: easy in practice, hard in theory.

The debate among developmental psychologists over how best to combine longitudinal and cross-sectional data sequences can be traced back at least four decades. During the 1970s, a variation of this theme received much attention: could the developmental influences of age, cohort, and time be unraveled by sufficiently ingenious application of combined data sequences? We believe discussion of this question has been needlessly parochial and confused. Substantive and methodological contributions from other disciplines have, until recently, been largely ignored by developmental psychologists. Moreover, the solutions debated by psychologists have generally been formulated in language that obscured, rather than explicated, the formal indeterminacy implicit in models of age, cohort, and time parameters. Methodologists have outlined formal solutions to problems of indeterminacy in the contexts of model identification and estimability theory. Sociologists have proposed solutions, based on explicitly theoretical assumptions, that permit model identification and unambiguous interpretation. We review these contributions, and suggest a hierarchy of solutions to the problem of age, cohort, and time indeterminacy.

Aging↗