Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Protein language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

GENPRO: automatic generation of Prolog clause files for knowledge-based systems in the biomedical sciences.

With the increasing interest in using knowledge-based approaches for protein structure prediction and modelling, there is a requirement for general techniques to convert molecular biological data into structures that can be interpreted by artificial intelligence programming languages (e.g. Prolog). We describe here an interactive program that generates files in Prolog clausal form from the most commonly distributed protein structural data collections. The program is flexible and enables a variety of clause structures to be defined by the user through a general schema definition system. Our method can be extended to include other types of molecular biological database or those containing non-structural information, thus providing a uniform framework for handling the increasing volume of data available to knowledge-based systems in biomedicine.

Database Management Systems↗

Learning anchor verbs for biological interaction patterns from published text articles.

Much of knowledge modeling in the molecular biology domain involves interactions between proteins, genes, various forms of RNA, small molecules, etc. Interactions between these substances are typically extracted and codified manually, increasing the cost and time for modeling and substantially limiting the coverage of the resulting knowledge base. In this paper, we describe an automatic system that learns from text interaction verbs; these verbs can then form the core of automatically retrieved patterns which model classes of biological interactions. We investigate text features relating verbs with genes and proteins, and apply statistical tests and a logistic regression statistical model to determine whether a given verb belongs to the class of interaction verbs. Our system, AVAD, achieves over 87% precision and 82% recall when tested on an 11 million word corpus of journal articles. In addition, we compare the automatically obtained results with a manually constructed database of interaction verbs and show that the automatic approach can significantly enrich the manual list by detecting rarer interaction verbs that were omitted from the database.

Artificial Intelligence↗

Linear programming based approach to the derivation of a contact potential for protein threading.

This paper proposes a novel method of deriving a contact potential (a pair score function) for protein threading. In this method, the constraint that the score of the native threading is minimum over all possible threadings is expressed in a form of linear inequalities, and then parameters defining the contact potential are determined by applying a program package of linear programming. The most important advantage of this method over the previous methods is that this method can learn a score function from a small number of training data. The proposed method was evaluated using Lathrop and Smith's algorithm for finding optimal threadings and was shown to be effective for computing nearly correct threadings.

Amino Acid Sequence↗

A query language for biological networks.

MOTIVATION: Many areas of modern biology are concerned with the management, storage, visualization, comparison and analysis of networks, but no appropriate query language for such complex data structures yet exists. RESULTS: We have designed and implemented the pathway query language (PQL) for querying large protein interaction or pathway databases. PQL is based on a simple graph data model with extensions reflecting properties of biological objects. Queries match subgraphs in the database based on node properties and paths between nodes. The syntax is easy to learn for anybody familiar with SQL. As an important feature, a query may require a certain structure in the database to exist but return a different subgraph. We have tested PQL queries on networks of up to 16,000 nodes and found it to scale very well. AVAILABILITY: The code is available on request from the author.

Computational Biology↗

Grammatical inference in bioinformatics.

Bioinformatics is an active research area aimed at developing intelligent systems for analyses of molecular biology. Many methods based on formal language theory, statistical theory, and learning theory have been developed for modeling and analyzing biological sequences such as DNA, RNA, and proteins. Especially, grammatical inference methods are expected to find some grammatical structures hidden in biological sequences. In this article, we give an overview of a series of our grammatical approaches to biological sequence analyses and related researches and focus on learning stochastic grammars from biological sequences and predicting their functions based on learned stochastic grammars.

Algorithms↗

Highly significant linkage to the SLI1 locus in an expanded sample of individuals affected by specific language impairment.

Specific language impairment (SLI) is defined as an unexplained failure to acquire normal language skills despite adequate intelligence and opportunity. We have reported elsewhere a full-genome scan in 98 nuclear families affected by this disorder, with the use of three quantitative traits of language ability (the expressive and receptive tests of the Clinical Evaluation of Language Fundamentals and a test of nonsense word repetition). This screen implicated two quantitative trait loci, one on chromosome 16q (SLI1) and a second on chromosome 19q (SLI2). However, a second independent genome screen performed by another group, with the use of parametric linkage analyses in extended pedigrees, found little evidence for the involvement of either of these regions in SLI. To investigate these loci further, we have collected a second sample, consisting of 86 families (367 individuals, 174 independent sib pairs), all with probands whose language skills are >/=1.5 SD below the mean for their age. Haseman-Elston linkage analysis resulted in a maximum LOD score (MLS) of 2.84 on chromosome 16 and an MLS of 2.31 on chromosome 19, both of which represent significant linkage at the 2% level. Amalgamation of the wave 2 sample with the cohort used for the genome screen generated a total of 184 families (840 individuals, 393 independent sib pairs). Analysis of linkage within this pooled group strengthened the evidence for linkage at SLI1 and yielded a highly significant LOD score (MLS = 7.46, interval empirical P<.0004). Furthermore, linkage at the same locus was also demonstrated to three reading-related measures (basic reading [MLS = 1.49], spelling [MLS = 2.67], and reading comprehension [MLS = 1.99] subtests of the Wechsler Objectives Reading Dimensions).

Adaptor Proteins, Signal Transducing↗

The application of molecular genetic approaches to the study of human evolution.

The past decade of advances in molecular genetic technology has heralded a new era for all evolutionary studies, but especially the science of human evolution. Data on various kinds of DNA variation in human populations have rapidly accumulated. There is increasing recognition of the importance of this variation for medicine and developmental biology and for understanding the history of our species. Haploid markers from mitochondrial DNA and the Y chromosome have proven invaluable for generating a standard model for evolution of modern humans. Conclusions from earlier research on protein polymorphisms have been generally supported by more sophisticated DNA analysis. Co-evolution of genes with language and some slowly evolving cultural traits, together with the genetic evolution of commensals and parasites that have accompanied modern humans in their expansion from Africa to the other continents, supports and supplements the standard model of genetic evolution. The advances in our understanding of the evolutionary history of humans attests to the advantages of multidisciplinary research.

Animals↗

BIOZON: a system for unification, management and analysis of heterogeneous biological data.

BACKGROUND: Integration of heterogeneous data types is a challenging problem, especially in biology, where the number of databases and data types increase rapidly. Amongst the problems that one has to face are integrity, consistency, redundancy, connectivity, expressiveness and updatability. DESCRIPTION: Here we present a system (Biozon) that addresses these problems, and offers biologists a new knowledge resource to navigate through and explore. Biozon unifies multiple biological databases consisting of a variety of data types (such as DNA sequences, proteins, interactions and cellular pathways). It is fundamentally different from previous efforts as it uses a single extensive and tightly connected graph schema wrapped with hierarchical ontology of documents and relations. Beyond warehousing existing data, Biozon computes and stores novel derived data, such as similarity relationships and functional predictions. The integration of similarity data allows propagation of knowledge through inference and fuzzy searches. Sophisticated methods of query that span multiple data types were implemented and first-of-a-kind biological ranking systems were explored and integrated. CONCLUSION: The Biozon system is an extensive knowledge resource of heterogeneous biological data. Currently, it holds more than 100 million biological documents and 6.5 billion relations between them. The database is accessible through an advanced web interface that supports complex queries, "fuzzy" searches, data materialization and more, online at http://biozon.org.

Animals↗

The PowerAtlas: a power and sample size atlas for microarray experimental design and research.

BACKGROUND: Microarrays permit biologists to simultaneously measure the mRNA abundance of thousands of genes. An important issue facing investigators planning microarray experiments is how to estimate the sample size required for good statistical power. What is the projected sample size or number of replicate chips needed to address the multiple hypotheses with acceptable accuracy? Statistical methods exist for calculating power based upon a single hypothesis, using estimates of the variability in data from pilot studies. There is, however, a need for methods to estimate power and/or required sample sizes in situations where multiple hypotheses are being tested, such as in microarray experiments. In addition, investigators frequently do not have pilot data to estimate the sample sizes required for microarray studies. RESULTS: To address this challenge, we have developed a Microrarray PowerAtlas. The atlas enables estimation of statistical power by allowing investigators to appropriately plan studies by building upon previous studies that have similar experimental characteristics. Currently, there are sample sizes and power estimates based on 632 experiments from Gene Expression Omnibus (GEO). The PowerAtlas also permits investigators to upload their own pilot data and derive power and sample size estimates from these data. This resource will be updated regularly with new datasets from GEO and other databases such as The Nottingham Arabidopsis Stock Center (NASC). CONCLUSION: This resource provides a valuable tool for investigators who are planning efficient microarray studies and estimating required sample sizes.

Algorithms↗

Identifying biological concepts from a protein-related corpus with a probabilistic topic model.

BACKGROUND: Biomedical literature, e.g., MEDLINE, contains a wealth of knowledge regarding functions of proteins. Major recurring biological concepts within such text corpora represent the domains of this body of knowledge. The goal of this research is to identify the major biological topics/concepts from a corpus of protein-related MEDLINE titles and abstracts by applying a probabilistic topic model. RESULTS: The latent Dirichlet allocation (LDA) model was applied to the corpus. Based on the Bayesian model selection, 300 major topics were extracted from the corpus. The majority of identified topics/concepts was found to be semantically coherent and most represented biological objects or concepts. The identified topics/concepts were further mapped to the controlled vocabulary of the Gene Ontology (GO) terms based on mutual information. CONCLUSION: The major and recurring biological concepts within a collection of MEDLINE documents can be extracted by the LDA model. The identified topics/concepts provide parsimonious and semantically-enriched representation of the texts in a semantic space with reduced dimensionality and can be used to index text.

Abstracting and Indexing↗

GAPSCORE: finding gene and protein names one word at a time.

MOTIVATION: New high-throughput technologies have accelerated the accumulation of knowledge about genes and proteins. However, much knowledge is still stored as written natural language text. Therefore, we have developed a new method, GAPSCORE, to identify gene and protein names in text. GAPSCORE scores words based on a statistical model of gene names that quantifies their appearance, morphology and context. RESULTS: We evaluated GAPSCORE against the Yapex data set and achieved an F-score of 82.5% (83.3% recall, 81.5% precision) for partial matches and 57.6% (58.5% recall, 56.7% precision) for exact matches. Since the method is statistical, users can choose score cutoffs that adjust the performance according to their needs. AVAILABILITY: GAPSCORE is available at http://bionlp.stanford.edu/gapscore/

Abstracting and Indexing↗

Vestige: maximum likelihood phylogenetic footprinting.

BACKGROUND: Phylogenetic footprinting is the identification of functional regions of DNA by their evolutionary conservation. This is achieved by comparing orthologous regions from multiple species and identifying the DNA regions that have diverged less than neutral DNA. Vestige is a phylogenetic footprinting package built on the PyEvolve toolkit that uses probabilistic molecular evolutionary modelling to represent aspects of sequence evolution, including the conventional divergence measure employed by other footprinting approaches. In addition to measuring the divergence, Vestige allows the expansion of the definition of a phylogenetic footprint to include variation in the distribution of any molecular evolutionary processes. This is achieved by displaying the distribution of model parameters that represent partitions of molecular evolutionary substitutions. Examination of the spatial incidence of these effects across regions of the genome can identify DNA segments that differ in the nature of the evolutionary process. RESULTS: Vestige was applied to a reference dataset of the SCL locus from four species and provided clear identification of the known conserved regions in this dataset. To demonstrate the flexibility to use diverse models of molecular evolution and dissect the nature of the evolutionary process Vestige was used to footprint the Ka/Ks ratio in primate BRCA1 with a codon model of evolution. Two regions of putative adaptive evolution were identified illustrating the ability of Vestige to represent the spatial distribution of distinct molecular evolutionary processes. CONCLUSION: Vestige provides a flexible, open platform for phylogenetic footprinting. Underpinned by the PyEvolve toolkit, Vestige provides a framework for visualising the signatures of evolutionary processes across the genome of numerous organisms simultaneously. By exploiting the maximum-likelihood statistical framework, the complex interplay between mutational processes, DNA repair and selection can be evaluated both spatially (along a sequence alignment) and temporally (for each branch of the tree) providing visual indicators to the attributes and functions of DNA sequences.

Algorithms↗

Plasma von Willebrand Factor and ADAMTS13 Interact With APOE-&#x3b5;4 in Predicting Longitudinal Brain Atrophy and Cognitive Decline Over a 9-Year Follow-Up.

BACKGROUND: Von Willebrand factor (VWF) and ADAMTS13 (a disintegrin and metalloproteinase with thrombospondin type 1 motif, 13) are linked to dementia risk, and limited evidence suggests apolipoprotein E (APOE)-&#x3b5;4 alters VWF release. This study assessed whether baseline VWF and ADAMTS13 levels predict neurodegeneration and cognitive decline and evaluated effect modification by APOE-&#x3b5;4 carriership. METHODS: Vanderbilt Memory and Aging Project cohort participants (n=332, 73&#xb1;7&#x2009;years, 59% male) completed serial blood draw, neuropsychological assessment, and brain magnetic resonance imaging over 6.4&#x2009;years (range 1.4-9.7&#x2009;years). Baseline plasma VWF and ADAMTS13 levels were quantified using mass spectrometry and Olink. Fully adjusted linear mixed-effects models related protein&#xd7;time and protein&#xd7;APOE-&#x3b5;4&#xd7;time interaction terms to longitudinal brain magnetic resonance imaging and neuropsychological outcomes. RESULTS: Lower baseline ADAMTS13 predicted faster declines in language (&#x3b2;=0.11, P=0.01), information processing speed (&#x3b2;=0.27, P=0.001), executive function (&#x3b2;=0.01, P=0.03), episodic memory (&#x3b2;=0.01, P=0.03), and visuospatial ability (&#x3b2;=0.11, P=0.001) and faster increases in global (&#x3b2;=-0.29, P=0.01) and frontal (&#x3b2;=-0.17, P=0.01) white matter hyperintensity volumes. Associations between ADAMTS13 and faster rates of cognitive decline and white matter injury were driven by APOE-&#x3b5;4 carriers. Models relating VWF to longitudinal outcomes were null. APOE-&#x3b5;4 interacted with VWF on longitudinal gray matter volumetric outcomes, such that faster rates of global gray matter atrophy were observed with higher baseline VWF levels among APOE-&#x3b5;4 noncarriers only (&#x3b2;=-1530.5, P<0.001). CONCLUSIONS: ADAMTS13 shows promise as a potential plasma biomarker for brain aging outcomes, but additional research is warranted to understand the performance of VWF in the presence versus absence of an APOE-&#x3b5;4 allele.

Humans↗

Post-translational modification of proteins in the human testis development pathway.

BACKGROUND: The foetal testes produce the androgens necessary to masculinise the developing embryo and support the maturation of germ cells, that will eventually develop into sperm, thus ensuring future reproductive capacity. The testes develop from the bi-potential gonads in a highly orchestrated process resulting in the differentiation of a complex tissue with multiple cellular lineages. While recent transcriptomic and chromatin-based analyses of human foetal testes have provided an unprecedented level of insight into signalling pathways activated during this process, proteomic studies of the human foetal gonads remain limited. Proteins are active molecules and post-translational modification (PTM) of proteins influences protein activity, stability and localisation. Studies have shown that PTMs regulate critical proteins in testis development, and their disruptions are implicated in congenital disorders including differences of sex development (DSD), in which sex development is atypical. Despite this, the role and regulation of protein PTM during human testis development remains poorly understood due to limited access to human foetal gonadal tissue, a paucity of large-scale proteomics studies, and a lack of robust of human gonad in vitro models. OBJECTIVE AND RATIONALE: This review aims to provide a comprehensive analysis of validated PTMs affecting proteins critical for testicular development. We discuss PTMs with evidence for a role in normal testis development, and highlight those disrupted in DSD. We review emerging techniques, including proteomic technologies and organ modelling systems that may advance our understanding of PTMs in foetal testis development. We discuss challenges that have restricted the application of these technologies and how overcoming these will significantly improve our understanding of testis development and disease, diagnostics and patient outcomes. SEARCH METHODS: We searched PubMed and the University of Melbourne library for peer-reviewed English-language studies using keywords such as phosphorylation, SUMOylation, acetylation, ubiquitination alongside each protein of interest. PTM sites in proteins involved in testis development were identified using the PhosphoSitePlus database focusing those confirmed in in vitro or animal model studies. ClinVar and the Human Gene Mutation Database were used to identify patient variants that may disrupt PTM sites. OUTCOMES: Our review finds that proteins required for human foetal testis development are subject to extensive PTM. Several PTM sites and PTM-mediated pathways [e.g. MAPK (mitogen-activated protein kinase) pathway] are disrupted in patients with DSD or related conditions. While recent advances in proteomics technologies hold considerable promise, their application to human foetal gonads has been constrained by technical, ethical, and logistical challenges. Encouragingly, emerging high-sensitivity and low-input technologies, alongside stem cell-based approaches, offer viable pathways to overcoming these barriers. WIDER IMPLICATIONS: The relationship between gene regulation, protein expression, and cellular outcome is inherently non-linear, shaped by additional regulatory layers-most notably PTMs. The contribution of PTMs to human testis development in both typical and atypical contexts is a major knowledge gap. Addressing this gap has broad clinical and biological relevance: it may help improve genetic diagnosis or shed light on how proteins or pathways critical for testis development respond to environmental signals-an increasingly pressing question as declining global fertility rates bring testicular function under greater scrutiny. REGISTRATION NUMBER: N/A.

Humans↗

Biomolecular visualization using AVS.

Dataflow systems for scientific visualization are becoming increasingly sophisticated in their architecture and functionality. AVS, from Advanced Visual Systems Inc., is a powerful dataflow environment that has been applied to many computation and visualization tasks. An important, yet complex, application area is molecular modeling and biomolecular visualization. Problems in biomolecular visualization tax the capability of dataflow systems because of the diversity of operations that are required and because many operations do not fit neatly into the dataflow paradigm. Here we describe visualization strategies and auxiliary programs developed to enhance the applicability of AVS for molecular modelling. Our visualization strategy is to use general-purpose AVS modules and a small number of chemistry-specific modules. We have developed methods to control AVS using AVS-tool, a programmable interface to the AVS Command Line Interpreter (CLI), and have also developed NAB, a C-like language for writing AVS modules that has extensions for operating on proteins and nucleic acids. This strategy provides a flexible and extensible framework for a wide variety of molecular modeling tasks.

Artificial Intelligence↗

Importing statistical measures into Artemis enhances gene identification in the Leishmania genome project.

BACKGROUND: Seattle Biomedical Research Institute (SBRI) as part of the Leishmania Genome Network (LGN) is sequencing chromosomes of the trypanosomatid protozoan species Leishmania major. At SBRI, chromosomal sequence is annotated using a combination of trained and untrained non-consensus gene-prediction algorithms with ARTEMIS, an annotation platform with rich and user-friendly interfaces. RESULTS: Here we describe a methodology used to import results from three different protein-coding gene-prediction algorithms (GLIMMER, TESTCODE and GENESCAN) into the ARTEMIS sequence viewer and annotation tool. Comparison of these methods, along with the CODONUSAGE algorithm built into ARTEMIS, shows the importance of combining methods to more accurately annotate the L. major genomic sequence. CONCLUSION: An improvised and powerful tool for gene prediction has been developed by importing data from widely-used algorithms into an existing annotation platform. This approach is especially fruitful in the Leishmania genome project where there is large proportion of novel genes requiring manual annotation.

Algorithms↗

Visualisation and integration of G protein-coupled receptor related information help the modelling: description and applications of the Viseur program.

G Protein-Coupled Receptors (GPCRs) constitute a superfamily of receptors that forms an important therapeutic target. The number of known GPCR sequences and related information increases rapidly. For these reasons, we are developing the Viseur program to integrate the available information related to GPCRs. The Viseur program allows one to interactively visualise and/or modify the sequences, transmembrane areas, alignments, models and results of mutagenesis experiments in an integrated environment. This integration increases the ease of modelling GPCRs: visualisation and manipulation improvements enable easier databank interrogation and interpretation. Unique program features include: (i) automatic construction of 'Snake-like' diagrams or hyperlinked GPCR molecular models to HTML or VRML and (ii) automatic access to a mutagenesis data server through the Internet. The novel algorithms or methods involved are presented, followed by the overall complementary features of the program. Finally, we present two applications of the program: (i) an automatic construction of GPCR snake-like diagrams for the GPCRDB WWW server, and (ii) a preparation of the modelling of the 5HT receptor subtypes. The interest of the direct access to mutagenesis results through an alignment and a molecular model are discussed. The Viseur program, which runs on SGI workstations, is freely available and can be used for preparing the modelling of integral membrane proteins or as an alignment editor tool.

Algorithms↗