Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

XFINGER: a tool for searching and visualising protein fingerprints and patterns.

A tool for searching pattern and fingerprint databases is described. Fingerprints are groups of motifs excised from conserved regions of sequence alignments and used for iterative database scanning. The constituent motifs are thus encoded as small alignments in which sequence information is maximised with each database pass; they therefore differ from regular-expression patterns, in which alignments are reduced to single consensus sequences. Different database formats have evolved to store these disparate types of information, namely the PROSITE dictionary of patterns and the PRINTS fingerprint database, but programs have not been available with the flexibility to search them both. We have developed a facility to do this: the system allows query sequences to be scanned against either PROSITE, the full PRINTS database, or against individual fingerprints. The results of fingerprint searches are displayed simultaneously in both text and graphical windows to render them more tangible to the user. Where structural coordinates are available, identified motifs may be visualised in a 3D context. The program runs on Silicon Graphics machines using GL graphics libraries and on machines with X servers supporting the PEX extension: its use is illustrated here by depicting the location of low-density lipoprotein-binding (LDL) motifs and leucine-rich repeats in a mosaic G-protein-coupled receptor (GPCR).

Amino Acid Sequence↗

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology↗

A new method to predict the consensus secondary structure of a set of unaligned RNA sequences.

MOTIVATION: To predict the consensus secondary structure, possibly including pseudoknots, of a set of RNA unaligned sequences. RESULTS: We have designed a method based on a new representation of any RNA secondary structure as a set of structural relationships between the helices of the structure. We refer to this representation as a structural pattern. In a first step, we use thermodynamic parameters to select, for each sequence, the best secondary structures according to energy minimization and we represent each of them using its corresponding structural pattern. In a second step, we search for the repeated structural patterns, i.e. the largest structural patterns that occur in at least one sequence, i.e. included in at least one of the structural patterns associated to each sequence. Thanks to an efficient encoding of structural patterns, this search comes down to identifying the largest repeated word suffixes in a dictionary. In a third step, we compute the plausibility of each repeated structural pattern by checking if it occurs more frequently in the studied sequences than in random RNA sequences. We then suppose that the consensus secondary structure corresponds to the repeated structural pattern that displays the highest plausibility. We present several experiments concerning tRNA, fragments of 16S rRNA and 10Sa RNA (including pseudoknots); in each of them, we found the putative consensus secondary structure.

Algorithms↗

Prediction of protein secondary structures by a neural network.

We have studied the prediction of globular protein secondary structures by neural networks. Protein secondary structures are allocated to amino acid residues using Kabsch and Sander's dictionary of protein secondary structures and the neural network is taught the protein secondary structures. The input layer of the neural network allows sequences of residues including 20 amino acids, chain break, B, X and Z. We consider classifying secondary structures into groups of 3, 4 and 8. In each case, we calculate the percentage of correct predictions. We discuss the effect of overlearning on the protein secondary structure prediction. In addition, we include the application of a neural network with a modular architecture to prediction of protein secondary structures. We compare the results from neural networks with a modular architecture and with a simple three-layer structure.

Algorithms↗

Accelerating screening of 3D protein data with a graph theoretical approach.

MOTIVATION: The Dictionary of Interfaces in Proteins (DIP) is a database collecting the 3D structure of interacting parts of proteins that are called patches. It serves as a repository, in which patches similar to given query patches can be found. The computation of the similarity of two patches is time consuming and traversing the entire DIP requires some hours. In this work we address the question of how the patches similar to a given query can be identified by scanning only a small part of DIP. The answer to this question requires the investigation of the distribution of the similarity of patches. RESULTS: The score values describing the similarity of two patches can roughly be divided into three ranges that correspond to different levels of spatial similarity. Interestingly, the two iso-score lines separating the three classes can be determined by two different approaches. Applying a concept of the theory of random graphs reveals significant structural properties of the data in DIP. These can be used to accelerate scanning the DIP for patches similar to a given query. Searches for very similar patches could be accelerated by a factor of more than 25. Patches with a medium similarity could be found 10 times faster than by brute-force search.

Algorithms↗

Discovering patterns to extract protein-protein interactions from full texts.

MOTIVATION: Although there are several databases storing protein-protein interactions, most such data still exist only in the scientific literature. They are scattered in scientific literature written in natural languages, defying data mining efforts. Much time and labor have to be spent on extracting protein pathways from literature. Our aim is to develop a robust and powerful methodology to mine protein-protein interactions from biomedical texts. RESULTS: We present a novel and robust approach for extracting protein-protein interactions from literature. Our method uses a dynamic programming algorithm to compute distinguishing patterns by aligning relevant sentences and key verbs that describe protein interactions. A matching algorithm is designed to extract the interactions between proteins. Equipped only with a dictionary of protein names, our system achieves a recall rate of 80.0% and precision rate of 80.5%. AVAILABILITY: The program is available on request from the authors.

Algorithms↗

Automatic extraction of gene/protein biological functions from biomedical text.

MOTIVATION: With the rapid advancement of biomedical science and the development of high-throughput analysis methods, the extraction of various types of information from biomedical text has become critical. Since automatic functional annotations of genes are quite useful for interpreting large amounts of high-throughput data efficiently, the demand for automatic extraction of information related to gene functions from text has been increasing. RESULTS: We have developed a method for automatically extracting the biological process functions of genes/protein/families based on Gene Ontology (GO) from text using a shallow parser and sentence structure analysis techniques. When the gene/protein/family names and their functions are described in ACTOR (doer of action) and OBJECT (receiver of action) relationships, the corresponding GO-IDs are assigned to the genes/proteins/families. The gene/protein/family names are recognized using the gene/protein/family name dictionaries developed by our group. To achieve wide recognition of the gene/protein/family functions, we semi-automatically gather functional terms based on GO using co-occurrence, collocation similarities and rule-based techniques. A preliminary experiment demonstrated that our method has an estimated recall of 54-64% with a precision of 91-94% for actually described functions in abstracts. When applied to the PUBMED, it extracted over 190 000 gene-GO relationships and 150 000 family-GO relationships for major eukaryotes.

Abstracting and Indexing↗

Applying GIFT, a Gene Interactions Finder in Text, to fly literature.

UNLABELLED: A number of freely available text mining tools have been put together to extract highly reliable Drosophila gene interaction data from text. The system has been tested with The Interactive Fly, showing low recall (27-34%), but very high precision (93-97%). AVAILABILITY: The extracted data and a web interface for submission of texts to GIFT analysis are available at http://gift.cryst.bbk.ac.uk/gift CONTACT: n.domedel_puig@cryst.bbk.ac.uk SUPPLEMENTARY INFORMATION: Additional documentation, such as the dictionaries and the reference sets, are available at the GIFT website.

Artificial Intelligence↗

The effects of very early Alzheimer's disease on the characteristics of writing by a renowned author.

Iris Murdoch (I.M.) was among the most celebrated British writers of the post-war era. Her final novel, however, received a less than enthusiastic critical response on its publication in 1995. Not long afterwards, I.M. began to show signs of insidious cognitive decline, and received a diagnosis of Alzheimer's disease, which was confirmed histologically after her death in 1999. Anecdotal evidence, as well as the natural history of the condition, would suggest that the changes of Alzheimer's disease were already established in I.M. while she was writing her final work. The end product was unlikely, however, to have been influenced by the compensatory use of dictionaries or thesauri, let alone by later editorial interference. These facts present a unique opportunity to examine the effects of the early stages of Alzheimer's disease on spontaneous written output from an individual with exceptional expertise in this area. Techniques of automated textual analysis were used to obtain detailed comparisons among three of her novels: her first published work, a work written during the prime of her creative life and the final novel. Whilst there were few disparities at the levels of overall structure and syntax, measures of lexical diversity and the lexical characteristics of these three texts varied markedly and in a consistent fashion. This unique set of findings is discussed in the context of the debate as to whether syntax and semantics decline separately or in parallel in patients with Alzheimer's disease.

Aged↗

Economic aspects of addiction policy.

One definition of policy or government action in the Oxford English Dictionary is "craftiness" i.e. cunning or deceit. Such qualities have to be employed by governments because of the potential vote-losing effects of radical addiction policies. Health promotion, in relation to addictive substances such as alcohol and tobacco in particular, involves a trade-off between the costs of such policies, especially to industry (which seeks regulation to protect itself from competitors), and the benefits--improvements in the quality and length of life. Measures of such benefits (quality-adjusted life-years or QALYs) are available now to use in the evaluation of competing health promotion policies to determine their efficiency at the margin. Analysis of the market for tobacco indicates that consumption has been falling generally in the UK except among teenagers who appear to be the target of the industry's advertising and sponsorship efforts. This fall in consumption appears to be explained by health promotion rather than the active use of fiscal instruments of control. The recognition of the health effects of passive smoking and the impact of advertising and sponsorship, especially on the young, are policy areas requiring careful review and the evaluation of the costs and benefits of competing policies.(ABSTRACT TRUNCATED AT 250 WORDS)

Alcoholism↗

Identifying achievable benchmarks of care: concepts and methodology.

Webster's Dictionary defines a benchmark as 'something that serves as a standard by which others can be measured'. Benchmarking pervades the health care quality improvement literature, and benchmarks are usually based on subjective assessment rather than on measurements derived from data. As such, benchmarks may fail to yield an achievable level of excellence that can be replicated under specific conditions. In this paper, we provide an overview of benchmarking in health care. We then describe the evolution of our data-driven method for identifying an Achievable Benchmark of Care (ABC) on the basis of process-of-care indicators. Here, our experience leads us to postulate the following premises for sound benchmarks: (i) benchmarks should represent a level of excellence; (ii) benchmarks should be demonstrably attainable; (iii) providers with high performance should be selected from among all providers in a predefined way using reliable data; (iv) all providers with high performance levels should contribute to the benchmark level; and (v) providers with high performance levels but small numbers of cases should not unduly influence the level of the benchmark. An example of an ABC applied to the cooperative cardiovascular project leads the reader through the computation of an ABC. Finally, we consider several refinements of the original ABC concept that are in progress, e.g. how to approach the special problems posed by very small denominators. The ABC methodology has been well accepted in multiple quality improvement projects. This approach lends objectivity and reliability to benchmarks that have been a widely used, but until now, arbitrarily defined tool.

Alabama↗

Assessing residents' prescribing behavior in renal impairment.

OBJECTIVE: Although fitting orders to renal function avoids overdosage and therefore iatrogenic risk, dosage adjustment is rarely made. The objective of this study was to assess residents' prescribing behavior in renal impairment, through a standardized simulated clinical setting. METHOD: This criterion-referenced study was carried out in a French teaching hospital. The hospital had 118 residents; 71 of them were asked to complete a questionnaire including four vignettes, simulating drug prescription in four 'patients' with various degrees of renal impairment (16 orders). The patients had an order of gentamicin sulfate, diclofenac sodium, and amlodipine bensylate. For each drug, the resident could maintain the order, discontinue the order, or change the dosage. A fourth drug, enalapril maleate, was to be started, with three possible dosages and the possibility of not prescribing it. The reference chosen for assessment was the Vidal dictionary, which corresponds to the Physician's Desk Reference and is the French reference for prescription. RESULTS: All the residents approached for the survey accepted the offer to complete the questionnaire. Among the 16 simulated orders, the median number of appropriate orders per resident was nine. Considering the renal function of their patients, 62% of residents wrote an inappropriate order for gentamicin, 42% wrote an inappropriate order for didofenac, and 52% wrote an inappropriate order for enalapril. Although no adjustment to renal function was required, 28% of the residents decreased the dosage of amlodipine and ordered an underdose. CONCLUSION: Considering the iatrogenic risk related to the lack of dosage adjustment, attention should be drawn to increasing residents' awareness of dosage adjustment in renal impairment and to providing them with better information on patients' renal function.

Aged↗

Signalling crosstalk in plants: emerging issues.

The Oxford English Dictionary defines crosstalk as 'unwanted transfer of signals between communication channels'. How does this definition relate to the way in which we view the organization and function of signalling pathways? Recent advances in the field of plant signalling have challenged the traditional view of a signalling transduction cascade as isolated linear pathways. Instead the picture emerging of the mechanisms by which plants transduce environmental signals is of the interaction between transduction chains. The manner in which these interactions occur (and indeed whether the transfer of these signals is 'unwanted' or beneficial) is currently the topic of intense research.

Plant Physiological Phenomena↗

The nucleotide sequence of a nematode vitellogenin gene.

The nematode, Caenorhabditis elegans, contains a family of six genes that code for vitellogenins. Here we report the complete nucleotide sequence of one of these genes, vit-5. The gene specifies a mRNA of 4869 nucleotides, including untranslated regions of 9 bases at the 5' end and 51 bases at the 3' end. Vit-5 contains four short introns totalling 218 bp. The predicted vitellogenin, yp170A, has a molecular weight of 186,430. At its N terminus it is clearly related to the vitellogenins of vertebrates. However, the vit-5-encoded protein does not contain a serine-rich sequence related to the vertebrate vitellin, phosvitin. In fact, the amino acid composition of the nematode protein is very similar to that of the vertebrate protein without phosvitin. Vit-5 has a highly asymmetric codon choice dictionary. The favored codons are different from those favored in other organisms, but are characteristic of highly expressed C. elegans genes. The strong selection against rare codons is not as great near the 5' end of the gene; rare codons are 15 times more frequent within the first 54 bp than in the next 4.8 kb.

Amino Acid Sequence↗

Progress with the PRINTS protein fingerprint database.

PRINTS is a compendium of protein motif 'fingerprints' derived from the OWL composite sequence database. Fingerprints are groups of motifs within sequence alignments whose conserved nature allows them to be used as signatures of family membership. To date, 400 fingerprints have been constructed and stored in Prints, the size of which has doubled in the last year. The current version, 9.0, encodes approximately 2000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. Fingerprints inherently offer improved diagnostic reliability over single motif methods by virtue of the mutual context provided by motif neighbours. PRINTS thus provides a useful adjunct to the widely used PROSITE dictionary of patterns. The database is now accessible via the Database Browser on the UCL Bioinformatics server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser .

Amino Acid Sequence↗

Novel developments with the PRINTS protein fingerprint database.

The PRINTS database of protein family 'fingerprints' is a diagnostic resource that complements the PROSITE dictionary of sites and patterns. Unlike regular expressions, fingerprints exploit groups of conserved motifs within sequence alignments to build characteristic signatures of family membership. Thus fingerprints inherently offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 600 fingerprints have been constructed and stored in PRINTS, representing a 50% increase in the size of the database in the last year. The current version, 13.0, encodes approximately 3000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via UCL's Bioinformatics World Wide Web (WWW) server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser / . We describe here progress with the database, its Web interface, and a recent exciting development: the integration of a novel colour alignment editor (http://www.biochem.ucl.ac.uk/bsm/dbbrowser++ +/CINEMA ), which allows visualisation and interactive manipulation of PRINTS alignments over the Internet.

Amino Acid Sequence↗

Touring protein fold space with Dali/FSSP.

The FSSP database and its new supplement, the Dali Domain Dictionary, present a continuously updated classification of all known 3D protein structures. The classification is derived using an automatic structure alignment program (Dali) for the all-against-all comparison of structures in the Protein Data Bank. From the resulting enumeration of structural neighbours (which form a surprisingly continuous distribution in fold space) we derive a discrete fold classification in three steps: (i) sequence-related families are covered by a representative set of protein chains; (ii) protein chains are decomposed into structural domains based on the recurrence of structural motifs; (iii) folds are defined as tight clusters of domains in fold space. The fold classification, domain definitions and test sets for sequence-structure alignment (threading) are accessible on the web at www.embl-ebi.ac.uk/dali . The web interface provides a rich network of links between neighbours in fold space, between domains and proteins, and between structures and sequences leading, for example, to a database of explicit multiple alignments of protein families in the twilight zone of sequence similarity. The Dali/FSSP organization of protein structures provides a map of the currently known regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination.

Computer Communication Networks↗

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗