Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Optimizing sparse and skew hashing: faster k-mer dictionaries.

MOTIVATION: Representing a set of k-mers-strings of length k-in small space under fast lookup queries is a fundamental requirement for several applications in Bioinformatics. A data structure based on sparse and skew hashing (SSHash) was recently proposed for this purpose (Pibiri 2022): it combines good space effectiveness with fast lookup and streaming queries. It is also order-preserving, i.e. consecutive k-mers (sharing a prefix-suffix overlap of length k-1) are assigned consecutive hash codes which helps compressing satellite data typically associated with k-mers, like abundances and color sets in colored De Bruijn graphs. RESULTS: We study the problem of accelerating queries under the sparse and skew hashing indexing paradigm, without compromising its space effectiveness. We propose a refined data structure with less complex lookups and fewer cache misses. We give a simpler and faster algorithm for streaming lookup queries. The refined architecture translates to substantial performance gains, outperforming the original version of SSHash in both index construction speed and query efficiency. Compared to indexes with similar capabilities and based on the Burrows-Wheeler transform, like SBWT and FMSI, SSHash is significantly faster to build and query. SSHash is competitive in space with the fast (and default) modality of SBWT when both k-mer strands are indexed. While larger than FMSI, it is also more than one order of magnitude faster to query. AVAILABILITY AND IMPLEMENTATION: The SSHash software is available at https://github.com/jermp/sshash, and also distributed via Bioconda. A benchmark of data structures for k-mer sets is available at https://github.com/jermp/kmer_sets_benchmark. The datasets used in this article are described and available at https://zenodo.org/records/17582116.

Algorithms↗

Cortical anatomy of mental imagery of concrete nouns based on their dictionary definition.

The functional anatomy of the interactions between spoken language and visual mental imagery was investigated with PET in eight normal volunteers during a series of three conditions: listening to concrete word definitions and generating their mental images (CONC), listening to abstract word definitions (ABST) and silent REST. The CONC task specifically elicited activations of the bilateral inferior temporal gyri, of the left premotor and left prefrontal regions, while activations in the bilateral superior temporal gyri were smaller than during the ABST task, during which an additional activation of the anterior part of the right middle temporal gyrus was observed. No activation of the occipital areas was observed during the CONC task when compared either to the REST or to the ABST task. The present study demonstrates that a network including part of the bilateral ventral stream and the frontal working memory areas is recruited when mental imagery of concrete words is performed on the basis of continuous spoken language.

Adult↗

Automated production of small-molecule dictionaries for use in crystallographic refinements.

Many macromolecules are now being studied crystallographically in complexes with a range of ligands and other associated molecules. It is necessary to have templates describing the expected geometry of such molecules before refinement and model building can be carried out. This paper describes a method for generating templates beginning from the SMILES description of the molecule, the final format of the molecular template being based on the mmCIF definitions for chemical composition. Additionally, the program SMILE2DICT, which converts the SMILES string to a more extended format, is described. The description details the input required, the output produced and how the program relates to attempts to automate the procedure of model building for crystallographic refinement. Examples of input to and output from the program are given.

Automation↗

A consensus view of fold space: combining SCOP, CATH, and the Dali Domain Dictionary.

We have determined consensus protein-fold classifications on the basis of three classification methods, SCOP, CATH, and Dali. These classifications make use of different methods of defining and categorizing protein folds that lead to different views of protein-fold space. Pairwise comparisons of domains on the basis of their fold classifications show that much of the disagreement between the classification systems is due to differing domain definitions rather than assigning the same domain to different folds. However, there are significant differences in the fold assignments between the three systems. These remaining differences can be explained primarily in terms of the breadth of the fold classifications. Many structures may be defined as having one fold in one system, whereas far fewer are defined as having the analogous fold in another system. By comparing these folds for a nonredundant set of proteins, the consensus method breaks up broad fold classifications and combines restrictive fold classifications into metafolds, creating, in effect, an averaged view of fold space. This averaged view requires that the structural similarities between proteins having the same metafold be recognized by multiple classification systems. Thus, the consensus map is useful for researchers looking for fold similarities that are relatively independent of the method used to compare proteins. The 30 most populated metafolds, representing the folds of about half of a nonredundant subset of the PDB, are presented here. The full list of metafolds is presented on the Web.

Amino Acid Motifs↗