Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dictionary”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

A rapid classification protocol for the CATH Domain Database to support structural genomics.

In order to support the structural genomic initiatives, both by rapidly classifying newly determined structures and by suggesting suitable targets for structure determination, we have recently developed several new protocols for classifying structures in the CATH domain database (http://www.biochem.ucl.ac.uk/bsm/cath). These aim to increase the speed of classification of new structures using fast algorithms for structure comparison (GRATH) and to improve the sensitivity in recognising distant structural relatives by incorporating sequence information from relatives in the genomes (DomainFinder). In order to ensure the integrity of the database given the expected increase in data, the CATH Protein Family Database (CATH-PFDB), which currently includes 25,320 structural domains and a further 160,000 sequence relatives has now been installed in a relational ORACLE database. This was essential for developing more rigorous validation procedures and for allowing efficient querying of the database, particularly for genome analysis. The associated Dictionary of Homologous Superfamilies [Bray,J.E., Todd,A.E., Pearl,F.M.G., Thornton,J.M. and Orengo,C.A. (2000) Protein Eng., 13, 153-165], which provides multiple structural alignments and functional information to assist in assigning new relatives, has also been expanded recently and now includes information for 903 homologous superfamilies. In order to improve coverage of known structures, preliminary classification levels are now provided for new structures at interim stages in the classification protocol. Since a large proportion of new structures can be rapidly classified using profile-based sequence analysis [e.g. PSI-BLAST: Altschul,S.F., Madden,T.L., Schaffer,A.A., Zhang,J., Zhang,Z., Miller,W. and Lipman,D.J. (1997) Nucleic Acids Res., 25, 3389-3402], this provides preliminary classification for easily recognisable homologues, which in the latest release of CATH (version 1.7) represented nearly three-quarters of the non-identical structures.

Computational Biology↗

The Zebrafish Information Network (ZFIN): a resource for genetic, genomic and developmental research.

The Zebrafish Information Network, ZFIN, is a WWW community resource of zebrafish genetic, genomic and developmental research information (http://zfin.org). ZFIN provides an anatomical atlas and dictionary, developmental staging criteria, research methods, pathology information and a link to the ZFIN relational database (http://zfin. org/ZFIN/). The database, built on a relational, object-oriented model, provides integrated information about mutants, genes, genetic markers, mapping panels, publications and contact information for the zebrafish research community. The database is populated with curated published data, user submitted data and large dataset uploads. A broad range of data types including text, images, graphical representations and genetic maps supports the data. ZFIN incorporates links to other genomic resources that provide sequence and ortholog data. Zebrafish nomenclature guidelines and an automated registration mechanism for new names are provided. Extensive usability testing has resulted in an easy to learn and use forms interface with complex searching capabilities.

Animals↗

MEPD: a Medaka gene expression pattern database.

The Medaka Expression Pattern Database (MEPD) stores and integrates information of gene expression during embryonic development of the small freshwater fish Medaka (Oryzias latipes). Expression patterns of genes identified by ESTs are documented by images and by descriptions through parameters such as staining intensity, category and comments and through a comprehensive, hierarchically organized dictionary of anatomical terms. Sequences of the ESTs are available and searchable through BLAST. ESTs in the database are clustered upon entry and have been blasted against public data-bases. The BLAST results are updated regularly, stored within the database and searchable. The MEPD is a project within the Medaka Genome Initiative (MGI) and entries will be interconnected to integrated genomic map databases. MEPD is accessible through the WWW at http://medaka.dsp.jst.go.jp/MEPD.

Animals↗

The Zebrafish Information Network (ZFIN): the zebrafish model organism database.

The Zebrafish Information Network (ZFIN) is a web based community resource that serves as a centralized location for the curation and integration of zebrafish genetic, genomic and developmental data. ZFIN is publicly accessible at http://zfin.org. ZFIN provides an integrated representation of mutants, genes, genetic markers, mapping panels, publications and community contact data. Recent enhancements to ZFIN include: (i) an anatomical dictionary that provides a controlled vocabulary of anatomical terms, grouped by developmental stages, that may be used to annotate and query gene expression data; (ii) gene expression data; (iii) expanded support for genome sequence; (iv) gene annotation using the standardized vocabulary of Gene Ontology (GO) terms that can be used to elucidate relationships between gene products in zebrafish and other organisms; and (v) collaborations with other databases (NCBI, Sanger Institute and SWISS-PROT) to provide standardization and interconnections based on shared curation.

Animals↗

Gene3D: structural assignments for the biologist and bioinformaticist alike.

The Gene3D database (http://www.biochem.ucl.ac.uk/bsm/cath_new/Gene3D/) provides structural assignments for genes within complete genomes. These are available via the internet from either the World Wide Web or FTP. Assignments are made using PSI-BLAST and subsequently processed using the DRange protocol. The DRange protocol is an empirically benchmarked method for assessing the validity of structural assignments made using sequence searching methods where appropriate assignment statistics are collected and made available. Gene3D links assignments to their appropriate entries in relevent structural and classification resources (PDBsum, CATH database and the Dictionary of Homologous Superfamilies). Release 2.0 of Gene3D includes 62 genomes, 2 eukaryotes, 10 archaea and 40 bacteria. Currently, structural assignments can be made for between 30 and 40 percent of any given genome. In any genome, around half of those genes assigned a structural domain are assigned a single domain and the other half of the genes are assigned multiple structural domains. Gene3D is linked to the CATH database and is updated with each new update of CATH.

Animals↗

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗

Computational identification of protein coding potential of conserved sequence tags through cross-species evolutionary analysis.

The identification of conserved sequence tags (CSTs) through comparative genome analysis may reveal important regulatory elements involved in shaping the spatio-temporal expression of genetic information. It is well known that the most significant fraction of CSTs observed in human-mouse comparisons correspond to protein coding exons, due to their strong evolutionary constraints. As we still do not know the complete gene inventory of the human and mouse genomes it is of the utmost importance to establish if detected conserved sequences are genes or not. We propose here a simple algorithm that, based on the observation of the specific evolutionary dynamics of coding sequences, efficiently discriminates between coding and non-coding CSTs. The application of this method may help the validation of predicted genes, the prediction of alternative splicing patterns in known and unknown genes and the definition of a dictionary of non-coding regulatory elements.

Algorithms↗

NLProt: extracting protein names and sequences from papers.

Automatically extracting protein names from the literature and linking these names to the associated entries in sequence databases is becoming increasingly important for annotating biological databases. NLProt is a novel system that combines dictionary- and rule-based filtering with several support vector machines (SVMs) to tag protein names in PubMed abstracts. When considering partially tagged names as errors, NLProt still reached a precision of 75% at a recall of 76%. By many criteria our system outperformed other tagging methods significantly; in particular, it proved very reliable even for novel names. Names encountered particularly frequently in Drosophila, such as white, wing and bizarre, constitute an obvious limitation of NLProt. Our method is available both as an Internet server and as a program for download (http://cubic.bioc.columbia.edu/services/NLProt/). Input can be PubMed/MEDLINE identifiers, authors, titles and journals, as well as collections of abstracts, or entire papers.

Algorithms↗

The CATH Domain Structure Database and related resources Gene3D and DHS provide comprehensive domain family information for genome analysis.

The CATH database of protein domain structures (http://www.biochem.ucl.ac.uk/bsm/cath/) currently contains 43,229 domains classified into 1467 superfamilies and 5107 sequence families. Each structural family is expanded with sequence relatives from GenBank and completed genomes, using a variety of efficient sequence search protocols and reliable thresholds. This extended CATH protein family database contains 616,470 domain sequences classified into 23,876 sequence families. This results in the significant expansion of the CATH HMM model library to include models built from the CATH sequence relatives, giving a 10% increase in coverage for detecting remote homologues. An improved Dictionary of Homologous superfamilies (DHS) (http://www.biochem.ucl.ac.uk/bsm/dhs/) containing specific sequence, structural and functional information for each superfamily in CATH considerably assists manual validation of homologues. Information on sequence relatives in CATH superfamilies, GenBank and completed genomes is presented in the CATH associated DHS and Gene3D resources. Domain partnership information can be obtained from Gene3D (http://www.biochem.ucl.ac.uk/bsm/cath/Gene3D/). A new CATH server has been implemented (http://www.biochem.ucl.ac.uk/cgi-bin/cath/CathServer.pl) providing automatic classification of newly determined sequences and structures using a suite of rapid sequence and structure comparison methods. The statistical significance of matches is assessed and links are provided to the putative superfamily or fold group to which the query sequence or structure is assigned.

Databases, Nucleic Acid↗

The mouse genome database (MGD): new features facilitating a model system.

The mouse genome database (MGD, http://www.informatics.jax.org/), the international community database for mouse, provides access to extensive integrated data on the genetics, genomics and biology of the laboratory mouse. The mouse is an excellent and unique animal surrogate for studying normal development and disease processes in humans. Thus, MGD's primary goals are to facilitate the use of mouse models for studying human disease and enable the development of translational research hypotheses based on comparative genotype, phenotype and functional analyses. Core MGD data content includes gene characterization and functions, phenotype and disease model descriptions, DNA and protein sequence data, polymorphisms, gene mapping data and genome coordinates, and comparative gene data focused on mammals. Data are integrated from diverse sources, ranging from major resource centers to individual investigator laboratories and the scientific literature, using a combination of automated processes and expert human curation. MGD collaborates with the bioinformatics community on the development of data and semantic standards, and it incorporates key ontologies into the MGD annotation system, including the Gene Ontology (GO), the Mammalian Phenotype Ontology, and the Anatomical Dictionary for Mouse Development and the Adult Anatomy. MGD is the authoritative source for mouse nomenclature for genes, alleles, and mouse strains, and for GO annotations to mouse genes. MGD provides a unique platform for data mining and hypothesis generation where one can express complex queries simultaneously addressing phenotypic effects, biochemical function and process, sub-cellular location, expression, sequence, polymorphism and mapping data. Both web-based querying and computational access to data are provided. Recent improvements in MGD described here include the incorporation of single nucleotide polymorphism data and search tools, the addition of PIR gene superfamily classifications, phenotype data for NIH-acquired knockout mice, images for mouse phenotypic genotypes, new functional graph displays of GO annotations, and new orthology displays including sequence information and graphic displays.

Animals↗

The CATH domain structure database: new protocols and classification levels give a more comprehensive resource for exploring evolution.

We report the latest release (version 3.0) of the CATH protein domain database (http://www.cathdb.info). There has been a 20% increase in the number of structural domains classified in CATH, up to 86 151 domains. Release 3.0 comprises 1110 fold groups and 2147 homologous superfamilies. To cope with the increases in diverse structural homologues being determined by the structural genomics initiatives, more sensitive methods have been developed for identifying boundaries in multi-domain proteins and for recognising homologues. The CATH classification update is now being driven by an integrated pipeline that links these automated procedures with validation steps, that have been made easier by the provision of information rich web pages summarising comparison scores and relevant links to external sites for each domain being classified. An analysis of the population of domains in the CATH hierarchy and several domain characteristics are presented for version 3.0. We also report an update of the CATH Dictionary of homologous structures (CATH-DHS) which now contains multiple structural alignments, consensus information and functional annotations for 1459 well populated superfamilies in CATH. CATH is directly linked to the Gene3D database which is a projection of CATH structural data onto approximately 2 million sequences in completed genomes and UniProt.

Classification↗

Occupational titles as risk factors for Parkinson's disease.

BACKGROUND: Job title or employment sector may be associated with Parkinson's disease (PD). METHODS: In a case-control study, in four European centres, lifetime occupational histories were coded using modified International Standard Industrial Classification (ISIC) and Dictionary of Occupational Titles (DOT). We employed multiple logistic regression analyses adjusting for age, gender, smoking and family history of PD. RESULTS: A total of 649 cases and 1587 controls were recruited. Scottish data showed a non-significant increased risk for agriculture (DOT: OR 1.32, 95% CI 0.81-2.16; ISIC: OR 1.30, 95% CI 0.84-2.02) and reduced risk for 'transport and communication' (ISIC: OR 0.60, 95% CI 0.37-0.97). Subsequent four-centre analyses showed reduced risk for processing occupations (DOT: OR 0.69, 95% CI 0.5-0.95). An association with pesticide exposure, found using detailed exposure assessment, was not apparent using job classification. CONCLUSIONS: In contrast to retrospective exposure assessment, job or industrial sector is a weak indicator of toxic exposures such that true associations may be missed.

Agricultural Workers' Diseases↗

Biologic synergism and parallelism.

In epidemiologic studies of two binary exposure factors, much attention has been given to the concept of synergism of the factors. The leading dictionary of epidemiology offers two definitions of synergism, one of which this author labels statistical and the other biologic. The epidemiologic literature has been largely concerned with statistical synergism, which is typically measured using additive or multiplicative interaction. This paper focuses on biologic synergism, on the related concept of biologic parallelism, and on the question of how much information can be gleaned about population amounts of biologic synergism and parallelism--information which is of vital interest to epidemiologists. A fundamental identity equates the difference between the amounts of biologic synergism and parallelism to the additive interaction. Two biologic models, the multistage model and the no-hit or immunity model, enhance the interpretation of multiplicative interaction as a measure of statistical synergism, but it is pointed out here that, unfortunately, both models incorporate the strong assumption that there is no parallelism. A third model, the single-hit or vulnerability model, makes the even stronger assumption that there is no biologic synergism and consequently that the additive interaction is equal to minus the amount of parallelism. A consequence of this fact is that a link which has been perceived in the literature to exist between the single-hit model and the additive interaction is false.

Epidemiologic Methods↗

Factors associated with osteoarthritis of the knee in the first national Health and Nutrition Examination Survey (HANES I). Evidence for an association with overweight, race, and physical demands of work.

The authors used data from the United States first national Health and Nutrition Examination Survey of 1971-1975 (HANES I) to explore the cross-sectional associations between radiographic osteoarthritis of the knee and a variety of putative risk factors. A total of 5,193 black and white study participants aged 35-74 years, 315 of whom had x-ray-diagnosed osteoarthritis of the knee, were available for analysis. After controlling for confounders, the authors found significant associations of knee osteoarthritis with overweight, race, and occupation, all of which have been suggested by smaller cross-sectional studies. They then focused specifically on those factors. For overweight, they found a strong association between current obesity and osteoarthritis of the knee, with a dose-response effect not previously assessed. This association was also seen for self-reported minimum adult weight, a proxy for long-term obesity, and was present in persons with asymptomatic osteoarthritis of the knee. These findings strongly suggest that obesity is causative. HANES I was the first study in which racial differences in osteoarthritis of the knee could be assessed within the same country. The black women who were studied had an increased risk of disease (odds ratio (OR) = 2.12, 95% confidence interval (CI) = 1.39-3.23) after controlling for age and weight, although the black men did not. Finally, the authors used the US Department of Labor Dictionary of Occupational Titles to obtain characterizations of the physical demands and knee-bending stress associated with occupations and to study the relation between physical demands of jobs and osteoarthritis of the knee. They found for persons aged 55-64 years an association between knee-bending demands and osteoarthritis of the knee (men, OR = 2.45, 95% CI = 1.21-4.97; women, OR = 3.49, 95% CI = 1.22-10.52). Since such occupational physical demands are common, the authors conclude that they may be associated with a substantial proportion of osteoarthritis of the knee.

Adult↗

PRINTS--a protein motif fingerprint database.

The PRINTS database of protein 'fingerprints' is described. Fingerprints comprise sets of motifs excised from conserved regions of sequence alignments, their diagnostic power or potency being refined by iterative database scanning (in this case the OWL composite sequence database). Generally, the motifs do not overlap, but are separated along a sequence, though they may be contiguous in 3-D space. The use of groups of independent, linearly or spatially separate motifs allows particular protein folds and functionalities to be characterized more flexibly and powerfully than conventional single-component patterns or regular expressions. The current version of the database (4.0) contains 150 entries (encoding > 700 motifs), covering a wide range of globular and membrane proteins, modular polypeptides and so on. The growth of the database is influenced by a number of factors, e.g. the use of multiple motifs, the maximization of sequence information through iterative database scanning and the fact that the database searched is a large composite. The information contained within PRINTS is distinct from but complementary to the single consensus expressions stored in the widely used PROSITE dictionary of patterns.

Amino Acid Sequence↗

Resources for vocational planning applied to the job title of PHYSICAL THERAPIST.

Industrial rehabilitation is a rapidly developing area of health care. As a result, physical therapists need to become functionally familiar with common vocational planning processes and resources. Therefore, the purpose of this article is to describe a process called Vocational Diagnosis and Assessment of Residual Employability (VDARE), which is based on the Dictionary of Occupational Titles (DOT) and Classification of Jobs (COJ) resources. We have provided the DOT and COJ classifications for the job title of PHYSICAL THERAPIST as an example of their terminology. A critique of the DOT and COJ, applied to several occupational examples, suggests these resources be used with supplemental task analyses for a given job. The physical therapist, however, can use the VDARE process and the DOT and COJ resources to identify specific and achievable job targets for clients rather than relying solely on traditional trial and error, on-the-job evaluation.

Ergonomics↗

A novel use for the word "trend" in the clinical trial literature.

BACKGROUND: A new meaning of the word "trend" is appearing in reports of clinical trials. METHODS: Abstracts of all clinical trials in PubMed with English abstracts that contained the word "trend" for each decade from 1971 to 2001 were reviewed. RESULTS: "Trend" was used 36 times in 1981, 170 times in 1991, to 375 times in 2001, most often to refer to a judgment about statistical significance. When the expression "significant trend" was accompanied by a P value, the P value was always less than 0.05; when the expression "nonsignificant trend" was accompanied by a P value, all P values were greater than 0.05; 25% of the time, P values were greater than 0.19, and 5% of the time, they were greater than 0.45. When the unmodified word "trend" was accompanied by a P value, most P values were greater than 0.05. However, on 16 occasions, the P value was less than 0.05, and 5% of the time, P values were greater than 0.20; on 1 occasion it was used with a P value of 0.6. About 30% of the time, the exact meaning of the word could not be determined from the abstract. CONCLUSION: A novel use of the word "trend" has developed in the medical literature over the last several decades. This novel use should be documented in technical dictionaries, and further discussion should occur about the implications of this new connotation. In the meantime, clinical trials reporting can be improved by accompanying the word "trend" with associated P values or point estimates and 95% confidence intervals.

Analysis of Variance↗

An alternative approach to defining the role of the clinical teacher.

BACKGROUND: A number of studies have attempted to identify the components of the clinical teacher role by examining learners' numerical ratings of items on researcher-generated lists. Some of these studies have also compared different groups' perceptions of clinical teaching, but have not directly compared the perceptions of first- and third-year residents. This study addressed two questions: (1) What do residents consider important components of the clinical teacher role? (2) Do first- and third-year residents perceive this role similarly? METHOD: A content analysis was performed on the comments written on evaluation forms by 268 residents about 490 clinical teachers over a five-year period (1980-81 through 1984-85) at a large family practice residency. Of 5,664 forms completed by the residents, 2,388 (42%) contained written comments; comments were on 1,024 (46%) of the first-year resident's forms, 701 (41%) of the second-year residents' forms, and 663 (39%) of the third-year residents' forms. Themes in these comments were coded into a coding dictionary of 157 categories, within 37 clusters, within four roles. RESULTS: The ten highest-ranked categories (Global; Teaching: General; Knowledgeable; Gives Resident Responsibility; Supportive; Miscellaneous; Interested in Teaching; Clinical Competence; Makes Effort to Teach; and Gives Resident Opportunity to Do Procedures) accounted for 41% of the themes coded. The first- and third-year residents differed in the clusters they used to describe their clinical teachers on evaluation forms (chi 2 = 149.86, df = 36, p < .0001). CONCLUSION: The results suggest that content analysis can be used to validly and reliably study residents' written evaluative comments about their teachers. This study contributes to the definition of the clinical teacher role, showing the relative importances of its components, and also supports Stritter's Learning Vector theory, finding the anticipated differences between the comments made by first- and third-year residents.

Clinical Competence↗