Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Building Codes”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Building an ontology of pulmonary diseases with natural language processing tools using textual corpora.

Pathologies and acts are classified in thesauri to help physicians to code their activity. In practice, the use of thesauri is not sufficient to reduce variability in coding and thesauri are not suitable for computer processing. We think the automation of the coding task requires a conceptual modeling of medical items: an ontology. Our task is to help lung specialists code acts and diagnoses with software that represents medical knowledge of this concerned specialty by an ontology. The objective of the reported work was to build an ontology of pulmonary diseases dedicated to the coding process. To carry out this objective, we develop a precise methodological process for the knowledge engineer in order to build various types of medical ontologies. This process is based on the need to express precisely in natural language the meaning of each concept using differential semantics principles. A differential ontology is a hierarchy of concepts and relationships organized according to their similarities and differences. Our main research hypothesis is to apply natural language processing tools to corpora to develop the resources needed to build the ontology. We consider two corpora, one composed of patient discharge summaries and the other being a teaching book. We propose to combine two approaches to enrich the ontology building: (i) a method which consists of building terminological resources through distributional analysis and (ii) a method based on the observation of corpus sequences in order to reveal semantic relationships. Our ontology currently includes 1550 concepts and the software implementing the coding process is still under development. Results show that the proposed approach is operational and indicates that the combination of these methods and the comparison of the resulting terminological structures give interesting clues to a knowledge engineer for the building of an ontology.

France↗

Locating protein coding regions in human DNA using a decision tree algorithm.

Genes in eukaryotic DNA cover hundreds or thousands of base pairs, while the regions of those genes that code for proteins may occupy only a small percentage of the sequence. Identifying the coding regions is of vital importance in understanding these genes. Many recent research efforts have studied computational methods for distinguishing between coding and noncoding regions, and several promising results have been reported. We describe here a new approach, using a machine learning system that builds decision trees from the data. This approach combines several coding measures to produce classifiers with consistently higher accuracies than previous methods, on DNA sequences ranging from 54 to 162 base pairs in length. The algorithm is very efficient, and it can easily be adapted to different sequence lengths. Our conclusion is that decision trees are a highly effective tool for identifying protein coding regions.

Algorithms↗

Using compound codes for automatic classification of clinical diagnoses.

Classification of diagnoses (a.k.a. coding) is the central part of current concept based medical IR systems. Some classification systems contain over 30,000 distinct codes which makes classifying clinical documents a time consuming labor intensive and error prone process. This paper presents a simple methodology for cleaning up and reusing existing manually coded diagnostic statements mainly extracted from clinical notes to build predictive models using a sparse-feature implementation of a Naïve Bayes classifier. One of the problems addressed is that diagnostic statements often contain several diagnoses and are assigned several codes resulting in a multi-class classification problem. We investigate one possible way of addressing this problem by introducing compound (multiple code) categories. We present experimental results of classifying >16,000 randomly selected diagnostic strings into 19 top level categories. A small improvement (3%) with using compound categories over simple categories indicates that using multiple code categories is a promising solution, although clearly in need of further research and refinement.

Abstracting and Indexing↗

Differential codon usage for conserved amino acids: evidence that the serine codons TCN were primordial.

The availability of specialized sequence databanks for Escherichia coli, Saccharomyces cerevisiae and Bacillus subtilis made it possible to build a set of 105 protein-coding genes that are homologous in these three species. An analysis of the triplets at both the nucleotide and amino acid level revealed that the codon bias of some amino acids are significantly higher at conserved rather than at non-conserved positions. Comparisons of homologous genes in E. coli and Salmonella typhimurium, and in S. cerevisiae and Drosophila melanogaster, led to the same conclusion. A special case was made for serine in E. coli, whose major codon is AGC for non-conserved and TCC for conserved residues. We interpret this observation as evidence that the primordial codons for serine were TCN, while codons AGY appeared later. This conclusion is substantiated by an analysis of the codon usage of catalytic serine residues in ancient, ubiquitous and essential proteins (ATP synthases and topoisomerases). It is shown that in these proteins the proportion of the catalytic serine residues coded by TCN is significantly higher than the one expected from the overall codon usage of serine residues.

Amino Acid Sequence↗

Assessment and requirements of nuclear reaction databases for GCR transport in the atmosphere and structures.

The transport properties of galactic cosmic rays (GCR) in the atmosphere, material structures, and human body (self-shielding) are of interest in risk assessment for supersonic and subsonic aircraft and for space travel in low-Earth orbit and on interplanetary missions. Nuclear reactions, such as knockout and fragmentation, present large modifications of particle type and energies of the galactic cosmic rays in penetrating materials. We make an assessment of the current nuclear reaction models and improvements in these model for developing required transport code data bases. A new fragmentation data base (QMSFRG) based on microscopic models is compared to the NUCFRG2 model and implications for shield assessment made using the HZETRN radiation transport code. For deep penetration problems, the build-up of light particles, such as nucleons, light clusters and mesons from nuclear reactions in conjunction with the absorption of the heavy ions, leads to the dominance of the charge Z = 0, 1, and 2 hadrons in the exposures at large penetration depths. Light particles are produced through nuclear or cluster knockout and in evaporation events with characteristically distinct spectra which play unique roles in the build-up of secondary radiation's in shielding. We describe models of light particle production in nucleon and heavy ion induced reactions and make an assessment of the importance of light particle multiplicity and spectral parameters in these exposures.

Aerospace Medicine↗

Procedure to protect confidentiality of familial data in community genetics and genomic research.

The collection of familial data is an essential step for community genetics programs or genetic research. Ethical issues concerning privacy and confidentiality present a major challenge in such programs. In order to keep familial data confidential, we have developed a family-based numerical coding procedure which allows the use of confidential data and the determination of familial relationships without risk of disclosure. This procedure is composed of two parts: the physical separation of identifying information and individual data; and the use of a code containing all the information required to build family trees. This procedure has been used in Eastern Quebec since 1995, mainly for screening, genetic counseling, research on familial dyslipidemias, public health intervention, and research projects on the genetics of complex traits, such as arterial hypertension and coronary artery disease.

Abstracting and Indexing↗

Algorithms for the identification of prevalent diabetes in the All of Us Research Program validated using polygenic scores.

The All of Us Research Program (AoU) is an initiative designed to gather a comprehensive and diverse dataset from at least one million individuals across the USA. This longitudinal cohort study aims to advance research by providing a rich resource of genetic and phenotypic information, enabling powerful studies on the epidemiology and genetics of human diseases. One critical challenge to maximizing its use is the development of accurate algorithms that can efficiently and accurately identify well-defined disease and disease-free participants for case-control studies. This study aimed to develop and validate type 1 (T1D) and type 2 diabetes (T2D) algorithms in the AoU cohort, using electronic health record (EHR) and survey data. Building on existing algorithms and using diagnosis codes, medications, laboratory results, and survey data, we developed and implemented algorithms for identifying prevalent cases of type 1 and type 2 diabetes. The first set of algorithms used only EHR data (EHR-only), and the second set used a combination of EHR and survey data (EHR+). A universal algorithm was also developed to identify individuals without diabetes. The performance of each algorithm was evaluated by testing its association with polygenic scores (PSs) for type 1 and type 2 diabetes. We demonstrated the feasibility and utility of using AoU EHR and survey data to employ diabetes algorithms. For T1D, the EHR-only algorithm showed a stronger association with T1D-PS compared to the EHR + algorithm (DeLong p-value = 3 × 10-5). For T2D, the EHR + algorithm outperformed both the EHR-only and the existing T2D definition provided in the AoU Phenotyping Library (DeLong p-values = 0.03 and 1 × 10-4, respectively), identifying 25.79% and 22.57% more cases, respectively, and providing an improved association with T2D PS. We provide a new validated type 1 diabetes definition and an improved type 2 diabetes definition in AoU, which are freely available for diabetes research in the AoU. These algorithms ensure consistency of diabetes definitions in the cohort, facilitating high-quality diabetes research.

Humans↗

Beliefs about control in the physician-patient relationship: effect on communication in medical encounters.

OBJECTIVES: Effective communication is a critical component of quality health care, and to improve it we must understand its dynamics. This investigation examined the extent to which physicians' and patients' preferences for control in their relationship (e.g., shared control vs doctor control) were related to their communications styles and adaptations (i.e., how they responded to the communication of the other participant). DESIGN: Stratified case-controlled study. PATIENTS/PARTICIPANTS: Twenty family medicine and internal medicine physicians and 135 patients. MEASUREMENTS: Based on scores from the Patient-Practitioner Orientation Scale, 10 patient-centered physicians (5 male, 5 female) and 10 doctor-centered physicians (5 male, 5 female) each interacted with 5 to 8 patients, roughly half of whom preferred shared control and the other half of whom were oriented toward doctor control. Audiotapes of 135 consultations were coded for behaviors indicative of physician partnership building and active patient participation. MAIN RESULTS: Patients who preferred shared control were more active participants (i.e., expressed more opinions, concerns, and questions) than were patients oriented toward doctor control. Physicians' beliefs about control were not related to their use of partnership building. However, physicians did use more partnership building with male patients. Not only were active patient participation and physician partnership building mutually predictive of each other, but also approximately 14% of patient participation was prompted by physician partnership building and 33% of physician partnership building was in response to active patient participation. CONCLUSIONS: Communication in medical encounters is influenced by the physician's and patient's beliefs about control in their relationship as well as by one another's behavior. The relationship between physicians' partnership building and active patient participation is one of mutual influence such that increases in one often lead to increases in the other.

Case-Control Studies↗

The Ensembl automatic gene annotation system.

As more genomes are sequenced, there is an increasing need for automated first-pass annotation which allows timely access to important genomic information. The Ensembl gene-building system enables fast automated annotation of eukaryotic genomes. It annotates genes based on evidence derived from known protein, cDNA, and EST sequences. The gene-building system rests on top of the core Ensembl (MySQL) database schema and Perl Application Programming Interface (API), and the data generated are accessible through the Ensembl genome browser (http://www.ensembl.org). To date, the Ensembl predicted gene sets are available for the A. gambiae, C. briggsae, zebrafish, mouse, rat, and human genomes and have been heavily relied upon in the publication of the human, mouse, rat, and A. gambiae genome sequence analysis. Here we describe in detail the gene-building system and the algorithms involved. All code and data are freely available from http://www.ensembl.org.

Animals↗

A method for fast evaluation of neutron spectra for BNCT based on in-phantom figure-of-merit calculation.

In this paper a fast method to evaluate neutron spectra for brain BNCT is developed. The method is based on an algorithm to calculate dose distribution in the brain, for which a data matrix has been taken into account, containing weighted biological doses per position per incident energy and the incident neutron spectrum to be evaluated. To build the matrix, using the MCNP 4C code, nearly monoenergetic neutrons were transported into a head model. The doses were scored and an energy-dependent function to biologically weight the doses was used. To find the beam quality, dose distribution along the beam centerline was calculated. A neutron importance function for this therapy to bilaterally treat deep-seated tumors was constructed in terms of neutron energy. Neutrons in the energy range of a few tens of kilo-electron-volts were found to produce the best dose gain, defined as dose to tumor divided by maximum dose to healthy tissue. Various neutron spectra were evaluated through this method. An accelerator-based neutron source was found to be more reliable for this therapy in terms of therapeutic gain than reactors.

Algorithms↗

The effect of fluoridated milk on bovine dental enamel.

The concept of intra-oral cariogenicity tests, using naturally accumulated plaque, was pioneered by Koulourides (Koulourides et al., 1974). Tests rely on changes in microhardness of dental enamel after exposure to substrates. We used this model, with significant modifications, to determine the possible caries-protective effect of fluoridated milk. Custom-made cast chrome intra-oral appliances were made to fit the lower arches of volunteers. Four removable, highly polished 3 x 4 mm gauze-covered bovine enamel blocks were slotted into the appliances. These were worn for 48 h so that plaque would build up. The enamel was color-coded with composite (Kerr Kolors, Kerr, Romulus, MI) to ensure error-free removal and immersion extra-orally in the coded test substrates.

Animals↗

Engineering in software testing: statistical testing based on a usage model applied to medical device development.

When a population is too large for exhaustive study, as is the case for all possible uses of a software system, a statistically correct sample must be drawn as a basis for inferences about the population. A Markov chain usage model is an engineering formalism that represents the population of possible uses for which a product is to be tested. In statistical testing of software based on a Markov chain usage model, the rich body of analytical results available for Markov chains provides numerous insights that can be used in both product development and test planing. A usage model is based on specifications rather than code, so insights that result from model building can inform product decisions in the early stages of a project when the opportunity to prevent problems is the greatest. Statistical testing based on a usage model provides a sound scientific basis for quantifying the reliability of software.

Equipment Safety↗

Transposable elements as the key to a 21st century view of evolution.

Cells are capable of sophisticated information processing. Cellular signal transduction networks serve to compute data from multiple inputs and make decisions about cellular behavior. Genomes are organized like integrated computer programs as systems of routines and subroutines, not as a collection of independent genetic 'units'. DNA sequences which do not code for protein structure determine the system architecture of the genome. Repetitive DNA elements serve as tags to mark and integrate different protein coding sequences into coordinately functioning groups, to build up systems for genome replication and distribution to daughter cells, and to organize chromatin. Genomes can be reorganized through the action of cellular systems for cutting, splicing and rearranging DNA molecules. Natural genetic engineering systems (including transposable elements) are capable of acting genome-wide and not just one site at a time. Transposable elements are subject to regulation by cellular signal transduction/computing networks. This regulation acts on both the timing and extent of DNA rearrangements and (in a few documented cases so far) on the location of changes in the genomes. By connecting transcriptional regulatory circuits to the action of natural genetic engineering systems, there is a plausible molecular basis for coordinated changes in the genome subject to biologically meaningful feedback.

DNA Transposable Elements↗

BUILD: a program generator for modelling experimental biological data.

BUILD is a program generator acting at source code level. The generated code corresponds to a whole application in order to model a biological process of interest using an iterative adjustment of experimental data. The program is designed to be executed in command line mode for processing of multiple data files with an individual execution control for each file. The results are completed by modular statistical and graphical functions. This approach has been shown to reduce the time and the amount of work needed for program development, debugging and maintenance. To date, BUILD has been successfully used in mathematical analysis of phenomenological approaches, but other fields of activity, such as educational software, are also conceivable.

Algorithms↗

Medication therapy management services: a critical review.

OBJECTIVE: To identify and examine medication therapy management (MTM) practice and compensation models currently being used by public and private sector programs, develop a model for payers to consider in compensating pharmacists for the provision of MTM services, and review how a relative value-based payment system based upon Current Procedural Terminology (CPT) codes might apply to MTM services. DATA SOURCES: Peer-reviewed literature; study of existing MTM practice and compensation models; interviews with pharmacists, pharmacy benefit providers, health plans, and policy makers; structured discussions with industry experts. SUMMARY: Implementation of MTM represents an opportunity for pharmacists to provide public and private payers with examples of service packages and business models that improve patient therapeutic outcomes. MTM services can lead to overall cost reductions and improved health outcomes. Recommendations for pharmacists, health plans, and Medicare Part D prescription drug plan sponsors are provided. Pharmacists should standardize and package MTM services at varying levels of intensity; determine work values for MTM CPT codes, use standards for billing and service delivery, build supply capacity to meet demand for MTM services, and cultivate patient and provider support for pharmacist-provided MTM services. Plans and sponsors should develop mechanisms for measuring MTM impact on overall health care costs and develop payment systems to cover costs as well as sustain and provide for growth in the number of providers. CONCLUSION: The essential components of MTM business and payment models, as outlined in this article, can be effectively mapped to relative value-based CPT codes for pharmacist-provided MTM services. Pharmacy providers, after considering various factors and conditions in their own environment, can develop an optimal MTM service package and business model based on this information.

Drug Therapy↗

The emerging functionality of endogenous lectins: A primer to the concept and a case study on galectins including medical implications.

Biochemistry textbooks commonly make it appear that it is a foregone conclusion that the hardware of biological information storage and transfer is confined to nucleotides and amino acids, the letters of the genetic code. However, the remarkable talents of a third class of biomolecules are often overlooked. For example, one of them far surpasses the building blocks of nucleic acids and proteins in terms of theoretical coding capacity by oligomer formation. Although often exclusively assigned to duties in energy metabolism, carbohydrates as part of cellular glycoconjugates (glycoproteins, proteoglycans, glycolipids) have, in fact, other important tasks. Currently, they are increasingly gaining recognition as an operative high-density information coding system. An elaborate enzymatic machinery enables cells to be versatile enough to produce a glycan profile (glycome) that is as characteristic as a fingerprint. Moreover, swift modifications during dynamic processes, such as differentiation or malignant transformation, are readily possible. The translation of the information presented in oligosaccharide determinants to biological responses is carried out by lectins. Recognition of foreign glycosignatures in innate immunity, regulation of cell-cell/matrix interactions, cell migration or growth, and intra- and intercellular glycan routing etc represent physiologically far-reaching lectin-carbohydrate functionality. The classification of endogenous lectins is guided by sequence alignments and conservation of distinct structural traits. For example, a jelly-roll-like folding pattern and maintenance of key residue positioning involved in stacking and C-H/pi-interactions as well as directional hydrogen bonds to the 1-galactoside ligands are common denominators among galectins. Biochemical and biophysical studies are beginning to unravel the intricacies of the selection of a limited set of endogenous ligands, such as certain integrins or ganglioside GM1, and combined with biological cell experiments, its relevance for cell sociology, e.g. in growth regulation and tumor cell invasion or activated T cell apoptosis. Histopathological monitoring accompanies the biological cell investigations, linking expression of certain family members to tumor progression or suppression. Further insights into the functional consequences of the sugar code's translation are thus expected to have notable repercussions for diagnostic and therapeutic procedures.

Amino Acid Sequence↗

Exploitation of the selectivity-conferring code of nonribosomal peptide synthetases for the rational design of novel peptide antibiotics.

Recently, the solved crystal structure of a phenylalanine-activating adenylation (A) domain enlightened the structural basis for the specific recognition of the cognate substrate amino acid in nonribosomal peptide synthetases (NRPSs). By adding sequence comparisons and homology modeling, we successfully used this information to decipher the selectivity-conferring code of NRPSs. Each codon combines the 10 amino residues of a NRPS A domain that are presumed to build up the substrate-binding pocket. In this study, the deciphered code was exploited for the first time to rationally alter the substrate specificity of whole NRPS modules in vitro and in vivo. First, the single-residue Lys239 of the L-Glu-activating initiation module C-A(Glu)-PCP of the surfactin synthetase A was mutated to Gln239 to achieve a perfect match to the postulated L-Gln-activating binding pocket. Biochemical characterization of the mutant protein C-A(Glu)-PCP(Lys239 --> Gln) revealed the postulated alteration in substrate specificity from L-Glu to L-Gln without decrease in catalytic efficiency. Second, according to the selectivity-conferring code, the binding pockets of L-Asp and L-Asn-activating A domains differs in three positions: Val299 versus Ile, His322 versus Glu, and Ile330 versus Val, respectively. Thus, the binding pocket of the recombinant A domain AspA, derived from the second module of the surfactin synthetases B, was stepwisely adapted for the recognition of L-Asn. Biochemical characterization of single, double, and triple mutants revealed that His322 represents a key position, whose mutation was sufficient to give rise to the intended selectivity-switch. Subsequently, the gene fragment encoding the single-mutant AspA(His322 --> Glu) was introduced back into the surfactin biosynthetic gene cluster. The resulting Bacillus subtilis strain was found to produce the expected so far unknown lipoheptapeptide [Asn(5)]surfactin. This indicates that site-directed mutagenesis, guided by the selectivity-conferring code of NRPS A domains, represents a powerful alternative for the genetic manipulation of NRPS biosynthetic templates and the rational design of novel peptide antibiotics.

Anti-Bacterial Agents↗