Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

Sensitivity and positive predictive value of Medicare Part B physician claims for rheumatologic diagnoses and procedures.

OBJECTIVE: To examine the sensitivity and positive predictive value of Medicare physician claims for select rheumatic conditions managed in rheumatology specialty practices. METHODS: Eight rheumatologists in 3 states abstracted 378 patient office records to obtain information on diagnosis and office procedures. The Medicare Part B physician claims for these patient visits were obtained from the Health Care Financing Administration. The sensitivity of the claims data for a specific diagnosis was calculated as the proportion of all patients whose office records for a particular visit documented that diagnosis and who also had physician claims for that visit which identified that diagnosis. The positive predictive value was evaluated in a separate sample of 331 patient visits identified in Medicare physician claims. The positive predictive value of the claims data for a specific diagnosis was calculated as the proportion of patients with that diagnosis coded in the claims for a particular visit who also had the diagnosis documented in the medical record for that visit. RESULTS: Ninety percent of abstracted office medical records were matched successfully with Medicare physician claims. The sensitivity of the Medicare physician claims was 0.90 (95% confidence interval [CI] 0.85-0.95) for rheumatoid arthritis (RA), 0.85 (95% CI 0.73-0.97) for systemic lupus erythematosus (SLE), and 0.85 (95% CI 0.78-1.0) for aspiration or injection procedures. The sensitivity for osteoarthritis (OA) of the hip or knee was < or = 0.50 if 5-digit codes specifying anatomic site were required. The sensitivity for fibromyalgia (FM) was 0.48 (95% CI 0.28-0.68). The positive predictive values were at least 0.90 for RA, SLE, and aspiration or injection procedures. Positive predictive values for FM and the 5-digit site-specific codes for OA of the knee were 0.83 (95% CI 0.66-1.0) and 0.88 (95% CI 0.75-1.0), respectively, while the positive predictive value of the 5-digit site-specific codes for OA of the hip was zero (95% CI 0-0.26). The positive predictive value of OA at any site was 0.83 (95% CI 0.76-0.90). CONCLUSION: In specialty practice, Medicare physician claims had high sensitivity and positive predictive value for RA, SLE, OA without specification of anatomic site, and injection or aspiration procedures. The claims had lower sensitivity and predictive value for FM and for OA of the hip. The accuracy of Medicare physician claims for other conditions and in the primary care setting requires further investigation.

Aged↗

Localization and DNA sequence analysis of the C gene of bacteriophage Mu, the positive regulator of Mu late transcription.

The C gene of bacteriophage Mu, required for transcription of the phage late genes, was localized by construction and analysis of a series of deleted derivatives of pKN50, a plasmid containing a 9.4 kb Mu DNA fragment which complements Mu C amber mutant phages for growth. One such deleted derivative, pWM10, containing only 0.5 kb of Mu DNA, complements C amber phages and transactivates the mom gene, one of the Mu late genes dependent on C for activation. The DNA sequence of the 0.5 kb fragment predicts a single long open reading frame coding for a 140 amino acid protein. Sequence analysis of DNA containing a C amber mutation located the base change to the second codon of this reading frame. Generation of a frameshift mutation by filling in a BglII site spanning codon 114 of this reading frame resulted in the loss of C complementation and transactivation activity. These results indicate that this open reading frame encodes the Mu C gene product. Comparison of the predicted amino acid sequence of the C protein with those of other transcriptional regulatory proteins revealed some similarity to a region highly conserved among bacterial sigma factors.

Amino Acid Sequence↗

Cytoplasmic proteins interact with a translational control element in the protein-coding region of proopiomelanocortin mRNA.

Previous studies have indicated that proopiomelanocortin (POMC) is translationally regulated. We proposed that the regulatory mechanism involves an interaction between trans-acting protein factors and a cis-acting stem-loop structure in the coding region of POMC mRNA. Functional interactions were tested by examining the translation of mouse POMC mRNA in a rabbit reticulocyte system. Specific binding was demonstrated with ultraviolet-crosslinking and RNA gel mobility shift assays. The evidence presented supports our hypothesis that the translational regulation of POMC gene expression involves recognition of the stem-loop by RNA-binding proteins. Furthermore, POMC stem-loop RNA-binding proteins specifically recognized a predicted stem-loop found in the coding region of corticotropin-releasing hormone, suggesting a novel mechanism of gene regulation that may extend to other neuropeptides as well.

Animals↗

Influence of KfoG on capsular polysaccharide structure in Escherichia coli K4 strain.

An enzyme KfoG with unknown function is coded by the gene kfoG. Gene kfoG belongs to genes from region 2, which are responsible for structure of capsular polysaccharide. Only two enzymes, KfoG and KfoC, coded by genes from region 2, have a glycosyltransferase motif. KfoC is the bifunctional enzyme, which is able to add both GalNAc and GlcUA on nascent polysaccharide, termed chondroitin polymerase. KfoG was predicted to be a fructosyltransferase. The gene that codes the KfoG enzyme was disrupted using homological recombination and absence of this gene was confirmed on both DNA and RNA levels. After disruption no structural changes have been observed, what indicates that fructose branching of the chondroitin backbone is not caused by enzymes, which are coded by genes from region 2 of the K4 capsular gene cluster.

Bacterial Capsules↗

Structural analysis of Arabidopsis thaliana chromosome 5. X. Sequence features of the regions of 3,076,755 bp covered by sixty P1 and TAC clones.

In our ongoing project to deduce the nucleotide sequence of Arabidopsis thaliana chromosome 5, non-redundant P1 and TAC clones have been sequenced on the basis of the fine physical map, and as of January, 2000, the sequences of 16.6 Mb representing approximately 60% of chromosome 5 have been accumulated and released at our web site. Along with the sequence determination, structural features of the sequenced regions have been analyzed by applying a variety of computer programs, and we already predicted a total of 2697 potential protein coding genes in the 11,166,130 bp regions, which are covered by 159 P1 and TAC clones. In this paper, we describe the structural features of the 3,076,755 bp regions covered by newly analyzed 60 P1 and TAC clones. A total of 715 potential protein coding genes were identified, giving an average density of the genes identified of 1 gene per 4001 bp. Introns were observed in 80% of the genes, and the average number per gene and the average length of the introns were 4.5 and 147 bp, respectively. These sequence features are nearly identical to those in our latest report in which the data were compiled based on a new standard of gene assignment including the computer-predicted hypothetical genes. The regions also contained 12 tRNA genes when searched by similarity to reported tRNA genes and the tRNA scan-SE program. The sequence data and information on the potential genes are available through the World Wide Web database KAOS (Kazusa Arabidopsis data Opening Site) at http://www.kazusa.or.jp/kaos/.

Arabidopsis↗

The formation of neural codes in the hippocampus: trace conditioning as a prototypical paradigm for studying the random recoding hypothesis.

The trace version of classical conditioning is used as a prototypical hippocampal-dependent task to study the recoding sequence prediction theory of hippocampal function. This theory conjectures that the hippocampus is a random recoder of sequences and that, once formed, the neuronal codes are suitable for prediction. As such, a trace conditioning paradigm, which requires a timely prediction, seems by far the simplest of the behaviorally-relevant paradigms for studying hippocampal recoding. Parameters that affect the formation of these random codes include the temporal aspects of the behavioral/cognitive paradigm and certain basic characteristics of hippocampal region CA3 anatomy and physiology such as connectivity and activity. Here we describe some of the dynamics of code formation and describe how biological and paradigmatic parameters affect the neural codes that are formed. In addition to a backward cascade of coding neurons, we point out, for the first time, a higher-order dynamic growing out of the backward cascade-a particular forward and backward stabilization of codes as training progresses. We also observe that there is a performance compromise involved in the setting of activity levels due to the existence of three behavioral failure modes. Each of these behavioral failure modes exists in the computational model and, presumably, natural selection produced the compromise performance observed by psychologists. Thus, examining the parametric sensitivities of the codes and their dynamic formation gives insight into the constraints on natural computation and into the computational compromises ensuing from these constraints.

Algorithms↗

Ribosomal protein genes S23 and L35 from amphioxus Branchiostoma belcheri tsingtauense: identification and copy number.

The complete cDNA and deduced amino acid sequences of the ribosomal proteins S23 (AmphiS23) and L35 (AmphiL35) from amphioxus Branchiostoma belcheri tsingtauense were identified in this study. AmphiS23 cDNA is 546 bp long and encodes a protein of 143 amino acids. It has a predicted molecular mass of 15,851 Da and a pI of 10.7. AmphiL35 cDNA comprises 473 bp, and codes for a protein of 123 amino acids with a predicted molecular mass of 14,543 Da and a pI of 10.8. AmphiS23 shares more than 83% identity with its homologues in the vertebrates and more than 84% identity with those in the invertebrates. AmphiL35 is more than 63% identical to its counterparts in the vertebrates and more than 52% identical to those in the invertebrates. Southern blot analysis demonstrated the existence of 1-2 copies of the S23 gene and 2-3 copies of the L35 gene in the genome of amphioxus B. belcheri tsingtauense. This is in sharp contrast to the presence of 6-13 copies of the S23 gene and 15-17 copies of the L35 gene in the rat genome. It is clear that the housekeeping genes like S23 and L35 underwent a large-scale duplication in the vertebrate lineage, reinforcing the gene/genome duplication hypothesis.

Amino Acid Sequence↗

A new subfamily of short bacterial adenylate kinases with the Mycobacterium tuberculosis enzyme as a model: A predictive and experimental study.

The adk gene from Mycobacterium tuberculosis codes for an enzyme of 181 amino acids. A sequence comparison with 52 different forms of adenylate kinases (AK) suggests that the enzyme from M. tuberculosis belongs to a new subfamily of "short" bacterial AKs. The recombinant protein, overexpressed in Escherichia coli, exhibits a low catalytic activity and an unexpectedly high thermal stability (Tm = 64.8 degrees C). Based on various spectroscopic data, on the known three-dimensional structure of the AK from E. coli and on secondary structure predictions for various sequenced AKs, we propose a structural model for AK from M. tuberculosis (AKmt). Proteins 1999;36:238-248.

Adenylate Kinase↗

Universality and Shannon entropy of codon usage.

The distribution functions of codon usage probabilities, computed over all the available GenBank data for 40 eukaryotic biological species and five chloroplasts, are best fitted by the sum of a constant, an exponential, and a linear function in the rank of usage. For mitochondria the analysis is not conclusive. These functions are characterized by parameters that strongly depend on the total guanine and cytosine (GC) content of the coding regions of biological species. It is predicted that the codon usage is the same in all exonic genes with the same GC content. The Shannon entropy for codons, also strongly dependent on the exonic GC content, is computed.

Amino Acids↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

A comprehensive catalog of human KRAB-associated zinc finger genes: insights into the evolutionary history of a large family of transcriptional repressors.

Krüppel-type zinc finger (ZNF) motifs are prevalent components of transcription factor proteins in all eukaryotes. KRAB-ZNF proteins, in which a potent repressor domain is attached to a tandem array of DNA-binding zinc-finger motifs, are specific to tetrapod vertebrates and represent the largest class of ZNF proteins in mammals. To define the full repertoire of human KRAB-ZNF proteins, we searched the genome sequence for key motifs and then constructed and manually curated gene models incorporating those sequences. The resulting gene catalog contains 423 KRAB-ZNF protein-coding loci, yielding alternative transcripts that altogether predict at least 742 structurally distinct proteins. Active rounds of segmental duplication, involving single genes or larger regions and including both tandem and distributed duplication events, have driven the expansion of this mammalian gene family. Comparisons between the human genes and ZNF loci mined from the draft mouse, dog, and chimpanzee genomes not only identified 103 KRAB-ZNF genes that are conserved in mammals but also highlighted a substantial level of lineage-specific change; at least 136 KRAB-ZNF coding genes are primate specific, including many recent duplicates. KRAB-ZNF genes are widely expressed and clustered genes are typically not coregulated, indicating that paralogs have evolved to fill roles in many different biological processes. To facilitate further study, we have developed a Web-based public resource with access to gene models, sequences, and other data, including visualization tools to provide genomic context and interaction with other public data sets.

Computational Biology↗

Evolution of chloroplast mononucleotide microsatellites in Arabidopsis thaliana.

The level of variation and the mutation rate were investigated in an empirical study of 244 chloroplast microsatellites in 15 accessions of Arabidopsis thaliana. In contrast to SNP variation, microsatellite variation in the chloroplast was found to be common, although less common than microsatellite variation in the nucleus. No microsatellite variation was found in coding regions of the chloroplast. To evaluate different models of microsatellite evolution as possible explanations for the observed pattern of variation, the length distribution of microsatellites in the published DNA sequence of the A. thaliana chloroplast was subsequently used. By combining information from these two analyses we found that the mode of evolution of the chloroplast mononucleotide microsatellites was best described by a linear relation between repeat length and mutation rate, when the repeat lengths exceeded about 7 bp. This model can readily predict the variation observed in non-coding chloroplast DNA. It was found that the number of uninterrupted repeat units had a large impact on the level of chloroplast microsatellite variation. No other factors investigated--such as the position of a locus within the chromosome, or imperfect repeats--appeared to affect the variability of chloroplast microsatellites. By fitting the slippage models to the Genbank sequence of chromosome 1, we show that the difference between microsatellite variation in the nucleus and the chloroplast is largely due to differences in slippage rate.

Arabidopsis↗

New genotypes in Fy(a-b-) individuals: nonsense mutations (Trp to stop) in the coding sequence of either FY A or FY B.

Duffy blood group antigens are carried on a glycoprotein that is predicted to pass through the erythrocyte membrane seven times and is a promiscuous chemokine receptor. The Fy(a- b-) phenotype is present in two-thirds of African-American Blacks but is rare in Caucasians. In Blacks, the phenotype is due to a non-functional GATA-1 motif in the FY B, which silences the gene in erythrocytes but not in other tissues, and these patients do not generally make anti-Fyb or anti-Fy3. We describe here the molecular analysis of FY in three unrelated Caucasians who were studied because they had strong anti-Fy3 in their serum. Each was found to have a point mutation that was predicted to change a tryptophan to a premature stop codon in the coding sequence. In one patient (patient 1), the nonsense mutation was at nucleotide 287 of the major transcript in FY A; in another (patient 2), it was at nucleotide 407 in the major transcript of FY B; and in a third (patient 3), it was at nucleotide 408 of the major transcript of FY A.

Aged↗

The role of context-dependent mutations in generating compositional and codon usage bias in grass chloroplast DNA.

The influence of local base composition on mutations in chloroplast DNA (cpDNA) is studied in detail and the resulting, empirically derived, mutation dynamics are used to analyze both base composition and codon usage bias. A 4 x 4 substitution matrix is generated for each of the 16 possible flanking base combinations (contexts) using 17,253 noncoding sites, 1309 of which are variable, from an alignment of three complete grass chloroplast genome sequences. It is shown that substitution bias at these sites is correlated with flanking base composition and that the A+T content of these flanking sites as well as the number of flanking pyrimidines on the same strand appears to have general influences on substitution properties. The context-dependent equilibrium base frequencies predicted from these matrices are then applied to two analyses. The first examines whether or not context dependency of mutations is sufficient to generate average compositional differences between noncoding cpDNA and silent sites of coding sequences. It is found that these two classes of sites exist, on average, in very different contexts and that the observed mutation dynamics are expected to generate significant differences in overall composition bias that are similar to the differences observed in cpDNA. Context dependency, however, cannot account for all of the observed differences: although silent sites in coding regions appear to be at the equilibrium predicted, noncoding cpDNA has a significantly lower A+T content than expected from its own substitution dynamics, possibly due to the influence of indels. The second study examines the codon usage of low-expression chloroplast genes. When context is accounted for, codon usage is very similar to what is predicted by the substitution dynamics of noncoding cpDNA. However, certain codon groups show significant deviation when followed by a purine in a manner suggesting some form of weak selection other than translation efficiency. Overall, the findings indicate that a full understanding of mutational dynamics is critical to understanding the role selection plays in generating composition bias and sequence structure.

Base Composition↗

Further variability within the genus Crinivirus, as revealed by determination of the complete RNA genome sequence of Cucurbit yellow stunting disorder virus.

The complete nucleotide (nt) sequences of genomic RNAs 1 and 2 of Cucurbit yellow stunting disorder virus (CYSDV) were determined for the Spanish isolate CYSDV-AlLM. RNA1 is 9123 nt long and contains at least five open reading frames (ORFs). Computer-assisted analyses identified papain-like protease, methyltransferase, RNA helicase and RNA-dependent RNA polymerase domains in the first two ORFs of RNA1. This is the first study on the sequences of RNA1 from CYSDV. RNA2 is 7976 nt long and contains the hallmark gene array of the family Closteroviridae, characterized by ORFs encoding a heat shock protein 70 homologue, a 59 kDa protein, the major coat protein and a divergent copy of the coat protein. This genome organization resembles that of Sweet potato chlorotic stunt virus (SPCSV), Cucumber yellows virus (CuYV) and Lettuce infectious yellows virus (LIYV), the other three criniviruses sequenced completely to date. However, several differences were observed. The most striking novel features of CYSDV compared to SPCSV, CuYV and LIYV are a unique gene arrangement in the 3'-terminal region of RNA1, the identification in this region of an ORF potentially encoding a protein which has no homologues in any databases, and the prediction of an unusually long 5' non-coding region in RNA2. Additionally, the CYSDV genome resembles that of SPCSV in having very similar 3' regions in RNAs 1 and 2, although for CYSDV similarity in primary structures did not result in predictions of equivalent secondary structures. Overall, these data reinforce the view that the genus Crinivirus contains considerable genetic variation. Additionally, several subgenomic RNAs (sgRNAs) were detected in CYSDV-infected plants, suggesting that generation of sgRNAs is a strategy used by CYSDV for the expression of internal ORFs.

3' Untranslated Regions↗

Molecular cloning and chromosomal assignment of the human homologue of the rat cGMP-inhibited phosphodiesterase 1 (PDE3A)--a gene involved in fat metabolism located at 11p 15.1.

We have cloned the coding region of a human gene, whose predicted amino acid sequence shows 88% homology and higher correspondence in functional domains to the rat cGMP inhibited phosphodiesterase gene (PDE3A). In concordance with the expression data of the rat PDE3A gene, a 5.3-kb transcript of the human cGMP-inhibited phosphodiesterase gene is shown in Northern blot analysis to be highly expressed in adipose tissue. In addition, weaker expression is seen in pancreas, skeletal muscle, liver, placenta, and heart. cDNA clones from the homologue mouse gene were isolated and sequenced spanning a highly conserved region coding for a C-terminal located catalytic core region of this enzyme family. Using a genomic cosmid clone of human PDE3A for fluorescence in situ hybridization, the gene was mapped to chromosomal region 11p15 and regionally sublocalized by PCR on a human-hamster somatic hybrid-cell mapping panel to 11p15.1-p2. Based on comparative linkage data in mouse and rat this chromosomal location is suggested to contain genes involved in complex diseases like obesity and diabetes mellitus type II. Therefore, a possible involvement of the human PDE3A gene in these polygenic traits is discussed, taking into account the prominent role of the rat PDE3A gene product in the antilipolytic action of insulin in adipocytes.

3',5'-Cyclic-AMP Phosphodiesterases↗

Mutation of a family 8 glycosyltransferase gene alters cell wall carbohydrate composition and causes a humidity-sensitive semi-sterile dwarf phenotype in Arabidopsis.

The genome of Arabidopsis thaliana contains about 400 genes coding for glycosyltransferases, many of which are predicted to be involved in the synthesis and remodelling of cell wall components. We describe the isolation of a transposon-tagged mutant, parvus, which under low humidity conditions exhibits a severely dwarfed growth phenotype and failure of anther dehiscence resulting in semi-sterility. All aspects of the mutant phenotype were partially rescued by growth under high-humidity conditions, but not by the application of growth hormones or jasmonic acid. The mutation is caused by insertion of a maize Dissociation (Ds) element in a gene coding for a putative Golgi-localized glycosyltransferase belonging to family 8. Members of this family, originally identified on the basis of similarity to bacterial lipooligosaccharide glycosyltransferases, include enzymes known to be involved in the synthesis of bacterial and plant cell walls. Cell-wall carbohydrate analyses of the parvus mutant indicated reduced levels of rhamnogalacturonan I branching and alterations in the abundance of some xyloglucan linkages that may, however, be indirect consequences of the mutation.

Amino Acid Sequence↗

Signal sequence and keyword trap in silico for selection of full-length human cDNAs encoding secretion or membrane proteins from oligo-capped cDNA libraries.

We have developed an in silico method of selection of human full-length cDNAs encoding secretion or membrane proteins from oligo-capped cDNA libraries. Fullness rates were increased to about 80% by combination of the oligo-capping method and ATGpr, software for prediction of translation start point and the coding potential. Then, using 5'-end single-pass sequences, cDNAs having the signal sequence were selected by PSORT ('signal sequence trap'). We also applied 'secretion or membrane protein-related keyword trap' based on the result of BLAST search against the SWISS-PROT database for the cDNAs which could not be selected by PSORT. Using the above procedures, 789 cDNAs were primarily selected and subjected to full-length sequencing, and 334 of these cDNAs were finally selected as novel. Most of the cDNAs (295 cDNAs: 88.3%) were predicted to encode secretion or membrane proteins. In particular, 165(80.5%) of the 205 cDNAs selected by PSORT were predicted to have signal sequences, while 70 (54.2%) of the 129 cDNAs selected by 'keyword trap' preserved the secretion or membrane protein-related keywords. Many important cDNAs were obtained, including transporters, receptors, and ligands, involved in significant cellular functions. Thus, an efficient method of selecting secretion or membrane protein-encoding cDNAs was developed by combining the above four procedures.

5' Flanking Region↗