Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Analysis of the thymidine kinase genes of macropodid herpesviruses 1 and 2.

The nucleotide sequences of the entire protein coding regions of the thymidine kinase (TK) genes of macropodid herpesvirus type 1 (MaHV-1) and type 2 (MaHV-2) were determined. The coding region of the MaHV-1 TK gene was 984 bp long and was predicted to encode a polypeptide of 327 amino acids. The coding region of the MaHV-2 TK gene was 1020 bp long and encoded a polypeptide of 340 amino acids. Comparisons of their deduced amino acid sequences with those of fifteen other herpesviruses revealed close homology to those of other alphaherpesviruses, particularly to human herpesvirus type 1 (HHV-1) and type 2 (HHV-2).

Alphaherpesvirinae↗

Dimensionality of amino acid space and solvent accessibility prediction with neural networks.

Solvent accessibility prediction from amino acid sequences has been pursued by several researchers. Such a prediction typically starts by transforming the amino acid category (or type) information into numerical representations. All twenty amino acids can be completely and uniquely represented by 20-dimensional vectors. Here, we investigate if the amino acid space defined in this way really requires twenty dimensions. We tried to develop corresponding representations in fewer dimensions. A method for searching optimal codification schema in an arbitrary space using neural networks was developed. The method is used to obtain optimal encoding of amino acids at various levels of dimensionality, and applied to optimize the amino acid codifications for the prediction of the solvent accessibility values of the proteins using feed-forward neural networks. The traditional 20-dimensional codification seems to be redundant in solving the solvent accessibility prediction problem, since a 1-dimensional codification is able to achieve almost the same degree of accuracy as the 20-dimensional codification. Optimal coding in much fewer dimensions could be used to make the predictions of accessible surface area with almost the same degree of accuracy as that obtained by a fully unique 20-dimensional coding. The 1-dimensional amino acid codification for solvent accessibility prediction obtained by a purely mathematical way based on neural networks is highly correlated with a physical property of the amino acids, namely their average solvent accessibility. The method developed to find the optimal codification is general, although the codification thus produced is dependent on the type of estimated property.

Amino Acid Sequence↗

Expressed emotion, parenting stress, and adjustment in mothers of young children with behavior problems.

Expressed Emotion (EE), a measure of the emotional climate of the family, predicts subsequent adjustment of adults with mental disorder (Leff & Vaughn, 1985). Despite the acknowledged importance of the family in childhood disorders, there have been relatively few studies of expressed emotion with adolescents and school-aged children and virtually none focused on preschoolers. The present study utilized the Five Minute Speech Sample (FMSS) to examine how Expressed Emotion relates concurrently and longitudinally to child problem status in a community sample of 112 preschool-aged children. At preschool, the proportion of high EE increased significantly across three child groups: Comparison (8.1%), Borderline Problem (15.8%), and High Problem (41.2%); however, preschool EE was not predictive of subsequent child status at 1st grade. Expanded FMSS codes. tapping positive affect and worry about the child, were also related to child problem group at preschool and were predictive of subsequent child status at 1st grade. Because parents' stress and adjustment were also highly related to child problem group status, we examined whether the FMSS codes were essentially a proxy for these or whether they explained unique variance. In two stepwise regressions on preschool child group status (divided by total problems and by externalizing problems), maternal stress was the only variable to enter. Also, in predicting to 1st grade externalizing child group status, only maternal stress entered. Discussion focused on the extension of the EE construct and other FMSS coding to young children, and the need to recognize that to some extent these variables may reflect maternal stress and adjustment.

Adult↗

Value of ICD-9 coded chief complaints for detection of epidemics.

To assess the value of ICD-9 coded chief complaints for early detection of epidemics, we measured sensitivity, positive predictive value, and timeliness of Influenza detection using a respiratory set (RS) of ICD-9 codes and an Influenza set (IS). We also measured inherent timeliness of these data using the cross-correlation function. We found that, for a one-year period, the detectors had sensitivity of 100% (1/1 epidemic) and positive predictive values of 50% (1/2) for RS and 25% (1/4) for IS. The timeliness of detection using ICD-9 coded chief complaints was one week earlier than the detection using Pneumonia and Influenza deaths (the gold standard). The inherent timeliness of ICD-9 data measured by the cross-correlation function was two weeks earlier than the gold standard.

Disease↗

Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data.

OBJECTIVES: Implementation of the International Statistical Classification of Disease and Related Health Problems, 10th Revision (ICD-10) coding system presents challenges for using administrative data. Recognizing this, we conducted a multistep process to develop ICD-10 coding algorithms to define Charlson and Elixhauser comorbidities in administrative data and assess the performance of the resulting algorithms. METHODS: ICD-10 coding algorithms were developed by "translation" of the ICD-9-CM codes constituting Deyo's (for Charlson comorbidities) and Elixhauser's coding algorithms and by physicians' assessment of the face-validity of selected ICD-10 codes. The process of carefully developing ICD-10 algorithms also produced modified and enhanced ICD-9-CM coding algorithms for the Charlson and Elixhauser comorbidities. We then used data on in-patients aged 18 years and older in ICD-9-CM and ICD-10 administrative hospital discharge data from a Canadian health region to assess the comorbidity frequencies and mortality prediction achieved by the original ICD-9-CM algorithms, the enhanced ICD-9-CM algorithms, and the new ICD-10 coding algorithms. RESULTS: Among 56,585 patients in the ICD-9-CM data and 58,805 patients in the ICD-10 data, frequencies of the 17 Charlson comorbidities and the 30 Elixhauser comorbidities remained generally similar across algorithms. The new ICD-10 and enhanced ICD-9-CM coding algorithms either matched or outperformed the original Deyo and Elixhauser ICD-9-CM coding algorithms in predicting in-hospital mortality. The C-statistic was 0.842 for Deyo's ICD-9-CM coding algorithm, 0.860 for the ICD-10 coding algorithm, and 0.859 for the enhanced ICD-9-CM coding algorithm, 0.868 for the original Elixhauser ICD-9-CM coding algorithm, 0.870 for the ICD-10 coding algorithm and 0.878 for the enhanced ICD-9-CM coding algorithm. CONCLUSIONS: These newly developed ICD-10 and ICD-9-CM comorbidity coding algorithms produce similar estimates of comorbidity prevalence in administrative data, and may outperform existing ICD-9-CM coding algorithms.

Algorithms↗

Identification of the Rev transactivation and Rev-responsive elements of feline immunodeficiency virus.

Spliced messages encoded by two distinct strains of feline immunodeficiency virus (FIV) were identified. Two of the cDNA clones represented mRNAs with bicistronic capacity. The first coding exon contained a short open reading frame (orf) of unknown function, designated orf 2. After a translational stop, this exon contained the L region of the env orf. The L region resides 5' to the predicted leader sequence of env. The second coding exon contained the H orf, which began 3' to env and extended into the U3 region of the long terminal repeat. The in-frame splicing of the L and H orfs created the FIV rev gene. Site-directed antibodies to the L orf recognized a 23-kDa protein in infected cells. Immunofluorescence studies localized Rev to the nucleoli of infected cells. The Rev-responsive element (RRE) of FIV was initially identified by computer analysis. Three independent isolates of FIV were searched in their entirety for regions with unusual RNA-folding properties. An unusual RNA-folding region was not found at the Su-TM junction but instead was located at the end of env. Minimal-energy foldings of this region revealed a structure that was highly conserved among the three isolates. Transient expression assays demonstrated that both the Rev and RRE components of FIV were necessary for efficient reporter gene expression. Cells stably transfected with rev-deleted proviruses produced virion-associated reverse transcriptase activity only when FIV Rev was supplied in trans. Thus, FIV is dependent on a fully functional Rev protein and an RRE for productive infection.

Amino Acid Sequence↗

Rearrangement of side-chains in a Zif268 mutant highlights the complexities of zinc finger-DNA recognition.

Structural and biochemical studies of Cys(2)His(2) zinc finger proteins initially led several groups to propose a "recognition code" involving a simple set of rules relating key amino acid residues in the zinc finger protein to bases in its DNA site. One recent study from our group, involving geometric analysis of protein-DNA interactions, has discussed limitations of this idea and has shown how the spatial relationship between the polypeptide backbone and the DNA helps to determine what contacts are possible at any given position in a protein-DNA complex. Here we report a study of a zinc finger variant that highlights yet another source of complexity inherent in protein-DNA recognition. In particular, we find that mutations can cause key side-chains to rearrange at the protein-DNA interface without fundamental changes in the spatial relationship between the polypeptide backbone and the DNA. This is clear from a simple analysis of the binding site preferences and co-crystal structures for the Asp20-->Ala point mutant of Zif268. This point mutation in finger one changes the specificity of the protein from GCG TGG GCG to GCG TGG GC(G/T), and we have solved crystal structures of the D20A mutant bound to both types of sites. The structure of the D20A mutant bound to the GCG site reveals that contacts from key residues in the recognition helix are coupled in complex ways. The structure of the complex with the GCT site also shows an important new water molecule at the protein-DNA interface. These side-chain/side-chain interactions, and resultant changes in hydration at the interface, affect binding specificity in ways that cannot be predicted either from a simple recognition code or from analysis of spatial relationships at the protein-DNA interface. Accurate computer modeling of protein-DNA interfaces remains a challenging problem and will require systematic strategies for modeling side-chain rearrangements and change in hydration.

Base Sequence↗

A novel class of ectomycorrhiza-regulated cell wall polypeptides in Pisolithus tinctorius.

Development of the ectomycorrhizal symbiosis leads to the aggregation of fungal hyphae to form the mantle. To identify cell surface proteins involved in this developmental step, changes in the biosynthesis of fungal cell wall proteins were examined in Eucalyptus globulus-Pisolithus tinctorius ectomycorrhizas by two-dimensional polyacrylamide gel electrophoresis. Enhanced synthesis of several immunologically related fungal 31- and 32-kDa polypeptides, so-called symbiosis-regulated acidic polypeptides (SRAPs), was observed. Peptide sequences of SRAP32d were obtained after trypsin digestion. These peptides were found in the predicted sequence of six closely related fungal cDNAs coding for ectomycorrhiza up-regulated transcripts. The PtSRAP32 cDNAs represented about 10% of the differentially expressed cDNAs in ectomycorrhiza and are predicted to encode alanine-rich proteins of 28.2 kDa. There are no sequence homologies between SRAPs and previously identified proteins, but they contain the Arg-Gly-Asp (RGD) motif found in cell-adhesion proteins. SRAPs were observed on the hyphal surface by immunoelectron microscopy. They were also found in the host cell wall when P. tinctorius attached to the root surface. RNA blot analysis showed that the steady-state level of PtSRAP32 transcripts exhibited a drastic up-regulation when fungal hyphae form the mantle. These results suggest that SRAPs may form part of a cell-cell adhesion system needed for aggregation of hyphae in ectomycorrhizas.

Amino Acid Sequence↗

Effects of carrier frequency and background noise on the detection of mixed modulation.

This article is concerned with the mechanisms underlying the detection of amplitude modulation (AM), frequency modulation (FM), and mixed modulation (MM), i.e., simultaneously occurring AM and FM. In a previous study [B. C. J. Moore and A. Sek, J. Acoust. Soc. Am. 92, 3119-3131 (1992)], psychometric functions were measured for the detection of AM alone and FM alone, using a 10-Hz modulation rate and a 1-kHz carrier frequency. Detectability was then measured for combined AM and FM, with modulation depths selected so that each type of modulation would be equally detectable if presented alone. The detectability of the combined AM and FM was better than would be predicted if the two types of modulation were coded completely independently. Significant effects of relative modulator phase were found when detectability was relatively high, but these effects were not correctly predicted by either of two excitation-pattern models considered. The first experiment reported here was similar to the earlier experiment, but performance was compared for carrier frequencies of 1 and 6 kHz; at the latter frequency, neural synchrony to the stimulus fine structure (phase locking) does not occur. The results at both carrier frequencies were similar to those of our earlier experiment, suggesting that the presence or absence of phase-locking information plays little role in the detection of MM. The second experiment was again similar, but bands of noise were used to mask selectively either the upper or lower side of the excitation pattern of the modulated carrier. The phase effects in this case were in the direction predicted by excitation pattern models. The overall pattern of the results could be predicted reasonably well using a multichannel excitation pattern model based on the assumption that listeners use an unweighted sum of decision variables across all suprathreshold channels with a positive signal-to-noise ratio.

Auditory Perception↗

The relationship between base composition and codon usage in bacterial genes and its use for the simple and reliable identification of protein-coding sequences.

Bacterial genes that code for proteins appear to possess a codon usage characteristic of their overall base composition. This results in different but predictable non-random distributions of nucleotides within codons, permitting the recognition of protein-coding sequences in a wide range of bacterial species. The nature of this distribution depends on the base composition of the coding sequence. The position-specific differences are especially conspicuous in genes of extreme G + C content, allowing the particularly reliable prediction of the reading frame and coding strand of experimentally determined DNA sequences. This finding has been exploited to identify the coding sequence of the viomycin phosphotransferase (vph) gene of Streptomyces vinaceus. An easily applied computer program ("Frame") has been written to carry out and display such analyses.

Bacterial Proteins↗

Genomic structure and organization of the human rBAT gene (SLC3A1).

Cystinuria is an autosomal recessive disorder of amino acid transport, manifesting as three phenotypes (I, II, and III). An amino acid transport gene, rBAT, is responsible for cystinuria. Mutation and linkage analyses have demonstrated the disease to be heterogeneous, with rBAT being the defective gene in type I cystinuria. The genomic structure of the human rBAT gene (HGMW-approved symbol SLC 3A1) has been established via two strategies: (i) construction of two different genomic libraries by subcloning the Mega-YAC921B6 (CEPH), containing rBAT, in Lambda ZAP and screening using rBAT cDNA and different PCR products; and (ii) generation and sequencing of genomic fragments by long PCR using rBAT cDNA-derived primers. The rBAT gene spans approximately 45 kb and consists of 10 exons. The introns range from 500 to 13,000 bp. All splice sites conform to the GT/AG rule. The promoter region has been further analyzed, and a predicted TATA box 98 bp upstream of the first coding ATG was identified. In addition an Alu repeat has been detected 72 bp upstream of the predicted TATA box.

Amino Acid Transport Systems, Basic↗

International Classification of Diseases, 9th Revision, Clinical Modification codes in discharge abstracts are poor measures of complication occurrence in medical inpatients.

OBJECTIVES: The authors tested the ability of International Classification of Diseases, 9th Revision, Clinical Modification (ICD-9-CM) codes in discharge abstracts to identify medical inpatients who experienced an in-hospital complication, using complications identified through chart review as the gold standard. METHODS: Two sets of ICD-9-CM codes were used: an inclusive set including many medical diagnoses that may also be coexistent complicating conditions on admission rather than complications and an exclusive set consisting primarily of ICD-9-CM-specified complication and adverse drug event codes. RESULTS: Neither set performed well as a diagnostic test for complication occurrence according to receiver operating characteristic analysis (ROC areas were 0.61 for the inclusive set and 0.55 for the exclusive set). Sensitivities of the ICD-9-CM codes for complications were 0.34 for the inclusive set and 0.14 for the exclusive set. Corresponding positive predictive values were 0.32 and 0.37, respectively. Sensitivities of code definitions for individual complications were generally poor, less than 0.5 in most cases. CONCLUSIONS: The authors conclude that ICD-9-CM codes in discharge abstracts are poor measures of complication occurrence.

Abstracting and Indexing↗

Computational methods for the identification of genes in vertebrate genomic sequences.

Research into new methods to identify genes in anonymous genomic sequences has been going on for more than 15 years. Over this period of time, the field has evolved from the designing of programs to identify protein coding regions in compact mitochondrial or bacterial genomes, to the challenge of predicting the detailed organization of multi-exon vertebrate genes. The best program currently available perfectly locates more than 80% of the internal coding exons, and only 5% of the predictions do not overlap a real exon. Given such accuracy, computational methods are indeed very useful; however, they do not alleviate the need for experimental validation. If the performances are satisfactory for the identification of the coding moiety of genes (internal coding exons), the determination of the full extent of the transcript (5' and 3' extremities of the gene) and the location of promoter regions are still unreliable. As the human and mouse genome sequencing projects enter a production mode, the fully automated annotation of megabase-long anonymous genomic sequences is the next big challenge in bioinformatics.

Animals↗

JIGSAW: integration of multiple sources of evidence for gene prediction.

MOTIVATION: Computational gene finding systems play an important role in finding new human genes, although no systems are yet accurate enough to predict all or even most protein-coding regions perfectly. Ab initio programs can be augmented by evidence such as expression data or protein sequence homology, which improves their performance. The amount of such evidence continues to grow, but computational methods continue to have difficulty predicting genes when the evidence is conflicting or incomplete. Genome annotation pipelines collect a variety of types of evidence about gene structure and synthesize the results, which can then be refined further through manual, expert curation of gene models. RESULTS: JIGSAW is a new gene finding system designed to automate the process of predicting gene structure from multiple sources of evidence, with results that often match the performance of human curators. JIGSAW computes the relative weight of different lines of evidence using statistics generated from a training set, and then combines the evidence using dynamic programming. Our results show that JIGSAW's performance is superior to ab initio gene finding methods and to other pipelines such as Ensembl. Even without evidence from alignment to known genes, JIGSAW can substantially improve gene prediction accuracy as compared with existing methods. AVAILABILITY: JIGSAW is available as an open source software package at http://cbcb.umd.edu/software/jigsaw.

Algorithms↗

Polyoma virus. The early region and its T-antigens.

The DNA sequence of the early coding region of polyoma virus is presented. It consists of 2739 nucleotides. The sequence predicts that more than one reading frame can be used to code for the three known polyoma virus early proteins (designated small, middle and large T-antigens). From the DNA sequence, the 'splicing' signals used in the processing of viral RNA to functional messenger RNAs can be predicted, as well as the sizes and sequences of the three proteins. Other unusual aspects of the DNA sequence are noted. Comparisons are made between the DNA sequences and the predicted amino acid sequences of the respective large T-antigens of polyoma virus and the related virus Simian Virus (SV) 40.

Antigens, Viral↗

Intensity perception. XIII. Perceptual anchor model of context-coding.

In our preliminary theory of intensity resolution [e.g., see N. I. Durlach and L. D. Braida, J. Acoust. Soc. Am. 46, 372-383 (1969)], two modes of memory operation are postulated: the trace mode and the context-coding mode. In this paper, we present a revised model of the context-coding mode which describes explicitly a process by which sensations are coded relative to the context and which predicts a resolution edge effect [L. D. Braida and N. I. Durlach, J. Acoust. Soc. Am. 51, 483-502 (1972); J. E. Berliner, L. D. Braida, and N. I. Durlach, J. Acoust. Soc. Am. 61, 1256-1267 (1977)]. The sensation arising from a given stimulus presentation is coded by determining its distance from internal references or perceptual anchors. The noise in this process, combined with the sensation noise, constitutes the limitation on resolution in the model. In the revised model the probability density functions of the decision variable are not precisely Gaussian (and cannot be expressed analytically in closed form). This paper outlines the predictions of the model for one-interval paradigms and for fixed-level two-interval paradigms and derives estimates of the values of model parameters.

Discrimination Learning↗

Aspects of large-scale chromatin structures in mouse liver nuclei can be predicted from the DNA sequence.

The large amount of non-coding DNA present in mammalian genomes suggests that some of it may play a structural or functional role. We provide evidence that it is possible to predict computationally, from the DNA sequence, loci in mouse liver nuclei that possess distinctive nucleosome arrays. We tested the hypothesis that a 100 kb region of DNA possessing a strong, in-phase, dinucleosome period oscillation in the motif period-10 non-T, A/T, G, should generate a nucleosome array with a nucleosome repeat that is one-half of the dinucleosome oscillation period value, as computed by Fourier analysis of the sequence. Ten loci with short repeats, that would be readily distinguishable from the pervasive bulk repeat, were predicted computationally and then tested experimentally. We estimated experimentally that less than 20% of the chromatin in mouse liver nuclei has a nucleosome repeat length that is 15 bp, or more, shorter than the bulk repeat value of 195 +/- bp. All 10 computational predictions were confirmed experimentally with high statistical significance. Nucleosome repeats as short as 172 +/- 5 bp were observed for the first time in mouse liver chromatin. These findings may be useful for identifying distinctive chromatin structures computationally from the DNA sequence.

Animals↗

A microRNA expression signature of human solid tumors defines cancer gene targets.

Small noncoding microRNAs (miRNAs) can contribute to cancer development and progression and are differentially expressed in normal tissues and cancers. From a large-scale miRnome analysis on 540 samples including lung, breast, stomach, prostate, colon, and pancreatic tumors, we identified a solid cancer miRNA signature composed by a large portion of overexpressed miRNAs. Among these miRNAs are some with well characterized cancer association, such as miR-17-5p, miR-20a, miR-21, miR-92, miR-106a, and miR-155. The predicted targets for the differentially expressed miRNAs are significantly enriched for protein-coding tumor suppressors and oncogenes (P < 0.0001). A number of the predicted targets, including the tumor suppressors RB1 (Retinoblastoma 1) and TGFBR2 (transforming growth factor, beta receptor II) genes were confirmed experimentally. Our results indicate that miRNAs are extensively involved in cancer pathogenesis of solid tumors and support their function as either dominant or recessive cancer genes.

Gene Expression Profiling↗