Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Evaluation of the exon predictions of the GRAIL software.

Identification of transcribed sequences within genomic regions has been a major rate-limiting step in the pursuit of genes. We have tested the ability of GRAIL 1 and GRAIL 2 to locate coding exons in the 41 longest human genomic sequences available in release 38 of the EMBL database. Of the total annotated exons GRAIL 1 correctly identified 68% of coding exons. GRAIL 2 correctly identified 71%. Of the 686 predictions made by GRAIL 1, 46.4% did not correspond to an annotated exon. The corresponding figures for GRAIL 2 were 556 and 30.5%. The GRAIL 2 software predicted as excellent a significantly higher proportion of the coding exons than did GRAIL 1, and also gave a significantly lower figure for false predictions. GRAIL 2 thus represents a significant improvement over GRAIL 1.

Exons↗

A model-based approach to selection of tag SNPs.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are the most common type of polymorphisms found in the human genome. Effective genetic association studies require the identification of sets of tag SNPs that capture as much haplotype information as possible. Tag SNP selection is analogous to the problem of data compression in information theory. According to Shannon's framework, the optimal tag set maximizes the entropy of the tag SNPs subject to constraints on the number of SNPs. This approach requires an appropriate probabilistic model. Compared to simple measures of Linkage Disequilibrium (LD), a good model of haplotype sequences can more accurately account for LD structure. It also provides a machinery for the prediction of tagged SNPs and thereby to assess the performances of tag sets through their ability to predict larger SNP sets. RESULTS: Here, we compute the description code-lengths of SNP data for an array of models and we develop tag SNP selection methods based on these models and the strategy of entropy maximization. Using data sets from the HapMap and ENCODE projects, we show that the hidden Markov model introduced by Li and Stephens outperforms the other models in several aspects: description code-length of SNP data, information content of tag sets, and prediction of tagged SNPs. This is the first use of this model in the context of tag SNP selection. CONCLUSION: Our study provides strong evidence that the tag sets selected by our best method, based on Li and Stephens model, outperform those chosen by several existing methods. The results also suggest that information content evaluated with a good model is more sensitive for assessing the quality of a tagging set than the correct prediction rate of tagged SNPs. Besides, we show that haplotype phase uncertainty has an almost negligible impact on the ability of good tag sets to predict tagged SNPs. This justifies the selection of tag SNPs on the basis of haplotype informativeness, although genotyping studies do not directly assess haplotypes. A software that implements our approach is available.

Databases, Genetic↗

Fast evaluation of internal loops in RNA secondary structure prediction.

MOTIVATION: Though not as abundant in known biological processes as proteins, RNA molecules serve as more than mere intermediaries between DNA and proteins. Research in the last 15 years demonstrates that RNA molecules serve in many roles, including catalysis. Furthermore, RNA secondary structure prediction based on free energy rules for stacking and loop formation remains one of the few major breakthroughs in the field of structure prediction, as minimum free energy structures and related quantities can be computed with full mathematical rigor. However, with the current energy parameters, the algorithms used hitherto suffer the disadvantage of either employing heuristics that risk (though highly unlikely) missing the optimal structure or becoming prohibitively time consuming for moderate to large sequences. RESULTS: We present a new method to evaluate internal loops utilizing currently used energy rules. This method reduces the time complexity of this part of the structure prediction from O(n4) to O(n3), thus reducing the overall complexity to O(n3). Even when the size of evaluated internal loops is bounded by k (a commonly used heuristic), the method presented has a competitive edge by reducing the time complexity of internal loop evaluation from O(k2n2) to O(kn2). The method also applies to the calculation of the equilibrium partition function. AVAILABILITY: Source code for an RNA secondary structure prediction program implementing this method is available at ftp://www.ibc.wustl.edu/pub/zuker/zuker .tar.Z

Algorithms↗

Transcriptional organization and posttranscriptional regulation of the Bacillus subtilis branched-chain amino acid biosynthesis genes.

In Bacillus subtilis, the genes of the branched-chain amino acids biosynthetic pathway are organized in three genetic loci: the ilvBHC-leuABCD (ilv-leu) operon, ilvA, and ilvD. These genes, as well as ybgE, encoding a branched-chain amino acid aminotransferase, were recently demonstrated to represent direct targets of the global transcriptional regulator CodY. In the present study, the transcriptional organization and posttranscriptional regulation of these genes were analyzed. Whereas ybgE and ilvD are transcribed monocistronically, the ilvA gene forms a bicistronic operon with the downstream located ypmP gene, encoding a protein of unknown function. The ypmP gene is also directly preceded by a promoter sharing the regulatory pattern of the ilvA promoter. The ilv-leu operon revealed complex posttranscriptional regulation: three mRNA species of 8.5, 5.8, and 1.2 kb were detected. Among them, the 8.5-kb full-length primary transcript exhibits the shortest half-life (1.2 min). Endoribonucleolytic cleavage of this transcript generates the 5.8-kb mRNA, which lacks the coding sequences of the first two genes of the operon and is predicted to carry a stem-loop structure at its 5' end. This processing product has a significantly longer half-life (3 min) than the full-length precursor. The most stable transcript (half-life, 7.6 min) is the 1.2-kb mRNA generated by the processing event and exonucleolytic degradation of the large transcripts or partial transcriptional termination. This mRNA, which encompasses exclusively the ilvC coding sequence, is predicted to carry a further stable stem-loop structure at its 3' end. The very different steady-state amounts of mRNA resulting from their different stabilities are also reflected at the protein level: proteome studies revealed that the cellular amount of IlvC protein is 10-fold greater than that of the other proteins encoded by the ilv-leu operon. Therefore, differential segmental stability resulting from mRNA processing ensures the fine-tuning of the expression of the individual genes of the operon.

Amino Acid Sequence↗

The nucleotide sequence and proposed genome organization of oat chlorotic stunt virus, a new soil-borne virus of cereals.

The complete genomic sequence of a new virus, first found infecting oats in Wales, UK, has been determined. The genome is a positive-sense ssRNA molecule, 4114 nucleotides in length, examination of which indicates the presence of four ORFs. The first ORF initiating at the 5' terminus (ORF1) encodes a protein with a predicted M(r) of 23476 (p23). ORF2 extends through the amber termination codon of ORF1 to give a protein with a predicted M(r) of 84355 (p84). The readthrough domain of p84 contains amino acid sequence similarities with a number of putative RNA-dependent RNA polymerases. ORF3 is in a different reading frame from ORF1/2 and encodes a protein with an M(r) of 48231 (p48), identified as the coat protein by direct peptide sequencing. ORF4 nests within ORF3 but is in a different frame from it and codes for a protein with a predicted M(r) of 8220 (p8). Comparisons of peptide sequence, particularly within the putative polymerase region and within the S domain of the coat protein, highlight similarities with members of both the tombusvirus and carmovirus groups. The coat protein region shows most similarity with members of the tombusvirus group, whilst the size and predicted strategy of the genome seem to be intermediate between that of the carmovirus and tombusvirus groups. These features highlight possible evolutionary links with each group whilst being distinct from both. We propose the name of oat chlorotic stunt for this new virus.

Amino Acid Sequence↗

Sequence analysis of genome segments S4 and S8 of Mal de Río Cuarto virus (MRCV): evidence that the virus should be a separate Fijivirus species.

This is the first sequence-based characterization of Mal de Río Cuarto virus (MRCV), currently classified as a variant of Maize rough dwarf virus (MRDV) and exclusively found in South America. We sequenced and analyzed genome segments S4 and S8. MRCV S4 coded for a putative 131.67 kDa protein while MRCV S8 coded for a putative 68.26 kDa protein containing an ATP/GTP-binding motif. The 5' and 3' ends of MRCV segments, were 5'AAGUUUUU3' and 5'CAGCUnnnGUC3', respectively. Prediction of secondary structure of both segments coding strands showed that terminal regions were able to form structures that are proposed to be replication and packaging signals. MRCV S4 showed identity to members of Fijivirus as well as to two other genera of the Reoviridae family. MRCV S8 revealed identity with Rice black streaked dwarf virus (RBSDV) S8, MRDV S7, Oat sterile dwarf virus (OSDV) S9 and Nilaparvata lugens reovirus (NLRV) S7. While MRDV and RBSDV segments are highly homologous between each other, MRCV identity levels with them was considerably lower. We discussed the evolutionary relationships of MRCV to other Reoviridae, and based on phylogenetic analysis we proposed that although MRCV is related to MRDV, it could be regarded as a new species of the Fijivirus genus.

Amino Acid Sequence↗

Histone modifications: signalling receptors and potential elements of a heritable epigenetic code.

The genetic code epitomises simplicity, near universality and absolute predictive power. By contrast, epigenetic information, in the form of histone modifications, is characterised by complexity, diversity and an overall tendency to respond to changes in genomic function rather than to predict them. Perhaps the transient changes in histone modifications involved in intranuclear signalling and ongoing chromatin functions mask stable, predictive modifications that lie beneath. The current rapid progress in unravelling the diversity and complexity of epigenetic information might eventually reveal an underlying histone or epigenetic code. But whether it does or not, it will certainly provide unprecedented opportunities, both for understanding how the genome responds to environmental and metabolic change and for manipulating its activities for experimental and therapeutic benefit.

Chromatin↗

A probabilistic model for detecting coding regions in DNA sequences.

A probabilistic model is presented to predict whether or not an anonymous sequence of DNA contains exons. The method is shown to be at least as reliable as Grail, a well-known neural network solution to the problem, and to be significantly more amenable to customization for specific prediction problems.

Base Sequence↗

The effects of interaction with the device described by procedural text on recall, true/false, and task performance.

In two experiments, subjects interacted to different extents with relevant devices while reading two complex multistep procedural texts and were then tested with task performance time, true/false, and recall measures. While reading, subjects performed the task (read and do), saw the experimenter perform the task (read and see experimenter do), imagined doing the task (read and imagine), looked at the device while reading (read and see), or only read (read only). Van Dijk and Kintsch's (1983) text representation theory led to the prediction that exposure to the task device (in the read-and-do, read-and-see, and read-and-see-experimenter-do conditions) would lead to the development of a stronger situation model and therefore faster task performance, whereas the read-only and read-and-see conditions would lead to a better textbase, and therefore better performance on the true/false and recall tasks. Paivio's (1991) dual coding theory led to the opposite prediction for recall. The results supported the text representation theory with task performance and recall. The read-and-see condition produced consistently good performance on the true/false measure. Amount of text study time contributed to recall performance. These findings support the notion that information available while reading leads to differential development of representations in memory, which, in turn, causes differences in performance on various measures.

Adult↗

A computational and experimental approach to validating annotations and gene predictions in the Drosophila melanogaster genome.

Five years after the completion of the sequence of the Drosophila melanogaster genome, the number of protein-coding genes it contains remains a matter of debate; the number of computational gene predictions greatly exceeds the number of validated gene annotations. We have assembled a collection of >10,000 gene predictions that do not overlap existing gene annotations and have developed a process for their validation that allows us to efficiently prioritize and experimentally validate predictions from various sources by sequencing RT-PCR products to confirm gene structures. Our data provide experimental evidence for 122 protein-coding genes. Our analyses suggest that the entire collection of predictions contains only approximately 700 additional protein-coding genes. Although we cannot rule out the discovery of genes with unusual features that make them refractory to existing methods, our results suggest that the D. melanogaster genome contains approximately 14,000 protein-coding genes.

Animals↗

Comparison of two Caenorhabditis genes encoding FMRFamide(Phe-Met-Arg-Phe-NH2)-like peptides.

We have identified a gene encoding multiple FMRFamide-like peptides in the necromenic nematode Caenorhabditis vulgaris. This gene, Cv-flp-1, shares strong sequence homology in the coding regions with the flp-1 gene from the related free-living soil nematode C. elegans. The predicted neuropeptide precursor proteins differ by only four conservative amino acid changes, none of which affects sequences of the predicted neuropeptides. DNA sequences in the non-coding areas are less conserved, but areas of sequence homology are found in introns and in 3' and 5' non-translated regions, suggesting some functional significance for these conserved regions. In C. vulgaris, as was found in C. elegans, two transcripts are presumably produced as a result of use of an alternative 3' splice acceptor site. Lastly, an antibody specific for the RF-moiety of FMRFamide stains a similar subset of cells in C. elegans and C. vulgaris. These results indicate that the function and regulation of the peptides are likely to be conserved in both species.

Alternative Splicing↗

Depletion of retinal dopamine does not affect the ERG b-wave increment threshold function in goldfish in vivo.

Increment threshold functions of the electroretinogram (ERG) b-wave were obtained from goldfish using an in vivo preparation to study intraretinal mechanisms underlying the increase in perceived brightness induced by depletion of retinal dopamine by 6-hydroxydopamine (6-OHDA). Goldfish received unilateral intraocular injections of 6-OHDA plus pargyline on successive days. Depletion of retinal dopamine was confirmed by the absence of tyrosine-hydroxylase immunoreactivity at 2 to 3 weeks postinjection as compared to sham-injected eyes from the same fish. There was no difference among normal, sham-injected or 6-OHDA-injected eyes with regard to ERG waveform, intensity-response functions or increment threshold functions. Dopamine-depleted eyes showed a Purkinje shift, that is, a transition from rod-to-cone dominated vision with increasing levels of adaptation. We conclude (1) dopamine-depleted eyes are capable of photopic vision; and (2) the ERG b-wave is not diagnostic for luminosity coding at photopic backgrounds. We also predict that (1) dopamine is not required for the transition from scotopic to photopic vision in goldfish; (2) the ERG b-wave in goldfish is influenced by chromatic interactions; (3) horizontal cell spinules, though correlated with photopic mechanisms in the fish retina, are not necessary for the transition from scotopic to photopic vision; and (4) the OFF pathway, not the ON pathway, is involved in the action of dopamine on luminosity coding in the retina.

Animals↗

Nucleotide sequence and genetic organization of barley stripe mosaic virus RNA gamma.

The complete nucleotide sequences of RNA gamma from the Type and ND18 strains of barley stripe mosaic virus (BSMV) have been determined. The sequences are 3164 (Type) and 2791 (ND18) nucleotides in length. Both sequences contain a 5'-noncoding region (87 or 88 nucleotides) which is followed by a long open reading frame (ORF1). A 42-nucleotide intercistronic region separates ORF1 from a second, shorter open reading frame (ORF2) located near the 3'-end of the RNA. There is a high degree of homology between the Type and ND18 strains in the nucleotide sequence of ORF1. However, the Type strain contains a 366 nucleotide direct tandem repeat within ORF1 which is absent in the ND18 strain. Consequently, the predicted translation product of Type RNA gamma ORF1 (mol wt 87,312) is significantly larger than that of ND18 RNA gamma ORF1 (mol wt 74,011). The amino acid sequence of the ORF1 polypeptide contains homologies with putative RNA polymerases from other RNA viruses, suggesting that this protein may function in replication of the BSMV genome. The nucleotide sequence of RNA gamma ORF2 is nearly identical in the Type and ND18 strains. ORF2 codes for a polypeptide with a predicted molecular weight of 17,209 (Type) or 17,074 (ND18) which is known to be translated from a subgenomic (sg) RNA. The initiation point of this sgRNA has been mapped to a location 27 nucleotides upstream of the ORF2 initiation codon in the intercistronic region between ORF1 and ORF2. The sgRNA is not coterminal with the 3'-end of the genomic RNA, but instead contains heterogeneous poly(A) termini up to 150 nucleotides long (J. Stanley, R. Hanau, and A. O. Jackson, 1984, Virology 139, 375-383). In the genomic RNA gamma, ORF2 is followed by a short poly(A) tract and a 238-nucleotide tRNA-like structure.

Amino Acid Sequence↗

A measure for predicting audibility discrimination thresholds for spectral envelope distortions in vowel sounds.

Both in speech synthesis and in sound coding it is often beneficial to have a measure that predicts whether, and to what extent, two sounds are different. This paper addresses the problem of estimating the perceptual effects of small modifications to the spectral envelope of a harmonic sound. A recently proposed auditory model is investigated that transforms the physical spectrum into a pattern of specific loudness as a function of critical band rate. A distance measure based on the concept of partial loudness is presented, which treats detectability in terms of a partial loudness threshold. This approach is adapted to the problem of estimating discrimination thresholds related to modifications of the spectral envelope of synthetic vowels. Data obtained from subjective listening tests using a representative set of stimuli in a 3IFC adaptive procedure show that the model makes reasonably good predictions of the discrimination threshold. Systematic deviations from the predicted thresholds may be related to individual differences in auditory filter selectivity. The partial loudness measure is compared with previously proposed distance measures such as the Euclidean distance between excitation patterns and between specific loudness applied to the same experimental data. An objective test measure shows that the partial loudness measure and the Euclidean distance of the excitation patterns are equally appropriate as distance measures for predicting audibility thresholds. The Euclidean distance between specific loudness is worse in performance compared with the other two.

Auditory Threshold↗

Contrasting effects of age of acquisition in lexical decision and letter detection.

Two experiments studied attention in beginning and skilled readers of Dutch to letter information in function words and content words. Early and late acquired nouns and function words were presented to third-grade students and skilled adolescent readers. Target words were presented in short story contexts, as in the study of Greenberg, Koriat, and Vellutino (1998). Target nouns were matched on word frequency. Predictions of the structural account hypothesis of letter detection (Koriat, Greenberg, & Goldshmid, 1991) were confirmed. No age-of-acquisition effect was found. In contrast, a separately conducted lexical decision experiment using the same content word stimulus sets showed shorter decision latencies for early acquired words. The combined results suggest that during silent reading, when attention is focused on meaning, phonological processes may play a less prominent role than in lexical decision tasks that demand explicit control of phonological codes. The letter detection results confirmed predictions of the structural account hypothesis for both beginning and skilled readers. Taken together, these studies demonstrate that phonological processes in silent reading may play a less prominent role and that the structural account of letter processing is valid for languages other than Hebrew and English but probably is not the unique mechanism involved in letter detection.

Adolescent↗

Dose distributions and dose rate constants for new ytterbium-169 brachytherapy seeds.

Relative dose distributions in the vicinities of two new prototypes of ytterbium-169 brachytherapy sources have been measured using lithium fluoride thermoluminescent dosimeters (TLD) placed in a solid water phantom. The type 6 seed consists of four Yb2O3 spheres contained in a 0.075-mm thick titanium tube, with overall dimensions of 4.5 mm in length and 0.8 mm in diameter. The type 8 seed contains a cylindrical segment of Yb wire, 2.5 mm long, sealed in a 0.075-mm thick titanium tube 4 mm in length and 0.5 mm in diameter. The relative dose distributions, measured with TLD, were compared with those predicted by Monte Carlo simulations (MCNP code). Experimental and theoretical dose distributions were generally in agreement within 5% of the local dose value. Source strength was determined from activity measurements using a high purity germanium (HPGe) spectrometer. With source strengths established, thermoluminescence was then used to measure the dose rate constant, lambda 0, for the type 8 seed. The measured lambda 0 of 1.34 +/- 0.10 cGy U-1 for the type 8 seed agrees, within statistical uncertainty, with the value of 1.25 +/- 0.05 cGy U-1 predicted through Monte Carlo simulation. Comparisons are made with experimental data reported by other investigators.

Brachytherapy↗

Improved lossless intra coding for H.264/MPEG-4 AVC.

A new lossless intra coding method based on sample-by-sample differential pulse code modulation (DPCM) is presented as an enhancement of the H.264/MPEG-4 AVC standard. The H.264/AVC design includes a multidirectional spatial prediction method to reduce spatial redundancy by using neighboring samples as a prediction for the samples in a block of data to be encoded. In the new lossless intra coding method, the spatial prediction is performed based on samplewise DPCM instead of in the block-based manner used in the current H.264/AVC standard, while the block structure is retained for the residual difference entropy coding process. We show that the new method, based on samplewise DPCM, does not have a major complexity penalty, despite its apparent pipeline dependencies. Experiments show that the new lossless intra coding method reduces the bit rate by approximately 12% in comparison with the lossless intra coding method previously included in the H.264/AVC standard. As a result, the new method is currently being adopted into the H.264/AVC standard in a new enhancement project.

Algorithms↗

Translation of the flavivirus kunjin NS3 gene in cis but not its RNA sequence or secondary structure is essential for efficient RNA packaging.

Our previous studies using trans-complementation analysis of Kunjin virus (KUN) full-length cDNA clones harboring in-frame deletions in the NS3 gene demonstrated the inability of these defective complemented RNAs to be packaged into virus particles (W. J. Liu, P. L. Sedlak, N. Kondratieva, and A. A. Khromykh, J. Virol. 76:10766-10775). In this study we aimed to establish whether this requirement for NS3 in RNA packaging is determined by the secondary RNA structure of the NS3 gene or by the essential role of the translated NS3 gene product. Multiple silent mutations of three computer-predicted stable RNA structures in the NS3 coding region of KUN replicon RNA aimed at disrupting RNA secondary structure without affecting amino acid sequence did not affect RNA replication and packaging into virus-like particles in the packaging cell line, thus demonstrating that the predicted conserved RNA structures in the NS3 gene do not play a role in RNA replication and/or packaging. In contrast, double frameshift mutations in the NS3 coding region of full-length KUN RNA, producing scrambled NS3 protein but retaining secondary RNA structure, resulted in the loss of ability of these defective RNAs to be packaged into virus particles in complementation experiments in KUN replicon-expressing cells. Furthermore, the more robust complementation-packaging system based on established stable cell lines producing large amounts of complemented replicating NS3-deficient replicon RNAs and infection with KUN virus to provide structural proteins also failed to detect any secreted virus-like particles containing packaged NS3-deficient replicon RNAs. These results have now firmly established the requirement of KUN NS3 protein translated in cis for genome packaging into virus particles.

Animals↗