Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Self-identification of protein-coding regions in microbial genomes.

A new method for predicting protein-coding regions in microbial genomic DNA sequences is presented. It uses an ab initio iterative Markov modeling procedure to automatically perform the partition of genomic sequences into three subsets shown to correspond to coding, coding on the opposite strand, and noncoding segments. In contrast to current methods, such as GENEMARK [Borodovsky, M. & McIninch, J. D. (1993) Comput. Chem. 17, 123-133], no training set or prior knowledge of the statistical properties of the studied genome are required. This new method tolerates error rates of 1-2% and can process unassembled sequences. It is thus ideal for the analysis of genome survey and/or fragmented sequence data from uncharacterized microorganisms. The method was validated on 10 complete bacterial genomes (from four major phylogenetic lineages). The results show that protein-coding regions can be identified with an accuracy of up to 90% with a totally automated and objective procedure.

Algorithms↗

Quantifying variability in neural responses and its application for the validation of model predictions.

A rate code assumes that a neuron's response is completely characterized by its time-varying mean firing rate. This assumption has successfully described neural responses in many systems. The noise in rate coding neurons can be quantified by the coherence function or the correlation coefficient between the neuron's deterministic time-varying mean rate and noise corrupted single spike trains. Because of the finite data size, the mean rate cannot be known exactly and must be approximated. We introduce novel unbiased estimators for the measures of coherence and correlation which are based on the extrapolation of the signal to noise ratio in the neural response to infinite data size. We then describe the application of these estimates to the validation of the class of stimulus-response models that assume that the mean firing rate captures all the information embedded in the neural response. We explain how these quantifiers can be used to separate response prediction errors that are due to inaccurate model assumptions from errors due to noise inherent in neuronal spike trains.

Action Potentials↗

[Positive predictive value of ICD-9-CM codes in hospital inpatient discharge abstract records, for identifying adverse events].

A retrospective study was conducted in the ambit of Risk Management research, in order to assess adverse events in patients hospitalised in hospitals in one Local Health Authority of the Piemonte region. Specifically, the aims of the study were to: evaluate the relative frequency of ICD-9-CM codes used to define adverse events, with respect to the total number of hospital discharge records submitted in 2003; identify true and false positives, by hospital chart review; estimate the positive predictive value (VPP) of the ICD-9-CM codes used, and determine, in each case, whether the adverse event had led to hospitalisation or if it had occurred during hospitalisation. Results show that the ICD-9-CM codes used effectively identify adverse events. In fact, the probability that an ICD-9-CM code will accurately identify an adverse event is 100% for codes in the "Misadventures of surgical and medical care" category of adverse events, 62.8% for codes indicating "Complications of medications (adverse drug events)" and 56.8% for the "Complications of surgical or medical procedures" category. In most cases the adverse event had occurred prior to hospital admission.

Adolescent↗

Effects of phonological and orthographic neighbourhood density interact in visual word recognition.

The present study investigated the role of phonological and orthographic neighbourhood density in visual word recognition. Three mechanisms were identified that predict distinct facilitatory or inhibitory effects of each variable. The lexical competition account predicts overall inhibitory effects of neighbourhood density. The global activation (familiarity) account predicts overall facilitatory effects of neighbourhood density. Finally, the cross-code consistency account predicts an interaction, with inhibition of phonological neighbours in sparse orthographic regions and facilitation of phonological neighbours in dense orthographic regions. In Experiment 1 (lexical decision), a cross-over interaction was indeed found, supporting the prediction of the cross-code consistency account. In Experiment 2, this cross-over interaction was exaggerated by adding pseudohomo-phone stimuli (e.g., brane) among the nonword targets. Finally, in Experiment 3 (progressive demasking), we tried to shift the balance between inhibitory and facilitatory mechanisms by using a perceptual identification task. As predicted, the inhibitory effects of phonological neighbourhood were amplified, whereas the facilitatory effects disappeared. We conclude that the level of compatibility across co-activated orthographic and phonological representations is a major causal factor underlying this pattern of effects.

Humans↗

Secondary DNA structure analysis of the coding strand switch regions of five Leishmania major Friedlin chromosomes.

As part of the EULEISH international genome project, a region of 74,674 nucleotides from chromosome 21 of Leishmania major Friedlin was subcloned and sequenced; and 31 new coding sequences were predicted. Of particular interest was a unique coding strand switching region covering 1.6 kb of DNA; and this was subjected to further investigation. Bioinformatic analysis of this region revealed an unusually high AT composition, a lack of putative hairpins and a strong curvature of the DNA in agreement with the structural characteristics of similar regions of other Leishmania chromosomes. These observations and a comparison with the secondary DNA structure of four other Leishmania chromosomes and chromosomes of different organisms could suggest a functional role of this region in transcription and mitotic division.

Animals↗

An odorant derivative as an antagonist for an olfactory receptor.

Different odorants are recognized by different combinations of G protein-coupled olfactory receptors, and thereby, odor identity is determined by a combinatorial receptor code for each odorant. We recently demonstrated that odorants appeared to compete for receptor sites to act as an agonist or an antagonist. Therefore, in natural circumstances where we always perceive a mixture of various odorants, olfactory receptor antagonism between odorants may result in a receptor code for the mixture that cannot be predicted from the codes for its individual components. Here we show that stored isoeugenol has an antagonistic effect on a mouse olfactory receptor, mOR-EG. However, freshly purified isoeugenol did not have an inhibitory effect. Instead, an isoeugenol derivative produced during storage turned out to be a potent competitive antagonist of mOR-EG. Structural analysis revealed that this derivative is an oxidatively dimerized isoeugenol that naturally occurs by oxidative reaction. The current study indicates that as odorants age, they decompose or react with other odorants, which in turn affects responsiveness of an olfactory receptor(s).

Calcium↗

The incidence of herpes zoster.

BACKGROUND: There are few population-based studies of the natural history and epidemiology of herpes zoster. Although a relatively common cause of morbidity, especially among the elderly, contemporary estimates of herpes zoster incidence are lacking. Herein we describe a population-based investigation of incident and recurrent herpes zoster from 1990 through 1992 in a health maintenance organization. METHODS: The health maintenance organization's automated medical records contain clinical and administrative information about care rendered to patients in ambulatory settings, emergency departments, and hospitals. Cases of herpes zoster were ascertained by screening the medical record for coded diagnoses. The predictive value of a herpes zoster diagnosis code was determined by review of a sample of patient records. Records from all patients with potential recurrences were also reviewed. RESULTS: The overall incidence, based on 1075 cases in 500,408 person-years, was 215 per 100,000 person-years (95% confidence interval, 192 to 240 per 100,000) and did not vary by gender. Although the rate increased sharply with age, approximately 5% of the cases occurred among children younger than 15 years. Infection with human immunodeficiency virus was documented in 5% of the persons with incident herpes zoster and cancer in 6%. Four persons had confirmed recurrences of herpes zoster (744 per 100,000 person-years; 95% confidence interval, 203 to 1907); three of these persons were infected with the human immunodeficiency virus. CONCLUSIONS: The recorded incidence of herpes zoster was 64% higher than that reported 30 years ago; the age-standardized rate was more than twofold higher. Immunosuppressive conditions had little impact on overall incidence, although they were strongly associated with early recurrences.

Adolescent↗

Do studies of wire code and childhood leukemia point towards or away from magnetic fields as the causal agent?

A long-standing point of controversy in the epidemiologic literature concerns the meaning of a wire code-childhood leukemia association for assessing the role of magnetic field exposure. Six studies of wire codes and childhood leukemia in North America were examined, three of which reported positive associations and all of which found some relation between wire codes and measured magnetic fields. Supporting magnetic fields as the basis for the wire code associations are the correspondence between those wire code levels which predict distinct magnetic fields and those which predict leukemia risk in the positive studies. Geographic locations and methods that refine wire codes as magnetic fields predictors also tend to strengthen the association with leukemia. Opposing arguments are based on the failure of the wire code-magnetic field association to predict the strength of association across studies, including the unexplained lack of association between wire codes and leukemia in the Midwest and in Canada. Alternatives to magnetic fields are less supported; residential mobility, social class, and neighborhood characteristics are unlikely to explain a wire code effect. Ambiguity persists because of the modest strength of the wire code-leukemia association, the complexity of the relation between wire codes and magnetic fields, lack of knowledge of risk factors for childhood leukemia, and the limited evaluation of wire code correlates other than magnetic fields.

Canada↗

Gene recognition in cyanobacterium genomic sequence data using the hidden Markov model.

We have developed a hidden Markov model (HMM) to detect the protein coding regions within one megabase contiguous sequence data, registered in a database called GenBank in eight entries, of the genome of cyanobacterium, Synechocystis sp. strain PCC6803. Detection of the coding regions in the database entry was performed by using HMM whose parameters were determined by taking the statistics from the rests of the entries. This HMM has states modeling the di-codons and their frequencies within coding regions and those modeling its base contents in the intergenic regions. Results of the cross-validation showed that the HMM recognized 92.1% of coding regions assigned in sequence annotation. In addition, it suggested 94 potential new coding regions whose length are longer than 90 bases. The recognition accuracy calculated at the level of individual bases was 90.7% for the coding regions and 88.1% for the intergenic regions. This corresponds to a correlation coefficient for coding region recognition of 0.784. Comparison with its prediction accuracy with that by GeneMark showed that the HMM has the same level of prediction accuracy as GeneMark on average. Since we can extend the HMM to utilize information such as SD sequences, the prediction accuracy of the HMM will be enhanced. It was observed that correlation was positive between the prediction rate of the coding regions and the G + C content at the third position of the codon. This suggests the possibility that the prediction rate of coding regions in the cyanobacteria sequence can be enhanced by improving the present HMM into that reflects the classification of coding regions based on the G + C content.

Cyanobacteria↗

Extensive sequence homology of the goldfish ras gene to mammalian ras genes.

We cloned ras-related sequences from goldfish genomic libraries constructed as recombinants using the lambda phage. Restriction enzyme mapping of the clones obtained revealed three kinds of ras-related sequences among approximately 350,000 genomic clones. One of these clones was partially sequenced. Comparison with the nucleotide sequences of mammalian ras genes showed that the determined sequences covered the predicted amino acid coding regions and parts of the intervening regions. The predicted amino acid sequences of the cloned ras-related goldfish gene suggested that the coding region is localized separately in DNA, and that its exon-intron boundaries are exactly the same as those of corresponding mammalian genes. The nucleotide and amino acid sequences of the goldfish ras-related gene may have extensive homologies to mammalian p 21 protein. Among the three mammalian ras proteins, the predicted amino acid sequence of the sequenced ras-related goldfish clone is most closely homologous (96%) to the Kirsten ras protein. Differences in the predicted amino acid sequence were greatest in the sequence predicted from the fourth exon; fewer differences were found in the sequence from the third exon, and only slight or no differences were found in the sequence predicted for the first and second exons. The 12th and 61st amino acids from the N-terminal of the protein, which are thought to be critical positions for GTP binding and catalysis, are both conserved in the goldfish protein.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

GENCODE: producing a reference annotation for ENCODE.

BACKGROUND: The GENCODE consortium was formed to identify and map all protein-coding genes within the ENCODE regions. This was achieved by a combination of initial manual annotation by the HAVANA team, experimental validation by the GENCODE consortium and a refinement of the annotation based on these experimental results. RESULTS: The GENCODE gene features are divided into eight different categories of which only the first two (known and novel coding sequence) are confidently predicted to be protein-coding genes. 5' rapid amplification of cDNA ends (RACE) and RT-PCR were used to experimentally verify the initial annotation. Of the 420 coding loci tested, 229 RACE products have been sequenced. They supported 5' extensions of 30 loci and new splice variants in 50 loci. In addition, 46 loci without evidence for a coding sequence were validated, consisting of 31 novel and 15 putative transcripts. We assessed the comprehensiveness of the GENCODE annotation by attempting to validate all the predicted exon boundaries outside the GENCODE annotation. Out of 1,215 tested in a subset of the ENCODE regions, 14 novel exon pairs were validated, only two of them in intergenic regions. CONCLUSION: In total, 487 loci, of which 434 are coding, have been annotated as part of the GENCODE reference set available from the UCSC browser. Comparison of GENCODE annotation with RefSeq and ENSEMBL show only 40% of GENCODE exons are contained within the two sets, which is a reflection of the high number of alternative splice forms with unique exons annotated. Over 50% of coding loci have been experimentally verified by 5' RACE for EGASP and the GENCODE collaboration is continuing to refine its annotation of 1% human genome with the aid of experimental validation.

Chromosome Mapping↗

Identification of coding regions in genomic DNA sequences: an application of dynamic programming and neural networks.

Dynamic programming (DP) is applied to the problem of precisely identifying internal exons and introns in genomic DNA sequences. The program GeneParser first scores the sequence of interest for splice sites and for these intron- and exon-specific content measures: codon usage, local compositional complexity, 6-tuple frequency, length distribution and periodic asymmetry. This information is then organized for interpretation by DP. GeneParser employs the DP algorithm to enforce the constraints that introns and exons must be adjacent and non-overlapping and finds the highest scoring combination of introns and exons subject to these constraints. Weights for the various classification procedures are determined by training a simple feed-forward neural network to maximize the number of correct predictions. In a pilot study, the system has been trained on a set of 56 human gene fragments containing 150 internal exons in a total of 158,691 bps of genomic sequence. When tested against the training data, GeneParser precisely identifies 75% of the exons and correctly predicts 86% of coding nucleotides as coding while only 13% of non-exon bps were predicted to be coding. This corresponds to a correlation coefficient for exon prediction of 0.85. Because of the simplicity of the network weighting scheme, generalization performance is nearly as good as with the training set.

Algorithms↗

Gene structure prediction using information on homologous protein sequence.

In this paper a new approach for the prediction of protein coding gene structures is described. The principal scheme of prediction is as follows: first, the exons with the best potential are predicted in a sequence with unknown functions and a list of potential amino acid fragments coded by these exons is formed. Second, testing the homology between each amino acid fragment from the list and proteins from the SWISS-PROT database of amino acid sequences. One protein with the best homology is chosen out of all the homologous sequences. Third, reconstruction of the exon-intron structure, basing it on its homology with the chosen protein sequences. The method was tested on an independent control set (20 genes). The results were as follows: 21% of real exons were lost and 3% of non-real exons were found. This system can be used to refine the results of gene prediction systems, especially if highly homologous proteins are found in the amino acid sequence database.

Algorithms↗

Comparative ab initio prediction of gene structures using pair HMMs.

We present a novel comparative method for the ab initio prediction of protein coding genes in eukaryotic genomes. The method simultaneously predicts the gene structures of two un-annotated input DNA sequences which are homologous to each other and retrieves the subsequences which are conserved between the two DNA sequences. It is capable of predicting partial, complete and multiple genes and can align pairs of genes which differ by events of exon-fusion or exon-splitting. The method employs a probabilistic pair hidden Markov model. We generate annotations using our model with two different algorithms: the Viterbi algorithm in its linear memory implementation and a new heuristic algorithm, called the stepping stone, for which both memory and time requirements scale linearly with the sequence length. We have implemented the model in a computer program called DOUBLESCAN. In this article, we introduce the method and confirm the validity of the approach on a test set of 80 pairs of orthologous DNA sequences from mouse and human. More information can be found at: http://www.sanger.ac.uk/Software/analysis/doublescan/

Algorithms↗

Separate brain regions code for salience vs. valence during reward prediction in humans.

Predicting rewards and avoiding aversive conditions is essential for survival. Recent studies using computational models of reward prediction implicate the ventral striatum in appetitive rewards. Whether the same system mediates an organism's response to aversive conditions is unclear. We examined the question using fMRI blood oxygen level-dependent measurements while healthy volunteers were conditioned using appetitive and aversive stimuli. The temporal difference learning algorithm was used to estimate reward prediction error. Activations in the ventral striatum were robustly correlated with prediction error, regardless of the valence of the stimuli, suggesting that the ventral striatum processes salience prediction error. In contrast, the orbitofrontal cortex and anterior insula coded for the differential valence of appetitive/aversive stimuli. Given its location at the interface of limbic and motor regions, the ventral striatum may be critical in learning about motivationally salient stimuli, regardless of valence, and using that information to bias selection of actions.

Adult↗

Cloning and sequencing the HinfI restriction and modification genes.

The HinfI restriction and modification genes were cloned on a 3.9-kb PstI fragment inserted into the PstI site of plasmid pBR322. Both genes are confined to an internal 2.3-kb BclI-AvaI subfragment. This subfragment was sequenced. Two large open reading frames (ORF's) are present. ORF1 codes for the methylase [predicted 359 amino acids (aa)] and ORF2 codes for the endonuclease (predicted 262 or 272 aa).

Amino Acid Sequence↗

Identification of human gene structure using linear discriminant functions and dynamic programming.

Development of advanced technique to identify gene structure is one of the main challenges of the Human Genome Project. Discriminant analysis was applied to the construction of recognition functions for various components of gene structure. Linear discriminant functions for splice sites, 5'-coding, internal exon, and 3'-coding region recognition have been developed. A gene structure prediction system FGENE has been developed based on the exon recognition functions. We compute a graph of mutual compatibility of different exons and present a gene structure models as paths of this directed acyclic graph. For an optimal model selection we apply a variant of dynamic programming algorithm to search for the path in the graph with the maximal value of the corresponding discriminant functions. Prediction by FGENE for 185 complete human gene sequences has 81% exact exon recognition accuracy and 91% accuracy at the level of individual exon nucleotides with the correlation coefficient (C) equals 0.90. Testing FGENE on 35 genes not used in the development of discriminant functions shows 71% accuracy of exact exon prediction and 89% at the nucleotide level (C = 0.86). FGENE compares very favorably with the other programs currently used to predict protein-coding regions. Analysis of uncharacterized human sequences based on our methods for splice site (HSPL, RNASPL), internal exons (HEXON), all type of exons (FEXH) and human (FGENEH) and bacterial (CDSB) gene structure prediction and recognition of human and bacterial sequences (HBR) (to test a library for E. coli contamination) is available through the University of Houston, Weizmann Institute of Science network server and a WWW page of the Human Genome Center at Baylor College of Medicine.

Algorithms↗

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny↗