Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

The MMPI as a predictor of response to conservative treatment for low back pain.

Studies that used the MMPI to predict the response of chronic low back pain patients to standard medical treatment have not produced definitive results. Patients seen in a university hospital orthopedic back pain clinic were given the MMPI before treatment, and 6 to 12 months later 76 patients completed follow-up forms that indicated their level of intensity during the previous week and their ratings of the success of treatment in relieving their pain as well as in enabling them to return to normal activities. Predictions of poor response were made in terms of either single MMPI scales or code types. Patients with poor outcome on two of the three criteria (level of pain intensity and ability to return to normal activities) had significantly higher scores on the Hs scale. The predicted high risk code types very accurately identified patients with poor response on the same two criteria; however, the code-type procedure overpredicted poor response in the good outcome group.

Back Pain↗

PCR differential display identifies a rat brain mRNA that is transcriptionally regulated by cocaine and amphetamine.

Neuronal plasticity associated with both short- and long-term administration of psychomotor stimulants involves alterations in specific patterns of gene expression. In order to screen for brain region specific mRNAs which are transcriptionally regulated by acute cocaine and amphetamine, PCR differential display was employed. This approach identified a previously uncharacterized mRNA whose relative levels in the striatum are induced four- to fivefold by acute psychomotor stimulant administration. Isolation and characterization of corresponding cDNA clones resulted in complete nucleotide sequence analysis, including prediction of the encoded protein product. Alternate polyA site utilization in the predicted 3' noncoding region results in the appearance of an RNA doublet, approximately 700 and 900 bases in length, following Northern analysis. A presumed alternate splicing event further generates diversity within the transcripts, and results in the presence or absence of an in-frame 39 base insert within the putative protein coding region. As a result, the predicted translation products are either 129 or 116 amino acids in length. A common hydrophobic leader sequence at the amino terminus is present within each predicted polypeptide, suggesting that the protein product is targeted for entry into the secretory pathway. Basal expression of the RNA doublet is limited to neuroendocrine tissues, further implying that the protein product plays a functional role in both neuronal and endocrine tissues.

Amino Acid Sequence↗

Photodosimetry of interstitial light delivery to solid tumors.

Both anaplastic and well-differentiated Dunning prostate adenocarcinomas were illuminated in anesthetized Fischer X Copenhagen rats by single-fiber and multiple-fiber illuminators. Each illuminator consisted of a 2-cm laterally diffusing optical fiber placed within a plastic brachytherapy needle which was implanted into a tumor. Light attenuation coefficients for various wavelengths were obtained from measures of the radial falloff of intensity with distance from single fibers. These coefficients served as input to a 2-dimensional (2-D) photodosimetry computer code which calculated relative light intensities in planes perpendicular to single-fiber and various multiple-fiber configurations. These calculations assumed uniform optical property of tissue throughout each tumor, uniform and equal illuminance from diffusing fibers, and precise needle implantation. Relative light intensities along specific tumor tracks were measured and compared with those predicted by the 2-D photodosimetry code. Agreement within +/- 14% was observed for all configurations studied. Variations in relative intensity in tumor planes perpendicular to a standard seven-fiber illuminator were determined as a function of the distance between the implanted needles. Light wavelengths of 700 nm and greater produced relatively uniform light fields (approximately +/- 20%) in the R3327-AT tumor with needle spacings of at least 1 cm. The addition of two fibers at the periphery of this illuminator (a nine-fiber illuminator) improved the uniformity of light delivery to the encompassed tumor volumes. The importance of precision photodosimetry for interstitial applications of photodynamic therapy is discussed.

Adenocarcinoma↗

Investigating biological response in the UVB as a function of ozone variation using perturbation theory.

In order to determine a biological response to ultraviolet radiation, calculations of biologically weighted dose rates are required, which in turn involve the integral over wavelength of an action spectrum multiplied by appropriate surface flux data. To determine a biologically weighted dose rate accurately, a reasonable wavelength resolution is required, involving a full radiative transfer solution to be performed for each wavelength in order to obtain the surface flux information. If biologically weighted dose rates are needed as a function of ozone variation, then the number of radiative transfer solutions quickly makes a large number of ozone variations cumbersome. This paper shows that the perturbation theory developed for atmospheric radiative transfer by Box and co-workers can predict surface fluxes and hence biologically weighted dose rates for a large range of ozone variations very efficiently. The method is then extended to calculate radiation amplification factors. Results for biologically weighted dose rates are presented for a large range of solar zenith angles and ozone loadings using perturbation theory and a full radiative transfer code and show that the perturbation predictions never deviate very far from the radiative transfer solutions.

Mathematical Computing↗

Lossless medical image compression by multilevel decomposition.

Lossless image coding is important for medical image compression because any information loss or error caused by the image compression process could affect clinical diagnostic decisions. This paper proposes a lossless compression algorithm for application to medical images that have high spatial correlation. The proposed image compression algorithm uses a multi-level decomposition scheme in conjunction with prediction and classification. In this algorithm, an image is divided into four subimages by subsampling. One subimage is used as a reference to predict the other three subimages. The prediction errors of the three subimages are classified into two or three groups by the characteristics of the reference subimage, and the classified prediction errors are encoded by entropy coding with corresponding code words. These subsampling and classified entropy coding procedures are repeated on the reference subimage in each level, and the reference subimage in the last repetition is encoded by conventional differential pulse code modulation and entropy coding. To verify this proposed algorithm, it was applied to several chest radiographs and computed tomography and magnetic resonance images, and the results were compared with those from well-known lossless compression algorithms.

Algorithms↗

A third approach to gene prediction suggests thousands of additional human transcribed regions.

The identification and characterization of the complete ensemble of genes is a main goal of deciphering the digital information stored in the human genome. Many algorithms for computational gene prediction have been described, ultimately derived from two basic concepts: (1) modeling gene structure and (2) recognizing sequence similarity. Successful hybrid methods combining these two concepts have also been developed. We present a third orthogonal approach to gene prediction, based on detecting the genomic signatures of transcription, accumulated over evolutionary time. We discuss four algorithms based on this third concept: Greens and CHOWDER, which quantify mutational strand biases caused by transcription-coupled DNA repair, and ROAST and PASTA, which are based on strand-specific selection against polyadenylation signals. We combined these algorithms into an integrated method called FEAST, which we used to predict the location and orientation of thousands of putative transcription units not overlapping known genes. Many of the newly predicted transcriptional units do not appear to code for proteins. The new algorithms are particularly apt at detecting genes with long introns and lacking sequence conservation. They therefore complement existing gene prediction methods and will help identify functional transcripts within many apparent "genomic deserts."

Algorithms↗

Investigating the accuracy of the FLUKA code for transport of therapeutic ion beams in matter.

In-beam positron emission tomography (PET) is currently used for monitoring the dose delivery at the heavy ion therapy facility at GSI Darmstadt. The method is based on the fact that carbon ions produce positron emitting isotopes in fragmentation reactions with the atomic nuclei of the tissue. The relation between dose and beta(+)-activity is not straightforward. Hence it is not possible to infer the delivered dose directly from the PET distribution. To overcome this problem and enable therapy monitoring, beta(+)-distributions are simulated on the basis of the treatment plan and compared with the measured ones. Following the positive clinical impact, it is planned to apply the method at future ion therapy facilities, where beams from protons up to oxygen nuclei will be available. A simulation code capable of handling all these ions and predicting the irradiation-induced beta(+)-activity distributions is desirable. An established and general purpose radiation transport code is preferred. FLUKA is a candidate for such a code. For application to in-beam PET therapy monitoring, the code has to model with high accuracy both the electromagnetic and nuclear interactions responsible for dose deposition and beta(+)-activity production, respectively. In this work, the electromagnetic interaction in FLUKA was adjusted to reproduce the same particle range as from the experimentally validated treatment planning software TRiP, used at GSI. Furthermore, projectile fragmentation spectra in water targets have been studied in comparison to available experimental data. Finally, cross sections for the production of the most abundant fragments have been calculated and compared to values found in the literature.

Carbon↗

Functional associations of proteins in entire genomes by means of exhaustive detection of gene fusions.

BACKGROUND: It has recently been shown that the detection of gene fusion events across genomes can be used for predicting functional associations of proteins, including physical interaction or complex formation. To obtain such predictions we have made an exhaustive search for gene fusion events within 24 available completely sequenced genomes. RESULTS: Each genome was used as a query against the remaining 23 complete genomes to detect gene fusion events. Using an improved, fully automatic protocol, a total of 7,224 single-domain proteins that are components of gene fusions in other genomes were detected, many of which were identified for the first time. The total number of predicted pairwise functional associations is 39,730 for all genomes. Component pairs were identified by virtue of their similarity to 2,365 multidomain composite proteins. We also show for the first time that gene fusion is a complex evolutionary process with a number of contributory factors, including paralogy, genome size and phylogenetic distance. On average, 9% of genes in a given genome appear to code for single-domain, component proteins predicted to be functionally associated. These proteins are detected by an additional 4% of genes that code for fused, composite proteins. CONCLUSIONS: These results provide an exhaustive set of functionally associated genes and also delineate the power of fusion analysis for the prediction of protein interactions.

Algorithms↗

Molecular modelling of the Norrie disease protein predicts a cystine knot growth factor tertiary structure.

The X-lined gene for Norrie disease, which is characterized by blindness, deafness and mental retardation has been cloned recently. This gene has been thought to code for a putative extracellular factor; its predicted amino acid sequence is homologous to the C-terminal domain of diverse extracellular proteins. Sequence pattern searches and three-dimensional modelling now suggest that the Norrie disease protein (NDP) has a tertiary structure similar to that of transforming growth factor beta (TGF beta). Our model identifies NDP as a member of an emerging family of growth factors containing a cystine knot motif, with direct implications for the physiological role of NDP. The model also sheds light on sequence related domains such as the C-terminal domain of mucins and of von Willebrand factor.

Amino Acid Sequence↗

A novel rat dentin mRNA coding only for dentin sialoprotein.

Dentin sialoprotein (DSP) is a major glycoprotein present in the mineralized dentin matrix that is expressed mainly by young and mature odontoblasts. Mutations in the DSP coding regions are linked to Dentinogenesis imperfecta I and II. indicating the importance of DSP in tooth formation. Previous studies have identified multiple mRNA transcripts in dentin that code for both DSP and phosphophoryns (PPs). Using reverse transcriptase-polymerase chain reaction (RT-PCR) to characterize these mRNA transcripts, we have identified a cDNA that codes for DSP, but not PP. This cDNA codes for a protein with 324 amino acids, 303 amino acids being identical to the published rat DSP sequence. However, the subsequent 21 amino acids are unique to this cDNA. Based on the coding sequence, the core protein is predicted to have a pI=4.24, a net charge of -34, and to contain four potential N-glycosylation sites and six potential sites for phosphorylation by casein kinase. That the corresponding mRNA was present in day 5 molar tooth germs was confirmed using RNA protection assays. These data, therefore, identify a novel transcript in rat tooth germs that codes only for DSP (designated as DSPII).

Age Factors↗

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames↗

Nucleotide sequence, evolution, and expression of the fetal globin gene of the spider monkey Ateles geoffroyi.

The single gamma-globin gene of the New World spider monkey Ateles geoffroyi is similar to other primate genes of the beta-globin gene cluster ("beta-like" globin genes). The number of nonsynonymous nucleotide substitutions between the coding regions of Ateles and other primate gamma-globin genes suggests that the Platyrrhine and Catarrhine evolutionary lines diverged approximately equal to 40 million years ago, an estimate reasonably consistent with the fossil record. However, the number of synonymous coding region and noncoding base differences is much smaller than predicted by various molecular "clocks." This suggests that the rate of synonymous coding and noncoding base substitution has not been constant per absolute time in primate lineages. Expression of the cloned Ateles gamma-globin gene in cultured monkey cells showed that the sequence AAUAAA near the mRNA 3' terminus is not sufficient to define the site of transcript polyadenylylation.

Amino Acid Sequence↗

Prediction of contact maps with neural networks and correlated mutations.

Contact maps of proteins are predicted with neural network-based methods, using as input codings of increasing complexity including evolutionary information, sequence conservation, correlated mutations and predicted secondary structures. Neural networks are trained on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well resolved three-dimensional structures. Proteins are selected from the Protein Data Bank database provided that they align with at least 15 similar sequences in the corresponding families. The predictors are trained to learn the association rules between the covalent structure of each protein and its contact map with a standard back propagation algorithm and tested on the same protein set with a cross-validation procedure. Our results indicate that the method can assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor >6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Furthermore, filtering the network outputs with a procedure based on the residue coordination numbers, the accuracy of predictions increases up to 0.25 for all the proteins, with an 8-fold deviation from a random predictor. These scores are the highest reported so far for predicting protein contact maps.

Algorithms↗

Dissociable systems for gain- and loss-related value predictions and errors of prediction in the human brain.

Midbrain dopaminergic neurons projecting to the ventral striatum code for reward magnitude and probability during reward anticipation and then indicate the difference between actual and predicted outcome. It has been questioned whether such a common system for the prediction and evaluation of reward exists in humans. Using functional magnetic resonance imaging and a guessing task in two large cohorts, we are able to confirm ventral striatal responses coding both reward probability and magnitude during anticipation, permitting the local computation of expected value (EV). However, the ventral striatum only represented the gain-related part of EV (EV+). At reward delivery, the same area shows a reward probability and magnitude-dependent prediction error signal, best modeled as the difference between actual outcome and EV+. In contrast, loss-related expected value (EV-) and the associated prediction error was represented in the amygdala. Thus, the ventral striatum and the amygdala distinctively process the value of a prediction and subsequently compute a prediction error for gains and losses, respectively. Therefore, a homeostatic balance of both systems might be important for generating adequate expectations under uncertainty. Prevalence of either part might render expectations more positive or negative, which could contribute to the pathophysiology of mood disorders like major depression.

Adult↗

DIANA-EST: a statistical analysis.

MOTIVATION: Expressed Sequence Tags (ESTs) are next to cDNA sequences as the most direct way to locate in silico the genes of the genome and determine their structure. Currently ESTs make up more than 60% of all the database entries. The goal of this work is the development of a new program called DNA Intelligent Analysis for ESTs (DIANA-EST) based on a combination of Artificial Neural Networks (ANN) and statistics for the characterization of the coding regions within ESTs and the reconstruction of the encoded protein. RESULTS: 89.7% of the nucleotides from an independent test set with 127 ESTs were predicted correctly as to whether they are coding or non coding. AVAILABILITY: The program is available upon request from the author. CONTACT: Present address: Department of Genetics, University of Pennsylvania, School of Medicine, 475 Clinical Research Building, 415 Curie Boulevard, Philadelphia, PA 19104-6145, USA. artemis@pcbi.upenn.edu.

Computational Biology↗

Cloning and characterization of a human orphan family C G-protein coupled receptor GPRC5D.

Recently three orphan G-protein coupled receptors, RAIG1, GPRC5B and GPRC5C, with homology to members of family C (metabotropic glutamate receptor-like) have been identified. Using the protein sequences of these receptors as queries we identified overlapping expressed sequence tags which were predicted to encode an additional subtype. The full length coding regions of mouse mGprc5d and human GPRC5D were cloned and shown to contain predicted open reading frames of 300 and 345 amino acids, respectively. GPRC5D has seven putative transmembrane segments and is expressed in the cell membrane. The four human receptor subtypes, which we assign to group 5 of family C GPCRs, show 31-42% amino acid sequence identity to each other and 20-25% sequence identity to the transmembrane domains of metabotropic glutamate receptor subtypes 2 and 3 and other family C members. In contrast to the remaining family C members, the group 5 receptors have short amino terminal domains of some 30-50 amino acids. GPRC5D was shown to be clustered with RAIG1 on chromosome 12p13.3 and like RAIG1 and GPRC5B to consist of three exons, the first exon being the largest containing all seven transmembrane segments. GPRC5D mRNA is widely expressed in the peripheral system but all four receptors show distinct expression patterns. Interestingly, mRNA levels of all four group 5 receptors were found in medium to high levels in the kidney, pancreas and prostate and in low to medium levels in the colon and the small intestine, whereas other organs only express a subset of the genes. In an attempt to delineate the signal transduction pathway(s) of the orphan receptors, a series of chimeric receptors containing the amino terminal domain of the calcium sensing receptor or metabotropic glutamate receptor subtype 1, and the seven transmembrane domain of the orphan receptors were constructed and tested in binding and functional assays.

Amino Acid Sequence↗

[Application of artificial neural network based on the genetic algorithm in predicting the root distribution of winter wheat].

In this study, a controlled experiment of winter wheat under water stress at the seedling stage was conducted in soil columns in greenhouse. Based on the data gotten from the experiment, a model to estimate root length density distribution was developed through optimizing the weights of neural network by genetic algorithm. The neural network model was constructed by using forward neural network framework, by applying the strategy of the roulette wheel selection and reserving the most optimizing series of weights, which were composed by real codes. This model was applied to predict the root length density distribution of winter wheat, and the predicted root length density had good agreement with experiment data. The way could save a lot of manpower and material resources for determining the root length density distribution of winter wheat.

Algorithms↗

Nucleotide sequence of the herpes simplex virus type 2 (HSV-2) thymidine kinase gene and predicted amino acid sequence of thymidine kinase polypeptide and its comparison with the HSV-1 thymidine kinase gene.

To analyze the boundaries of the functional coding region of the HSV-2(333) thymidine kinase gene (TK gene), deletion mutants of hybrid plasmid pMAR401 H2G, which contains the 17.5 kbp BglII-G fragment of HSV-2 DNA, were prepared and tested for capacity to transform LM(TK-) cells to the thymidine kinase-positive phenotype. These studies showed that hybrid plasmids containing 2.2-2.4 kbp subfragments of HSV-2 BglII-G DNA transformed LM(TK-) cells to the thymidine kinase-positive phenotype and suggested that the region critical for transformation might be less than 2 kbp. That the activity expressed in the transformants was HSV-2 thymidine kinase was shown by experiments with type-specific enzyme-inhibiting rabbit antisera and by disc-polyacrylamide gel electrophoresis analyses. DNA fragments of the HSV-2 TK gene were subcloned in phage M13mp9 and M13mp8. A sequence of 1656 bp containing the entire coding region of the TK gene and the flanking sequences was determined by the dideoxynucleotide chain termination method. Comparisons with the HSV-1(Cl 101) TK gene revealed that PstI, PvuII, and EcoRI cleavage sites had homologous locations as did promoter, translational start and stop, and polyadenylation signals. Extensive homology was observed in the nucleotide sequence preceding the ATG translational start signal and in portions of the coding region of the genes. Comparisons of the predicted amino acid sequences of the HSV-1 and HSV-2 thymidine kinase polypeptides revealed that both were enriched in alanine, arginine, glycine, leucine, and proline residues and that clear, but interrupted homology existed within several regions of the polypeptide chains. Stretches of 15-30 amino acid residues were identical in conserved regions. The possibility is suggested that domains containing some of the conserved amino acid sequences might have a role in substrate binding and as major antigenic determinants.

Amino Acid Sequence↗