Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

The murine complement receptor gene family. IV. Alternative splicing of Cr2 gene transcripts predicts two distinct gene products that share homologous domains with both human CR2 and CR1.

The murine Cr2 gene encodes at least two related proteins. The first of these is predicted to include 1408 amino acids from a transcript including 4224 coding nucleotides. This protein is predicted to contain 21 60-amino acid repeats plus those residues encoding transmembrane and cytoplasmic regions for a total peptide m.w. of 155,307. The first six of these repeats are similar to human CR1 in sequence and organization. The second protein is encoded by an alternatively spliced Cr2 transcript that is lacking those sequences which encode the first six 60-amino acid repeats of the larger protein. This second, smaller protein, encoded by a transcript of 3096 coding nucleotides, is predicted to include 15 60-amino acid repeats, plus the transmembrane and cytoplasmic regions. This smaller protein contains 1032 amino acids for a peptide m.w. of 113,328. This second protein is extremely homologous in size and sequence to human CR2. Both proteins share the same signal sequence for membrane insertion. DNA sequence analysis, RNA protection studies and genomic phage mapping indicate the transcripts which encode these proteins are derived from the Cr2 gene via alternative splicing.

Amino Acid Sequence↗

Report on antibodies submitted to the stromal cell section of HLDA8.

The paradigm for tissue specific homing of leukocytes is the "area code" hypothesis, which predicts that a specific combination of adhesive interactions and chemokine signals from the endothelium directs leukocyte migration into specific tissue sites. This area code hypothesis has been supported by studies from previous HLDA workshops where endothelial specific cell antigens have been studied. Similarly, a clear haematopoietic "stem cell code" comprising the chemokine SDF-1 (CXCL12) and the adhesion receptor VCAM-1 (CD106) has been shown to contribute to the stem cell niche within bone marrow [K. Tokoyoda, T. Egawa, T. Sugiyama, B.I. Chai, T. Nagasawa, Cellular niches controlling B lymphocyte behaviour within bone marrow during development, Immunity 20 (2004) 707-718]. HLDA 7 included a section devoted to stem cell antigens, which began to define additional antigens important in these processes. During the course of HLDA 8 we have extended these observations to determine whether a more global stromal address code defined by fibroblasts, exists in variety of different tissues [G. Parsonage, A.D. Filer, O. Haworth, G.B. Nash, G.E. Rainger, M. Salmon, C.D. Buckley, A stromal area postcode defined by fibroblasts, Trends Immunol. 26 (2005) 150-156]. The stromal cell section in HLDA 8 was designed to complement the malignant cell, endothelial cell, and stem cell/progenitor cell sections. Seven new CD numbers were assigned to antibodies included in this section at the HLDA 8 Workshop meeting held during December 2004.

Antibodies, Monoclonal↗

Genetic codes as evolutionary filters: subtle differences in the structure of genetic codes result in significant differences in patterns of nucleotide substitution.

The codon-degeneracy model (CDM) predicts that patterns of nucleotide substitution in protein-coding genes are largely determined by the relative frequencies of four-fold (4f), two-fold, and non-degenerate sites, the attributes of which are determined by the structure of the governing genetic code. The CDM thus further predicts that genetic codes with alternative structures will "filter" molecular evolution differentially. A method, therefore, is presented by which the CDM may be applied to the unique structure of any genetic code. The mathematical relationship between the proportion of transitions at 4f degenerate nucleotide sites and the transition-to-transversion ratio is described. Predictions for five individual genetic codes, relative to the relationship between code structure and expected patterns of nucleotide substitution, are clearly defined. To test this "filter" hypothesis of genetic codes, simulated DNA sequence data sets were generated with a variety of input parameter values to estimate the relationship between patterns of nucleotide substitution and best-fit estimates of transition bias at 4f degenerate sites for both the universal genetic code and the vertebrate mitochondrial genetic code. These analyses confirm the prediction of the CDM that, all else being equal, even small differences in the structure of alternative genetic codes may result in significant shifts in the overall pattern of nucleotide substitution.

Codon↗

On error minimization in a sequential origin of the standard genetic code.

Distances between amino acids were derived from the polar requirement measure of amino acid polarity and Benner and co-workers' (1994) 74-100 PAM matrix. These distances were used to examine the average effects of amino acid substitutions due to single-base errors in the standard genetic code and equally degenerate randomized variants of the standard code. Second-position transitions conserved all distances on average, an order of magnitude more than did second-position transversions. In contrast, first-position transitions and transversions were about equally conservative. In comparison with randomized codes, second-position transitions in the standard code significantly conserved mean square differences in polar requirement and mean Benner matrix-based distances, but mean absolute value differences in polar requirement were not significantly conserved. The discrepancy suggests that these commonly used distance measures may be insufficient for strict hypothesis testing without more information. The translational consequences of single-base errors were then examined in different codon contexts, and similarities between these contexts explored with a hierarchical cluster analysis. In one cluster of codon contexts corresponding to the RNY and GNR codons, second-position transversions between C and G and transitions between C and U were most conservative of both polar requirement and the matrix-based distance. In another cluster of codon contexts, second-position transitions between A and G were most conservative. Despite the claims of previous authors to the contrary, it is shown theoretically that the standard code may have been shaped by position-invariant forces such as mutation and base content. These forces may have left heterogeneous signatures in the code because of differences in translational fidelity by codon position. A scenario for the origin of the code is presented wherein selection for error minimization could have occurred multiple times in disjoint parts of the code through a phyletic process of competition between lineages. This process permits error minimization without the disruption of previously useful messages, and does not predict that the code is optimally error-minimizing with respect to modern error. Instead, the code may be a record of genetic process and patterns of mutation before the radiation of modern organisms and organelles.

Amino Acids↗

Simulation of CRISPR/Cas9-mediated gene editing for the Vitellogenin gene in Apis mellifera.

CRISPR/Cas9 genome editing provides a powerful framework for interrogating gene function in Apis mellifera. Yet, empirical application remains challenging due to biological constraints, including haplodiploid genetics, narrow embryonic injection window, and the social rearing requirements that complicate functional validation. These constraints necessitate in silico pre-screening to maximize editing success before resource-intensive wet-lab implementation. Within the omnigenic framework, which distinguishes core regulatory genes from peripheral loci buffered by network effects, vitellogenin (Vg) represents an optimal target which is ancestrally dedicated to yolk provisioning; it has been co-opted to orchestrate diverse non-reproductive functions including longevity, stress resistance, immunity, and social behavior. We developed a computational pipeline to design a list of 57 and 56 candidate guide RNAs (gRNA) for targeted Vg knockout, evaluating candidate sites in both functional exons 2 and 3 based on structural accessibility and frameshift efficiency. Comparative analysis revealed complementary strengths in two top-best candidates from initial target pool of predicted gRNAs. The gRNA targeting exon 2 exhibits weaker secondary structure (ΔG = -0.25 kcal/mol versus -2.10 kcal/mol for exon 3), aligning with empirical evidence that sites with ΔG > -1.0 kcal/mol achieve 2-5 × higher Cas9 binding efficiency. This site yielded moderate frameshift frequency (77.8%; 61.9 percentile). Conversely, the predicted editing outcome for the gRNA targeting exon 3, despite stronger structural constraints, demonstrated superior functional disruption metrics demonstrating very high frameshift frequency (88.3%; 95.2 percentile), high in silico editing precision, minimal microhomology-mediated repair bias, and reproducible outcomes wherein nearly all predicted indels disrupt the coding sequence. Protein structure and domain analyses further predict that frameshift edits will generate a truncated protein missing all downstream functional domains. We recommend parallel empirical validation of both exon 2 and exon 3 targets to resolve the trade-off between structural accessibility (favoring higher editing rates) and frameshift efficacy (favoring complete loss-of-function). This dual-target strategy accommodates uncertainty in in vivo performance while maximizing the probability of generating informative phenotypes. Our in silico framework enables rational CRISPR design in non-model organisms by computationally balancing biophysical accessibility with functional impact, accelerating functional genomics in species where empirical optimization faces substantial biological constraints.

Animals↗

Real-time distortionless high-factor compression scheme.

Nowadays, digital subtraction angiography systems must be able to sustain real-time acquisition (30 frames per second) of 512 x 512 x 8 bit images and store several sequences of such images on low cost and general-purpose mass memories. Concretely, that means a 7.8 Mbytes per second rate and about 780 Mbytes disk space to hold a 100-s cardiac examination. To fulfill these requirements at competitive cost, a distortionless compressor/decompressor system can be designed: during acquisition, the real-time compressor transforms the input images into a lower quantity of coded information through a predictive coder and a variable-length Huffman code. The process is fully reversible because during review, the real-time decompressor exactly recovers the acquired images from the stored compressed data. Test results on many raw images demonstrate that real-time compression is feasible and takes place with absolutely no loss of information. The designed system indifferently works on 512 or 1024 formats, and 256 or 1024 gray levels.

Angiography, Digital Subtraction↗

Injury severity and probability of survival assessment in trauma patients using a predictive hierarchical network model derived from ICD-9 codes.

UNLABELLED: Accurate assessment of injury severity is critical for decision making related to the prevention, triage, and treatment of injured patients. Presently, the standard method of controlling for variations of injury severity between groups has been based upon the Injury Severity Score (ISS) and the Trauma Score and the Trauma and Injury Severity Score (TRISS) methodology. The purpose of this study was to attempt to build upon previous work using International Classification of Diseases, ninth revision (ICD-9) coded diagnosis, and procedure information available from standard hospital discharge abstracts (UB-82 Billing format) to create a hierarchical network to provide a tool for predicting injury severity and probability of survival. METHODS: Data were obtained for this analysis from the North Carolina Medical Database. Data were available on all trauma patients admitted to hospitals in North Carolina from January 1, 1988 until June 30, 1992. The dependent variable of interest was the patient's survival after injury, coded as live or die. The independent variables used in the study included the ISS derived using the technique described by MacKenzie Abbreviated Injury Score (AIS) and body system maximum AIS scores, mortality risk ratios derived from the ICD-9-DM primary, secondary, and tertiary diagnoses, primary and secondary procedures as described in previous work, age and gender. Network generation used a commercial software package, AIM (Abtech Corp., Charlottesville, Va.), which is a numeric modeling tool that automatically "learns" knowledge from a data base of examples. RESULTS: In the test data set an ISS and a prediction of survival based upon the derived network were calculated for each and every patient. The relative predictive power of these two scores were compared by calculating the overall accuracy, sensitivity, and specificity and the false positive and false negative rates. The receiver operator characteristic curves demonstrate that the network is a more effective tool in predicting the outcome of trauma patients. All the measures of predictive power show that the network was the better predictor of outcome than the ISS. CONCLUSIONS: Given the recognized limitations of the ISS, the widespread availability of the ICD-9 coded diagnoses and procedures, and the availability of many state and regional data bases that have no ISS or Trauma Score, the purpose of this study was to assess the ability of a network derived from limited but widely available hospital discharge data to predict the outcome of injured patients. The study confirms previous work showing that the ICD-9 codes were strongly associated with outcome. The study demonstrated that the network created from these data was a better predictor of outcome than the derived ISS. When the results of the network were compared with other published series, the network, created without access to physiologic information, was almost as accurate, sensitive, and specific as reported values for TRISS and A Severity Characterization of Trauma (ASCOT). Because the present study is the first of its type, further investigations are needed to validate these findings. If other studies corroborate this study, a network model based upon ICD-9 codes could become the principal method for grading injury severity. This would provide superior predictive power of injury severity with important cost savings and universal application.

Adult↗

Identification of two subunit A isoforms of the vacuolar H(+)-ATPase in human osteoclastoma.

Subunit A is thought to be the main component of the catalytic site of the vacuolar-type H(+)-ATPase. Screening of a cDNA library made from human osteoclastoma tumor tissue revealed the presence of two isoforms of subunit A. HO68 is a cDNA of 3.1 kilobase pairs, corresponding to a mRNA of approximately 3.4 kilobases in osteoclastoma only, encoding a protein of 615 amino acids with a predicted molecular mass of 68177 Da. A second subtype, VA68, corresponding to a mRNA of approximately 4.8 kilobases was present in all tissues analyzed, and codes for a predicted protein of 617 residues and theoretical molecular mass of 68264 Da. These clones share homology with previously published subunit A sequences, and this, together with the tissue distribution of the mRNA, suggests there are ubiquitous (VA68-type) and tissue-specific (HO68-type) isoforms. HO68 shows the closest sequence homology (95% at the amino acid level) to subunit A of a proton-secreting vacuolar-type H(+)-ATPase located in the apical membrane of midgut goblet cells of tobacco hornworm larva (Manduca sexta). We propose that HO68 could correspond to an isoform of subunit A specific for a vacuolar-type H(+)-ATPase located in the osteoclast plasma membrane.

Adenosine Triphosphatases↗

The utility of Medicare claims data for measuring cancer stage.

BACKGROUND: The validity of using claims data for measuring tumor stage, one of the most important determinants of choice of therapy and long-term survival, is unknown. OBJECTIVES: To determine the relative accuracy of both inpatient and hospital Outpatient Medicare claims for measuring the stage of disease of six commonly diagnosed cancers. RESEARCH DESIGN: Analysis of a database linking Surveillance, Epidemiology, and End Results (SEER) registry data and Medicare claims in patients aged 65 years with cancer. SUBJECTS: Three hundred twenty thousand, six hundred and thirty seven cases of invasive breast, colorectal, endometrial, lung, pancreatic, and prostate cancers diagnosed between 1984 and 1993. MEASURES: Using SEER files as the "gold standard," concordance with Medicare claims, as well as sensitivity and positive predictive value of coding for each stage was measured. RESULTS: Although Medicare data correctly categorized local, regional, and distant stage tumors in 97%, 33%, and 65%, respectively, the data substantially overestimated the proportion of localized tumors and underestimated the rate of regional stage disease. The highest concordance was observed for breast and colorectal cancer. However, the sensitivity and positive predictive values were never simultaneously 80% within one stage of a specific cancer. The accuracy of coding for stage in Outpatient files was inferior to inpatient data. CONCLUSIONS: With few exceptions, Medicare claims have limited utility as a measure of cancer stage. If tumor registry data are not available, investigators should consider the trade offs in sensitivity and predictive value when considering a study that will use claims data.

Aged↗

International Classification of Diseases-9th revision coding for preeclampsia: how accurate is it?

OBJECTIVE: The purpose of this study was to evaluate the accuracy of the International Classification of Diseases-9th revision codes for preeclampsia and eclampsia. STUDY DESIGN: The University of Illinois Medical Center at Chicago discharge database was used to identify 135 women from 1999 through 2001 whose disease was coded as having preeclampsia or eclampsia. With American College of Obstetrics and Gynecology criteria as the gold standard, the diagnosis that was determined through chart review was compared with the International Classification of Diseases-9th revision code that was present in the discharge database. Patients were classified as true cases if the International Classification of Diseases-9th revision code matched the American College of Obstetricians and Gynecologists diagnosis; the positive predictive value of the code was then calculated. RESULTS: The overall positive predictive value for the complete sample was only 54%, but the positive predictive value for severe preeclampsia was 84.8%, which was high compared with mild preeclampsia (45.3%) and eclampsia (41.7%). Diagnostic (clinician) error was the most common reason for miscoding error. CONCLUSION: The findings suggest that International Classification of Diseases-9th revision codes for preeclampsia/eclampsia vary greatly in their accuracy of diagnosis. Therefore, a review of medical records is required when data are being gathered on the incidence of preeclampsia and eclampsia.

Adolescent↗

Loci for act recall: contextual influence on the processing of action events.

A problem-solving account of act memory predicts stronger impacts of context than theories that explain act memory by reference to automatic processing or by reference to the operation of modality-specific code systems. This prediction was tested in three experiments, all using the loci memotechnique to provide contexts for memorization of subject-performed tasks (SPTs). The results of the three experiments did not provide unambiguous evidence for or against any of the rival theories. Most consistent, however, was the observation that memory under motor-encoding conditions profits less on contexts than memory under nonmotor-encoding conditions, a finding which by itself lends more support to a multicode than to a problem-solving interpretation.

Adolescent↗

Nucleotide sequence of the complementary DNA for human Pit-1/GHF-1.

Human cDNA clones encoding Pit-1/GHF-1, a pituitary-specific DNA binding factor, were obtained by PCR following reverse transcription of human pituitary RNA. It is approx. 1.3 kb in size with 0.1 kb 5' non-coding region, 0.9 kb protein-coding region and 0.3 kb 3' non-coding region. The predicted human Pit-1/GHF-1 peptide structure has 291 amino acids and is highly conserved among mouse, rat and bovine. In addition, the 5' non-coding region is highly conserved with rat pit-1/GHF-1 sequence to the transcription start site.

Amino Acid Sequence↗

Molecular analysis of the isocitrate lyase gene (acu-7) of the mushroom Coprinus cinereus.

The nucleotide sequence of the structural gene for isocitrate lyase (acu-7) is presented and features of its coding sequence and predicted protein are described. Several motifs were identified within the promoter region which are potentially involved in transcriptional regulation. Surprisingly, some of these occur within the coding sequence of an adjacent gene of unrelated function that terminates within 371 bp upstream from acu-7. The sequence of this second gene identified an N-acetylglucosamine-1-phosphate transferase.

Amino Acid Sequence↗

Tropheryma whipplei Twist: a human pathogenic Actinobacteria with a reduced genome.

The human pathogen Tropheryma whipplei is the only known reduced genome species (<1 Mb) within the Actinobacteria [high G+C Gram-positive bacteria]. We present the sequence of the 927303-bp circular genome of T. whipplei Twist strain, encoding 808 predicted protein-coding genes. Specific genome features include deficiencies in amino acid metabolisms, the lack of clear thioredoxin and thioredoxin reductase homologs, and a mutation in DNA gyrase predicting a resistance to quinolone antibiotics. Moreover, the alignment of the two available T. whipplei genome sequences (Twist vs. TW08/27) revealed a large chromosomal inversion the extremities of which are located within two paralogous genes. These genes belong to a large cell-surface protein family defined by the presence of a common repeat highly conserved at the nucleotide level. The repeats appear to trigger frequent genome rearrangements in T. whipplei, potentially resulting in the expression of different subsets of cell surface proteins. This might represent a new mechanism for evading host defenses. The T. whipplei genome sequence was also compared to other reduced bacterial genomes to examine the generality of previously detected features. The analysis of the genome sequence of this previously largely unknown human pathogen is now guiding the development of molecular diagnostic tools and more convenient culture conditions.

Actinomycetales↗

Secondary structure computer prediction of the poliovirus 5' non-coding region is improved by a genetic algorithm.

Comparison of the secondary structure of the 5' non-coding region of poliovirus 3 RNA derived from the genetic algorithm with the model of Skinner et al. (J. Mol. Biol., 207, 379-392, 1989) demonstrates many of the confirmed structural elements. The genetic algorithm (Shapiro and Navetta, J. Supercomput., 8, 195-201, 1994) generates a population of all possible stems, then mixes, combines, and recombines these stems in multiple iterations on a massively parallel computer, ultimately selecting a most fit structure based on its energy. The secondary structure of the region containing the determinants of neurovirulence was better predicted using the genetic algorithm, whereas the dynamic programming algorithm (Zuker, Science, 244, 48-52, 1989) required phylogenetic comparative sequence analysis to arrive at the correct conclusion. In addition, artificial mutations were introduced throughout this region of the genome and although rearrangements in structure may occur, many structures persisted, suggesting that the given structures thus selected may have evolved to withstand isolated mutations. The genetic algorithm-derived structure for the 5' non-coding region compares favorably with the biological data and functions previously described, and contains all of the 'persistent' structures, suggesting also that the persistence factor may be an aid to validating structures.

Algorithms↗

Dopamine neurons report an error in the temporal prediction of reward during learning.

Many behaviors are affected by rewards, undergoing long-term changes when rewards are different than predicted but remaining unchanged when rewards occur exactly as predicted. The discrepancy between reward occurrence and reward prediction is termed an 'error in reward prediction'. Dopamine neurons in the substantia nigra and the ventral tegmental area are believed to be involved in reward-dependent behaviors. Consistent with this role, they are activated by rewards, and because they are activated more strongly by unpredicted than by predicted rewards they may play a role in learning. The present study investigated whether monkey dopamine neurons code an error in reward prediction during the course of learning. Dopamine neuron responses reflected the changes in reward prediction during individual learning episodes; dopamine neurons were activated by rewards during early trials, when errors were frequent and rewards unpredictable, but activation was progressively reduced as performance was consolidated and rewards became more predictable. These neurons were also activated when rewards occurred at unpredicted times and were depressed when rewards were omitted at the predicted times. Thus, dopamine neurons code errors in the prediction of both the occurrence and the time of rewards. In this respect, their responses resemble the teaching signals that have been employed in particularly efficient computational learning models.

Animals↗

Molecular cloning of an alpha-glucosidase-like gene from Penicillium minioluteum and structure prediction of its gene product.

The dexC cDNA, which is expressed in dextran-containing medium by the filamentous fungus Penicillium minioluteum, was cloned and sequence characterized. The cDNA sequence comprises 1859 bp plus a poly (A) tail, coding for a predicted protein of 597 amino acids. The genomic counterpart was isolated by PCR, finding three introns in its sequence. The dexC gene was located by Southern blot in the same 9-kb fragment that the previously isolated dextranase-encoding gene (dexA). Sequence analysis revealed that the deduced DexC protein belongs to glycosyl hydrolase family 13, showing a high sequence identity (58%) with Aspergillus parasiticus alpha-1,6-glucosidase. In addition, the high sequence identity (51%) between DexC protein and oligo-1,6-glucosidase of Bacillus cereus, with three-dimensional (3D) structure determined, leads us to proposed a 3D model for the structural core of DexC protein.

Amino Acid Sequence↗

Presence of ATG triplets in 5' untranslated regions of eukaryotic cDNAs correlates with a 'weak' context of the start codon.

MOTIVATION: The context of the start codon (typically, AUG) and the features of the 5' Untranslated Regions (5' UTRs) are important for understanding translation regulation in eukaryotic mRNAs and for accurate prediction of the coding region in genomic and cDNA sequences. The presence of AUG triplets in 5' UTRs (upstream AUGs) might effect the initiation rate and, in the context of gene prediction, could reduce the accuracy of the identification of the authentic start. To reveal potential connections between the presence of upstream AUGs and other features of 5' UTRs, such as their length and the start codon context, we undertook a systematic analysis of the available eukaryotic 5' UTR sequences. RESULTS: We show that a large fraction of 5' UTRs in the available cDNA sequences, 15-53% depending on the organism, contain upstream ATGs. A negative correlation was observed between the information content of the translation start signal and the length of the 5' UTR. Similarly, a negative correlation exists between the 'strength' of the start context and the number of upstream ATGs. Typically, cDNAs containing long 5' UTRs with multiple upstream ATGs have a 'weak' start context, and in contrast, cDNAs containing short 5' UTRs without ATGs have 'strong' starts. These counter-intuitive results may be interpreted in terms of upstream AUGs having an important role in the regulation of translation efficiency by ensuring low basal translation level via double negative control and creating the potential for additional regulatory mechanisms. One of such mechanisms, supported by experimental studies of some mRNAs, includes removal of the AUG-containing portion of the 5' UTR by alternative splicing. AVAILABILITY: An ATG_ EVALUATOR program is available upon request or at www.itba.mi.cnr.it/webgene. CONTACT: rogozin@ncbi.nlm.nih.gov, milanesi@itba.mi.cnr.it.

5' Untranslated Regions↗