Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Organization of the psbE, psbF, orf38, and orf42 gene loci on the Euglena gracilis chloroplast genome.

The genes for cytochrome b559, designated psbE and psbF, and two highly conserved open reading frames of 38 and 42 codons have been located and characterized on the chloroplast genome of Euglena gracilis. The organization of the genes is psbE - 8 bp spacer - psbF - 110 bp spacer - orf38 - 87 bp spacer - orf42. All genes are of the same polarity. The psbE gene contains two introns of 350 and 326 bp. The psbF gene contains a single large intron of 1,042 bp. The orf38 and orf42 loci lack introns. The introns are extremely AT rich with a pronounced base composition bias of T greater than A greater than G greater than C in the mRNA-like strand and group II-like boundary sequences at their 3' and 5' ends having the consensus 5'-GTGTG .. INTRON .. TTAATTTNAT-3'. The psbE gene consists of 82 codons and encodes a polypeptide with a predicted molecular weight of 9,212. The psbF gene consists of 42 codons, which specify a polypeptide with a predicted molecular weight of 4,785. The highly conserved open reading frames of 38 and 42 codons code for polypeptides with predicted molecular weights of 4,405 and 4,426, respectively. The gene products of psbE, psbF, orf38 and orf42 are, respectively, 69.5%, 70% and 61.5% identical to those found in higher plants. The predicted secondary structure of the proteins from hydropathy plots is consistent with each containing a single membrane-spanning domain of at least 20 amino acids. Each of the genes is preceded by sequences which may serve as ribosome binding sites. All four genes are transcribed.

Amino Acid Sequence↗

Molecular analysis of 10 coding regions from Arabidopsis that are homologous to the MUR3 xyloglucan galactosyltransferase.

Plant cell walls are composed of a large number of complex polysaccharides, which contain at least 13 different monosaccharides in a multitude of linkages. This structural complexity of cell wall components is paralleled by a large number of predicted glycosyltransferases in plant genomes, which can be grouped into several distinct families based on conserved sequence motifs (B. Henrissat, G.J. Davies [2000] Plant Physiol 124: 1515-1519). Despite the wealth of genomic information in Arabidopsis and several crop plants, the biochemical functions of these coding regions have only been established in a few cases. To lay the foundation for the genetic and biochemical characterization of putative glycosyltransferase genes, we conducted a phylogenetic and expression analysis on 10 predicted coding regions (AtGT11-20) that are closely related to the MUR3 xyloglucan galactosyltransferase of Arabidopsis. All of these proteins contain the conserved sequence motif pfam 03016 that is the hallmark of the beta-d-glucuronosyltransferase domain of exostosins, a class of animal enzymes involved in the biosynthesis of the extracellular polysaccharide heparan sulfate. Reverse transcriptase-polymerase chain reaction and promoter:beta-glucuronidase studies indicate that all AtGT genes are transcribed. Although six of the 10 AtGT genes were expressed in all major plant organs, the remaining four genes showed more restricted expression patterns that were either confined to specific organs or to highly specialized cell types such as hydathodes or pollen grains. T-DNA insertion mutants in AtGT13 and AtGT18 displayed reductions in the Gal content of total cell wall material, suggesting that the disrupted genes encode galactosyltransferases in plant cell wall synthesis.

Amino Acid Sequence↗

The nucleotide sequence of herpes simplex virus type 2 (333) glycoprotein gB2 and analysis of predicted antigenic sites.

The gene coding for glycoprotein B2 (gB2) of herpes simplex virus type 2 (HSV-2) strain 333 was mapped and its nucleotide sequence determined. Open reading frame analysis deduced a polypeptide consisting of 902 amino acids and having close homology to gB1 of HSV type 1. Several predicted features of gB2 are consistent with a membrane-bound glycoprotein, i.e., a signal peptide sequence, a hydrophilic extracellular domain containing possible N-linked glycosylation sites, a hydrophobic membrane spanning sequence, and a cytoplasmic domain. Computer analysis on hydrophilicity, accessibility, and flexibility of the gB2 amino acid sequence, produced a composite surface value plot. At least nine major antigenic regions were predicted on the extracellular domain. The amino acids between residues 59-74, 127-139, 199-205, 460-476, and 580-594 exhibited the highest surface values. Comparison of the primary sequence with gB1 revealed localized regions showing amino acid diversity. Several of these locations correspond to major antigenic regions. Chou and Fasman analyses indicated that the amino acid substitutions, between positions 57-66, 461-472, and 473-481, induced changes in the secondary structure of gB. These sites could represent site-specific epitopes in the gB polypeptide.

Amino Acid Sequence↗

Identifying acute myocardial infarction: effects on treatment and mortality, and implications for National Service Framework audit.

BACKGROUND: The National Service Framework (NSF) for Coronary Heart Disease requires annual clinical audit of the care of patients with myocardial infarction, with little guidance on how to achieve these standards and monitor practice. AIM: To assess which method of identification of acute myocardial infarction (AMI) cases is most suitable for NSF audit, and to determine the effect of the definition of AMI on the assessment of quality of care. DESIGN: Observational study. METHODS: Over a 3-month period, 2153 consecutive patients from 20 hospitals across the Yorkshire region, with confirmed AMI, were identified from coronary care registers, biochemistry records and hospital coding systems. The sensitivity and positive predictive value of AMI patient identification using clinical coding, biochemistry and coronary care registers were compared to a 'gold standard' (the combination of all three methods). RESULTS: Of 3685 possible cases of AMI singled out by one or more methods, 2153 patients were identified as having a final diagnosis of AMI. Hospital coding revealed 1668 (77.5%) cases, with a demographic profile similar to that of the total cohort. Secondary preventative measures required for inclusion in NSF were also of broadly similar distribution. The sensitivities and positive predictive values for patient identification were substantially less in the cohorts identified through biochemistry and coronary care unit register. Patients fulfilling WHO criteria (n=1391) had a 30-day mortality of 15.9%, vs. 24.2% for the total cohort. DISCUSSION: Hospital coding misses a substantial proportion (22.5%) of AMI cases, but without any apparent systematic bias, and thus provides a suitably representative and robust basis for NSF-related audit. Better still would be the routine use of multiple methods of case identification.

Aged↗

Sensitivity and uncertainty studies of the CRAC2 computer code.

We have studied the sensitivity of health impacts from nuclear reactor accidents, as predicted by the CRAC2 computer code, to the following sources of uncertainty: (1) the model for plume rise, (2) the model for wet deposition, (3) the meteorological bin-sampling procedure for selecting weather sequences with rain, (4) the dose conversion factors for inhalation as affected by uncertainties in the particle size of the carrier aerosol and the clearance rates of radionuclides from the respiratory tract, (5) the weathering half-time for external ground-surface exposure, and (6) the transfer coefficients for terrestrial foodchain pathways. Predicted health impacts usually showed little sensitivity to use of an alternative plume-rise model or a modified rain-bin structure in bin-sampling. Health impacts often were quite sensitive to use of an alternative wet-deposition model in single-trial runs with rain during plume passage, but were less sensitive to the model in bin-sampling runs. Uncertainties in the inhalation dose conversion factors had important effects on early injuries in single-trial runs. Latent cancer fatalities were moderately sensitive to uncertainties in the weathering half-time for ground-surface exposure, but showed little sensitivity to the transfer coefficients for terrestrial foodchain pathways. Sensitivities of CRAC2 predictions to uncertainties in the models and parameters also depended on the magnitude of the source term, and some of the effects on early health effects were comparable to those that were due only to selection of different sets of weather sequences in bin-sampling.

Accidents↗

Sectional neuroanatomy of the lower limb I: lower back and hip.

This series of two articles is structured to provide anatomically accurate functional schematics of the motor and sensory innervation of the lower back, hip, and lower limb. This first paper provides radiographically oriented schematic axial sections of the lower back and hip in which the muscles are appropriately color-coded to match the peripheral nerves. A companion color-coded summary table allows prediction of unique patterns of denervation from 25 lesion sites. These are divided into three categories (roots T12 to S4, four plexal quadrants, and 11 sectional levels). Correlation between an imaging abnormality at one of these lesion sites and the predicted denervation pattern ensures the lesion is, in fact, clinically significant. The next article will continue this color-coded approach into the lower limb.

Electromyography↗

Finding prokaryotic genes by the 'frame-by-frame' algorithm: targeting gene starts and overlapping genes.

MOTIVATION: Tightly packed prokaryotic genes frequently overlap with each other. This feature, rarely seen in eukaryotic DNA, makes detection of translation initiation sites and, therefore, exact predictions of prokaryotic genes notoriously difficult. Improving the accuracy of precise gene prediction in prokaryotic genomic DNA remains an important open problem. RESULTS: A software program implementing a new algorithm utilizing a uniform Hidden Markov Model for prokaryotic gene prediction was developed. The algorithm analyzes a given DNA sequence in each of six possible global reading frames independently. Twelve complete prokaryotic genomes were analyzed using the new tool. The accuracy of gene finding, predicting locations of protein-coding ORFs, as well as the accuracy of precise gene prediction, and detecting the whole gene including translation initiation codon were assessed by comparison with existing annotation. It was shown that in terms of gene finding, the program performs at least as well as the previously developed tools, such as GeneMark and GLIMMER. In terms of precise gene prediction the new program was shown to be more accurate, by several percentage points, than earlier developed tools, such as GeneMark.hmm, ECOPARSE and ORPHEUS. The results of testing the program indicated the possibility of systematic bias in start codon annotation in several early sequenced prokaryotic genomes. AVAILABILITY: The new gene-finding program can be accessed through the Web site: http:@dixie.biology.gatech.edu/GeneMark/fbf.cgi CONTACT: mark@amber.gatech.edu.

Algorithms↗

Murine erythropoietin gene: cloning, expression, and human gene homology.

The gene for murine erythropoietin (EPO) was isolated from a mouse genomic library with a human EPO cDNA probe. Nucleotide sequence analysis permitted the identification of the murine EPO coding sequence and the prediction of the encoded amino acid sequence based on sequence conservation between the mouse and human EPO genes. Both the coding DNA and the amino acid sequences were 80% conserved between the two species. Transformation of COS-1 cells with a mammalian cell expression vector containing the murine EPO coding region resulted in secretion of murine EPO with biological activity on both murine and human erythroid progenitor cells. The transcription start site for the murine EPO gene in kidneys was determined. This permitted tentative identification of the transcription control region. The region included 140 base pairs upstream of the cap site which was over 90% conserved between the murine and human genes. Surprisingly, the first intron and much of the 5'- and 3'-untranslated sequences were also substantially conserved between the genes of the two species.

Amino Acid Sequence↗

Validity of information on gynecological operations in the Swedish in-patient registry.

In order to validate information held at the Swedish In-patient Registry on oophorectomy and/or hysterectomy procedures, the register codes were compared with data from the medical records in a random sample of 1,338 women. Only 1% of these codes were missing but 5% were erroneous, which in most cases meant that the oophorectomy had been misclassified. The positive predictive value of operation codes was high, ranging from 86 to 100% of the registered events. The actual procedures among women registered with a code for hysterectomy with or without oophorectomy comprised hysterectomy alone in 47% of the women and hysterectomy with a bilateral or unilateral oophorectomy in 30% and 20%, respectively. The reliability of register codes for major gynecological surgical procedures is good, but when the code for hysterectomy is used, medical record data are needed to ascertain the ovarian status. Revised codes are therefore recommended.

Adult↗

Definition of interferon gamma-response elements in a novel human Fc gamma receptor gene (Fc gamma RIb) and characterization of the gene structure.

The human Fc gamma RI (CD64) is a high affinity receptor for the Fc portion of immunoglobulin (Ig), and its constitutively low expression on the cell surface of monocyte/macrophage and neutrophils is selectively upregulated by interferon gamma (IFN-gamma) treatment (Perussia, B., E. T. Dayton, R. Lazarus, V. Fanning, and G. Trinchieri. 1983. J. Exp. Med. 158:1092). Three distinct cDNAs have been cloned and code for proteins that predict three extracellular Ig-like domains (Allen, J.M., and B. Seed. 1989. Science [Wash. DC]. 243:378). Several differences in the coding region of these cDNAs suggest that in addition to polymorphic differences a second Fc gamma RI gene could possibly exist. This alternative Fc gamma RI gene (Fc gamma RIb) was defined by the lack of a genomic HindIII restriction site (van der Winkel, J. G. J., L. U. Ernst, C. L. Anderson, and I. M. Chiu. 1991. J. Biol. Chem. 266:13449). We describe the characterization a second gene (Fc gamma RIb) that has a termination codon in the third extracellular domain and therefore predicts a soluble form of a termination codon in the third extracellular domain and therefore predicts a soluble form of the receptor. We also define two distinct IFN-gamma-responsive regions in the 5' flanking sequence of Fc gamma RIb that resemble motifs that have been defined in the class II major histocompatibility complex promoter. The Fc gamma RIb promoter does not possess canonical TATA or CCAAT boxes, but does possess a palindromic motif that closely resembles the initiator sequence identified in the terminal deoxynucleotidyl transferase/human leukocyte IFN/adeno-associated virus type II P5 gene promoters (Smale, S. T., and D. Baltimore. 1989. Cell. 57:103; Seto, E., Y. Shi, and T. Shenk. 1991. Nature [Lond.]. 354:241; Roy, A. L., M. Meisterernst, P. Pognonec, and R. C. Roeder. 1991. Nature [Lond.]. 354:245) virus type II P5 gene promoters raising interesting questions as to its role in the basal and myeloid-specific transcription of this gene.

Amino Acid Sequence↗

Mutually symmetric and complementary triplets: differences in their use distinguish systematically between coding and non-coding genomic sequences.

The general property of asymmetry in word use in meaningful texts written in a variety of languages, motivates a quantification of the differences in the use of mutually symmetric triplets in genomic sequences. When this is done in the three reading frames, high values found for one of them are used as indication that the sequence is coding for a protein. Moreover, a similar quantification of the differences in the use of complementary triplets is introduced, again with predictive power of the coding character of a sequence. This method reflects the non-equivalence between sense and anti-sense strand of a coding segment. In both approaches, "linguistic asymmetry" in coding sequences is related to the form of the genetic code and to the bias in codon usage and amino acid use skews.

Algorithms↗

Proteome-scale prediction of molecular mechanisms underlying dominant genetic diseases.

Many dominant genetic disorders result from protein-altering mutations, acting primarily through dominant-negative (DN), gain-of-function (GOF), and loss-of-function (LOF) mechanisms. Deciphering the mechanisms by which dominant diseases exert their effects is often experimentally challenging and resource intensive, but is essential for developing appropriate therapeutic approaches. Diseases that arise via a LOF mechanism are more amenable to be treated by conventional gene therapy, whereas DN and GOF mechanisms may require gene editing or targeting by small molecules. Moreover, pathogenic missense mutations that act via DN and GOF mechanisms are more difficult to identify than those that act via LOF using nearly all currently available variant effect predictors. Here, we introduce a tripartite statistical model made up of support vector machine binary classifiers trained to predict whether human protein coding genes are likely to be associated with DN, GOF, or LOF molecular disease mechanisms. We test the utility of the predictions by examining biologically and clinically meaningful properties known to be associated with the mechanisms. Our results strongly support that the models are able to generalise on unseen data and offer insight into the functional attributes of proteins associated with different mechanisms. We hope that our predictions will serve as a springboard for researchers studying novel variants and those of uncertain clinical significance, guiding variant interpretation strategies and experimental characterisation. Predictions for the human UniProt reference proteome are available at https://osf.io/z4dcp/.

Humans↗

Picture-word differences in decision latency: a test of common-coding assumptions.

Three experiments examined the processing of pictures and words in two tasks of semantic decision--judgments of conceptual size and judgments of associative relatedness--in order to test the prediction from single-coding models of memory that different semantic decisions produce comparable picture-word latency differences. In Experiment 1 an interaction in decision latency was found such that picture-picture (P-P) pairs were significantly faster than word-word (W-W) pairs in decisions of size but not in decisions of associative relatedness. In Experiment 2 no latency differences were found in decisions of association for pairs presented in P-P, W-W, or mixed (P-W or W-P) forms. Decisions of size, however, were fastest for P-P pairs, intermediate for mixed pairs, and slowest for W-W pairs. In a third experiment, using a speeded inference task, the interaction obtained in the first two experiments was reproduced. In light of these results possible revisions to common-coding assumptions about the processing of pictures and words in semantic decisions are discussed.

Discrimination Learning↗

A cyanobacterial gene family coding for single-helix proteins resembling part of the light-harvesting proteins from higher plants.

In the cyanobacterium Synechocystis sp. PCC 6803 five genes were identified with significant sequence similarity to regions of members of the eukaryotic chlorophyll a/b binding gene family (Cab family) and to hliA, a gene coding for a small high-light-induced protein in Synechococcus sp. PCC 7942. Four of these five genes are 174-213 bp in length and code for small proteins predicted to have a single transmembrane helix. The fifth Cab-like gene in Synechocystis sp. PCC 6803 is much longer and codes for a protein of which the N-terminal 80% resemble ferrochelatase but the C-terminal domain has similarity to Cab regions. The small genes were expressed preferentially in the absence of photosystem I, but gene expression was not significantly enhanced at moderately high light intensity. Therefore they were not designated as hli (high-light-induced) as was done for the Synechococcus sp. PCC 7942 homolog. Instead, the genes have been named scp, as the corresponding polypeptides of Synechocystis sp. PCC 6803 are small Cab-like proteins (SCP). The scpA gene, which codes for ferrochelatase with a C-terminal Cab-like extension, was interrupted by the insertion of a kanamycin-resistance cassette between the ferrochelatase and Cab-like gene domains. In the PS I-less background, interruption of scpA was found to lead to increased tolerance to high light intensity and to the requirement of a slightly higher light intensity to drive photosystem II electron transfer, suggestive of decreased light-harvesting efficiency in the absence of the C-terminal extension of ScpA. Immunodetection of ScpC and ScpD indicated that either or both accumulated in PS I-less strains. These proteins were also detected in bands of more than 45 kDa on denaturing gels, raising the possibility that they may occur as stable oligomers. The SCPs represent a new group of cyanobacterial proteins that, in view of their primary structure and response to deletion of photosystem I, are likely to be involved in transient pigment binding.

Amino Acid Sequence↗

Adaptive prediction trees for image compression.

This paper presents a complete general-purpose method for still-image compression called adaptive prediction trees. Efficient lossy and lossless compression of photographs, graphics, textual, and mixed images is achieved by ordering the data in a multicomponent binary pyramid, applying an empirically optimized nonlinear predictor, exploiting structural redundancies between color components, then coding with hex-trees and adaptive runlength/Huffman coders. Color palettization and order statistics prefiltering are applied adaptively as appropriate. Over a diverse image test set, the method outperforms standard lossless and lossy alternatives. The competing lossy alternatives use block transforms and wavelets in well-studied configurations. A major result of this paper is that predictive coding is a viable and sometimes preferable alternative to these methods.

Algorithms↗

A simplified model of the source channel of the Leksell Gamma Knife: testing multisource configurations with PENELOPE.

A simplification of the source channel geometry of the Leksell Gamma Knife (GK), recently proposed by the authors and checked for a single source configuration (Al-Dweri F M O, Lallena A M and Vilches M 2004 Phys. Med. Biol. 49 2687-703), has been used to calculate the dose distributions along the x, y and z axes in a water phantom with a diameter of 160 mm, for different configurations of the Gamma Knife, including 201, 150 and 102 unplugged sources. The code PENELOPE (v. 2001) has been used to perform the Monte Carlo simulations. In addition, the output factors for the 14, 8 and 4 mm helmets have been calculated. The results found for the dose profiles show a qualitatively good agreement with previous ones obtained with EGS4 and PENELOPE (v. 2000) codes and with the predictions of GammaPlan. The output factors obtained with our model agree within the statistical uncertainties with those calculated with the same Monte Carlo codes and with those measured with different techniques. Owing to the accuracy of the results obtained and to the reduction in the computational time with respect to full geometry simulations (larger than a factor 15), this simplified model opens the possibility of using Monte Carlo tools for planning purposes in the Gamma Knife.

Algorithms↗

HETC radiation transport code development for cosmic ray shielding applications in space.

In order to facilitate three-dimensional analyses of space radiation shielding scenarios for future space missions, the Monte Carlo radiation transport code HETC is being extended to include transport of energetic heavy ions, such as are found in the galactic cosmic ray spectrum in space. Recently, an event generator capable of providing nuclear interaction data for use in HETC was developed and incorporated into the code. The event generator predicts the interaction product yields and production angles and energies using nuclear models and Monte Carlo techniques. Testing and validation of the extended transport code has begun. In this work, the current status of code modifications, which enable energetic heavy ions and their nuclear reaction products to be transported through thick shielding, are described. Also, initial results of code testing against available laboratory beam data for energetic heavy ions interacting in thick targets are presented.

Construction Materials↗

Computational methods for the identification of differential and coordinated gene expression.

With the first complete 'draft' of the human genome sequence expected for Spring 2000, the three basic challenges for today's bioinformatics are more than ever: (i) finding the genes; (ii) locating their coding regions; and (iii) predicting their functions. However, our capacity for interpreting vertebrate genomic and transcript (cDNA) sequences using experimental or computational means very much lags behind our raw sequencing power. If the performances of current programs in identifying internal coding exons are good, the precise 5'-->3' delineation of transcription units (and promoters) still requires additional experiments. Similarly, functional predictions made with reference to previously characterized homologues are leaving >50% of human genes unannotated or classified in uninformative categories ('kinase', 'ATP-binding', etc.). In the context of functional genomics, large-scale gene expression studies using massive cDNA tag sequencing, two-dimensional gel proteome analysis or microarray technologies are the only approaches providing genome-scale experimental information at a pace consistent with the progress of sequencing. Given the difficulty and cost of characterizing genes one by one, academic and industrial researchers are increasingly relying on those methods to prioritize their studies and choose their targets. The study of expression patterns can also provide some insight into the function, reveal regulatory pathways, indicate side effects of drugs or serve as a diagnostic tool. In this article, I review the theoretical and computational approaches used to: (i) identify genes differentially expressed (across cell types, developmental stages, pathological conditions, etc.); (ii) identify genes expressed in a coordinated manner across a set of conditions; and (iii) delineate clusters of genes sharing coherent expression features, eventually defining global biological pathways.

Animals↗