Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Decoupled evolution of coding region and mRNA expression patterns after gene duplication: implications for the neutralist-selectionist debate.

The neutralist perspective on molecular evolution maintains that the vast majority of mutations affecting gene function are neutral or deleterious. After a gene duplication where both genes are retained, it predicts that original and duplicate genes diverge at clock-like rates. This prediction is usually tested for coding sequences, but can also be applied to another important aspect of gene function, the genes' expression pattern. Moreover, if both sequence and expression pattern diverge at clock-like rates, a correlation between divergence in sequence and divergence in expression patterns is expected. Duplicate gene pairs with more highly diverged sequences should also show more highly diverged expression patterns. This prediction is tested for a large sample of duplicated genes in the yeast Saccharomyces cerevisiae, using both genome sequence and microarray expression data. Only a weak correlation is observed, suggesting that coding sequence and mRNA expression patterns of duplicate gene pairs evolve independently and at vastly different rates. Implications of this finding for the neutralist-selectionist debate are discussed.

Biological Evolution↗

cis-acting RNA signals in the NS5B C-terminal coding sequence of the hepatitis C virus genome.

The cis-replicating RNA elements in the 5' and 3' nontranslated regions (NTRs) of the hepatitis C virus (HCV) genome have been thoroughly studied before. However, no cis-replicating elements have been identified in the coding sequences of the HCV polyprotein until very recently. The existence of highly conserved and stable stem-loop structures in the RNA polymerase NS5B coding sequence, however, has been previously predicted (A. Tuplin, J. Wood, D. J. Evans, A. H. Patel, and P. Simmonds, RNA 8:824-841, 2002). We have selected for our studies a 249-nt-long RNA segment in the C-terminal NS5B coding region (NS5BCR), which is predicted to form four stable stem-loop structures (SL-IV to SL-VII). By deletion and mutational analyses of the RNA structures, we have determined that two of the stem-loops (SL-V and SL-VI) are essential for replication of the HCV subgenomic replicon in Huh-7 cells. Mutations in the loop and the top of the stem of these RNA elements abolished replicon RNA synthesis but had no effect on translation. In vitro gel shift and filter-binding assays revealed that purified NS5B specifically binds to SL-V. The NS5B-RNA complexes were specifically competed away by unlabeled homologous RNA, to a small extent by 3' NTR RNA, and only poorly by 5' NTR RNA. The other two stem-loops (SL-IV and SL-VII) of the NS5BCR domain were found to be important but not essential for colony formation by the subgenomic replicon. The precise function(s) of these cis-acting RNA elements is not known.

Base Sequence↗

Malaysian antenatal risk coding and the outcome of pregnancy.

OBJECTIVE: Measure the effectiveness of the colour coding system in Malaysia for the prediction of risk in pregnancy. METHOD: Cohort study of records and interviews of 253/279 women examined at first antenatal visit. RESULTS: Nurses' final coding showed poor concordance with guidelines; recoding produced a predictive value of high risk of 48%; 25% of those with low risk had 50% of complications. Complication rates were moderate and intervention rates low. The mothers had little appreciation of risk and preferred home delivery. Home deliveries gave excellent results except for the 17% requiring transfer to hospital during labour or delivery. CONCLUSION: The coding system is ineffective with Malaysia's relatively low reproductive risk. Women require more personalised counseling about risk to make appropriate choices. Better results depend on simpler but consistent selection for hospital delivery using reproductive history, combined with better communication and transport systems for home deliveries and a reorientation within hospitals to rapid emergency care.

Adult↗

An assessment of the validity of ICD Code 410 to identify hospital admissions for myocardial infarction: The Corpus Christi Heart Project.

BACKGROUND: The identification of myocardial infarction (MI) is typically based on finding events designated by a nosologist with the appropriate International Classification of Diseases (ICD) code, currently code 410. These codes are applied based on review of medical records or death certificates. However, other factors, including reimbursement considerations, may influence the coding process, especially for hospitalizations. Thus, the validity of using ICD code 410 to identify MI must be assessed. METHODS: The Corpus Christi Heart Project (CCHP) is a population-based surveillance programme for hospitalized MI. Patients were identified using concurrent ascertainment in coronary care units and retrospective review of medical records. Events were validated as definite or possible MI using data regarding chest pain, electrocardiographic changes and cardiac enzymes. The validity of using ICD code 410 to identify cases of MI was assessed by calculating the sensitivity, specificity, predictive values and efficiency of ICD code 410 versus the CCHP 'gold standard'. RESULTS: Use of ICD code 410 identified 80.9% (401/496) of definite MI, but only 19.0% (243/1280) of possible MI. Only 12.3% (90/734) of discharges with an ICD 410 code received a 'no MI' designation based on the 'gold standard'. The efficiency of ICD code 410 for identifying MI was 92.0% for definite MI and 77.1% for definite and possible MI. CONCLUSIONS: The use of ICD code 410 to identify hospitalized cases of MI results in a modestly biased overestimate of the number of definite MI hospitalizations; however, this approach warrants consideration due to the expense of validation procedures.

Adult↗

Semidwarf (sd-1), "green revolution" rice, contains a defective gibberellin 20-oxidase gene.

The introduction of semidwarf rice (Oryza sativa L.) led to record yield increases throughout Asia in the 1960s. The major semidwarfing allele, sd-1, is still extensively used in modern rice cultivars. The phenotype of sd-1 is consistent with dwarfism that results from a deficiency in gibberellin (GA) plant growth hormones. We propose that the semidwarf (sd-1) phenotype is the result of a deficiency of active GAs in the elongating stem arising from a defective 20-oxidase GA biosynthetic enzyme. Sequence data from the rice genome was combined with previous mapping studies to locate a putative GA 20-oxidase gene (Os20ox2) at the predicted map location of sd-1 on chromosome 1. Two independent sd-1 alleles contained alterations within Os20ox2: a deletion of 280 bp within the coding region of Os20ox2 was predicted to encode a nonfunctional protein in an indica type semidwarf (Doongara), whereas a substitution in an amino acid residue (Leu-266) that is highly conserved among dioxygenases could explain loss of function of Os20ox2 in a japonica semidwarf (Calrose76). The quantification of GAs in elongating stems by GC-MS showed that the initial substrate of GA 20-oxidase activity (GA53) accumulated, whereas the content of the major product (GA20) and of bioactive GA1 was lower in semidwarf compared with tall lines. We propose that the Os20ox2 gene corresponds to the sd-1 locus.

Amino Acid Sequence↗

How should we measure social disadvantage in clinic settings?

BACKGROUND: Despite a large research literature supporting their validity, deprivation indices derived from census data have not been routinely applied to clinic populations. METHOD: A case-note sample of 201 cases was examined, to identify whether such data (Jarman indices) predicted presenting disability separately from diagnostic class (behaviour, emotional, mixed, other, and no diagnosable disorder), or conventional clinic measures of social adversity (ICD-10 psychosocial diagnostic codes). RESULTS: Jarman index scores predicted disability in behaviour disorders or other disorders. Conventional clinic measures of adversity predicted disability in mixed disorders. For emotional disorders, and those cases with no diagnosed disorder, clinically measured adversity and Jarman scores interacted. CONCLUSIONS: Postcode related census data capture information about clinic children's presenting disability that is not available from routine clinic assessment of psychosocial adversity. It should therefore be collected as part of the routine clinical child psychiatry assessment.

Adolescent↗

Initiation of translation at an AUA codon for an archaebacterial protein gene expressed in E.coli.

Overexpression of the Sulfolobus solfataricus L12 ribosomal protein gene in E.coli cells yielded two products of different size. If the E.coli cells carrying the overexpression plasmid were induced in the early stage of bacterial growth, the smaller of the two products was almost exclusively produced. However, induction in a late stage of bacterial growth yielded the larger product in significant excess. The larger protein was identified as the translation product of the entire SsoL12 gene, while the smaller product was a N-terminally shortened version of the L12 protein (sh-SsoL12), starting with a N-terminal methionine at position 22 of the coded protein and continuing with the predicted protein sequence. Position 22 is an isoleucine in the complete SsoL12 protein sequence, coded by an AUA codon. A subclone (SsoL12**) of the SsoL12 gene containing overexpression plasmid, lacking the regular AUG start codon and the putative Shine Dalgarno sequence, was constructed to determine if E.coli ribosomes could initiate at this AUA codon. During overexpression the SsoL12** construct yielded exclusively the sh-SsoL12 product in significant amounts. An AUA start codon has never been found before in a natural message. However, experiments utilizing site directed mutagenesis to generate AUA start codons showed that this codon can be functional for initiation in prokaryotes and eukaryotes. The findings presented in this paper show that AUA acts as an initiation codon in a natural message expressed in a heterologous organism.

Amino Acid Sequence↗

An unsupervised classification scheme for improving predictions of prokaryotic TIS.

BACKGROUND: Although it is not difficult for state-of-the-art gene finders to identify coding regions in prokaryotic genomes, exact prediction of the corresponding translation initiation sites (TIS) is still a challenging problem. Recently a number of post-processing tools have been proposed for improving the annotation of prokaryotic TIS. However, inherent difficulties of these approaches arise from the considerable variation of TIS characteristics across different species. Therefore prior assumptions about the properties of prokaryotic gene starts may cause suboptimal predictions for newly sequenced genomes with TIS signals differing from those of well-investigated genomes. RESULTS: We introduce a clustering algorithm for completely unsupervised scoring of potential TIS, based on positionally smoothed probability matrices. The algorithm requires an initial gene prediction and the genomic sequence of the organism to perform the reannotation. As compared with other methods for improving predictions of gene starts in bacterial genomes, our approach is not based on any specific assumptions about prokaryotic TIS. Despite the generality of the underlying algorithm, the prediction rate of our method is competitive on experimentally verified test data from E. coli and B. subtilis. Regarding genomes with high G+C content, in contrast to some previously proposed methods, our algorithm also provides good performance on P. aeruginosa, B. pseudomallei and R. solanacearum. CONCLUSION: On reliable test data we showed that our method provides good results in post-processing the predictions of the widely-used program GLIMMER. The underlying clustering algorithm is robust with respect to variations in the initial TIS annotation and does not require specific assumptions about prokaryotic gene starts. These features are particularly useful on genomes with high G+C content. The algorithm has been implemented in the tool "TICO" (TIs COrrector) which is publicly available from our web site.

Algorithms↗

Validation of 1997 Partin Tables' lymph node invasion predictions in men treated with radical prostatectomy in Montreal Quebec.

OBJECTIVE: The accuracy of 1997 Partin Tables' lymph node invasion (LNI) predictions exhibits important variability in different testing populations. We explored the LNI predictive accuracy in radical prostatectomy (RP) patients from Montreal, Canada. Moreover, we assessed the extent of change in predictive accuracy related to a modification of PSA coding from categorical to continuous. METHODS: We used pretreatment serum PSA, clinical stage, and biopsy Gleason sum from 537 men treated with RP to compare predicted and observed rates of LNI. Accuracy was quantified with receiver-operating characteristics curves. RESULTS: Accuracy was 0.760 in 369 evaluable patients, when categorically coded pretreatment PSA (0-4, 4.1-10, 10.1-20, 20.1+) was combined with clinical stage and biopsy Gleason sum. A 2.7% accuracy increase was noted when categorically coded PSA was replaced with continuously coded values. CONCLUSION: Partin Tables' LNI predictions showed comparable accuracy to a community-based sample from the United States (0.766), and to a recent, multi-institutional sample (0.740). However, accuracy was lower than reported in internal (0.818), and external (0.837) academic, validation cohorts. Accuracy of LNI predictions was appreciably higher, when continuously coded PSA was used.

Humans↗

The mutational spectrum of single base-pair substitutions causing human genetic disease: patterns and predictions.

Reports of single base-pair substitutions that cause human genetic disease and that have been located and characterized in an unbiased fashion were collated; 32% of point mutations were CG----TG or CG----CA transitions consistent with a chemical model of mutation via methylation-mediated deamination. This represents a 12-fold higher frequency than that predicted from random expectation, confirming that CG dinucleotides are indeed hotspots of mutation causing human genetic disease. However, since CG also appears hypermutable irrespective of methylation-mediated deamination, a second mechanism may also be involved in generating CG mutations. The spectrum of point mutations occurring outwith CG dinucleotides is also non-random, at both the mono- and dinucleotide, levels. An intrinsic bias in clinical detection was excluded since frequencies of specific amino acid substitutions did not correlate with the 'chemical difference' between the amino acids exchanged. Instead, a strong correlation was observed with the mutational spectrum predicted from the experimentally measured mispairing frequencies of vertebrate DNA polymerases alpha and beta in vitro. This correlation appears to be independent of any difference in the efficiency of enzymatic proofreading/mismatch-repair mechanisms but is consistent with a physical model of mutation through nucleotide misincorporation as a result of transient misalignment of bases at the replication fork. This model is further supported by an observed correlation between dinucleotide mutability and stability, possibly because transient misalignment must be stabilized long enough for misincorporation to occur. Since point mutations in human genes causing genetic disease neither arise by random error nor are independent of their local sequence environment, predictive models may be considered. We present a computer model (MUTPRED) based upon empirical data; it is designed to predict the location of point mutations within gene coding regions causing human genetic disease. The mutational spectrum predicted for the human factor IX gene was shown to resemble closely the observed spectrum of point mutations causing haemophilia B. Further, the model was able to predict successfully the rank order of disease prevalence and/or mutation rates associated with various human autosomal dominant and sex-linked recessive conditions. Although still imperfect, this model nevertheless represents an initial attempt to relate the variable prevalence of human genetic disease to the mutability inherent in the nucleotide sequences of the underlying genes.

Amino Acid Sequence↗

A comparison of the COG and MCNP codes in computational neutron capture therapy modeling, Part II: gadolinium neutron capture therapy models and therapeutic effects.

The goal of this study was to evaluate the COG Monte Carlo radiation transport code, developed and tested by Lawrence Livermore National Laboratory, for gadolinium neutron capture therapy (GdNCT) related modeling. The validity of COG NCT model has been established for this model, and here the calculation was extended to analyze the effect of various gadolinium concentrations on dose distribution and cell-kill effect of the GdNCT modality and to determine the optimum therapeutic conditions for treating brain cancers. The computational results were compared with the widely used MCNP code. The differences between the COG and MCNP predictions were generally small and suggest that the COG code can be applied to similar research problems in NCT. Results for this study also showed that a concentration of 100 ppm gadolinium in the tumor was most beneficial when using an epithermal neutron beam.

Body Burden↗

Nucleotide sequence of the genes encoding the major tail sheath and tail tube proteins of bacteriophage P2.

The major structural components of the contractile tail of bacteriophage P2 are proteins FI and FII, which are believed to be the tail sheath and tube proteins, respectively. Both proteins were mapped previously to the P2 late gene F, based on the pattern of protein synthesis in various P2 amber mutants. In order to clarify the gene arrangement and to provide a basis for structural comparisons with other contractile phage tails, we have determined the nucleotide sequence of the region of the P2 genome encoding these two proteins. The coding regions were confirmed by location of the Fam4 mutation and by N-terminal amino acid sequencing of both proteins. The molecular weight and amino acid composition predicted by each of the coding regions correspond well to those determined experimentally for each protein. FII is encoded by a newly identified P2 late gene. These proteins bear little resemblance to their functional homologues in bacteriophage T4.

Amino Acid Sequence↗

Identification of Bcd, a novel proto-oncogene expressed in B-cells.

A novel B-cell derived (Bcd) oncogene has been isolated from the peripheral blood lymphocytes of one B-cell chronic lymphocytic leukemia (B-CLL) patient using DNA transfer and a mouse tumorigenicity assay. The Bcd proto-oncogene was activated by a truncation in the 5' UTR. It predicts for two open reading frames (ORFs). ORF1 consists of 240 bp that would encode 80 amino acids, while the major ORF2 consists of 648 bp capable of coding for 216 amino acids. Predicted peptide sequence of ORF2 contained a zinc finger domain which showed significant homology to GC box binding proteins BTEB2 and SP1. Transfection of an expression vector containing ORF2 but not full length cDNA was able to transform NIH3T3 cells and induce tumors in nude mice. Bcd mRNA transcripts of < or = 2.6 kb were selectively expressed in PBL and testis of healthy individuals. Within the PBL, Bcd gene expression was restricted to CD19+ B-cells and absent from CD14+ monocytes and T-cells. Bcd transcripts were detected in all normal PBL samples tested but not in several malignant human B-cell lines and not in 50% of B-cells from B-CLL patients. However, stimulation of B-cells from B-CLL patients under conditions which induced differentiation into plasma cells was associated with induction of Bcd gene expression. The Bcd gene may therefore play an important role in B-cell growth and development.

3T3 Cells↗

Towards the proteome of the marine bacterium Rhodopirellula baltica: mapping the soluble proteins.

The marine bacterium Rhodopirellula baltica, a member of the phylum Planctomycetes, has distinct morphological properties and contributes to remineralization of biomass in the natural environment. On the basis of its recently determined complete genome we investigated its proteome by 2-DE and established a reference 2-DE gel for the soluble protein fraction. Approximately 1000 protein spots were excised from a colloidal Coomassie-stained gel (pH 4-7), analyzed by MALDI-MS and identified by PMF. The non-redundant data set contained 626 distinct protein spots, corresponding to 558 different genes. The identified proteins were classified into role categories according to their predicted functions. The experimentally determined and the theoretically predicted proteomes were compared. Proteins, which were most abundant in 2-DE gels and the coding genes of which were also predicted to be highly expressed, could be linked mainly to housekeeping functions in glycolysis, tricarboxic acid cycle, amino acid biosynthesis, protein quality control and translation. Absence of predictable signal peptides indicated a localization of these proteins in the intracellular compartment, the pirellulosome. Among the identified proteins, 146 contained a predicted signal peptide suggesting their translocation. Some proteins were detected in more than one spot on the gel, indicating post-translational modification. In addition to identifying proteins present in the published sequence database for R. baltica, an alternative approach was used, in which the mass spectrometric data was searched against a maximal ORF set, allowing the identification of four previously unpredicted ORFs. The 2-DE reference map presented here will serve as framework for further experiments to study differential gene expression of R. baltica in response to external stimuli or cellular development and compartmentalization.

Bacteria↗

cDNA clones encoding bovine gamma-crystallins.

We have determined the nucleotide sequence of two bovine lens gamma-crystallin cDNA clones, pBL gamma II-1 and pBL gamma III-1. The 644 bp cDNA insert of pBL gamma II-1 contains coding information for the entire amino acid sequence of bovine gamma II-crystallin. The 497 bp cDNA insert of pBL gamma III-1 encodes a homologous but different gamma-crystallin polypeptide, and appears to lack the coding information for the C-terminal 17 amino acid residues. While the nucleotide and predicted amino acid sequences of the coding regions of the clones show a high degree of homology, the untranslated leader sequences are relatively dissimilar. The leader sequence of pBL gamma III-1 is strikingly homologous to a portion of a rabbit immunoglobulin alpha-heavy chain mRNA.

Amino Acid Sequence↗

A challenge for 21st century molecular biology and biochemistry: what are the causes of obligate autotrophy and methanotrophy?

We assess the use to which bioinformatics in the form of bacterial genome sequences, functional gene probes and the protein sequence databases can be applied to hypotheses about obligate autotrophy in eubacteria. Obligate methanotrophy and obligate autotrophy among the chemo- and photo-lithotrophic bacteria lack satisfactory explanation a century or more after their discovery. Various causes of these phenomena have been suggested, which we review in the light of the information currently available. Among these suggestions is the absence in vivo of a functional alpha-ketoglutarate dehydrogenase. The advent of complete and partial genome sequences of diverse autotrophs, methylotrophs and methanotrophs makes it possible to probe the reasons for the absence of activity of this enzyme. We review the role and evolutionary origins of the Krebs cycle in relation to autotrophic metabolism and describe the use of in silico methods to probe the partial and complete genome sequences of a variety of obligate genera for genes encoding the subunits of the alpha-ketoglutarate dehydrogenase complex. Nitrosomonas europaea and Methylococcus capsulatus, which lack the functional enzyme, were found to contain the coding sequences for the E1 and E2 subunits of alpha-ketoglutarate dehydrogenase. Comparing the predicted physicochemical properties of the polypeptides coded by the genes confirmed the putative gene products were similar to the active alpha-ketoglutarate dehydrogenase subunits of heterotrophs. These obligate species are thus genomically competent with respect to this enzyme but are apparently incapable of producing a functional enzyme. Probing of the full and incomplete genomes of some cyanobacterial and methanogenic genera and Aquifex confirms or suggests the absence of the genes for at least one of the three components of the alpha-ketoglutarate dehydrogenase complex in these obligate organisms. It is recognized that absence of a single functional enzyme may not explain obligate autotrophy in all cases and may indeed be only be one of a number of controls that impose obligate metabolism. Availability of more genome sequences from obligate genera will enable assessment of whether obligate autotrophy is due to the absence of genes for a few or many steps in organic compound metabolism. This problem needs the technologies and mindsets of the present generation of molecular microbiologists to resolve it.

Archaea↗

Isolation of a cDNA clone for the human lysosomal proteinase cathepsin B.

The cysteine proteinase cathepsin B is one member of the lysosomal acid hydrolases. Based on the peptide sequence of rat liver cathepsin B, an oligonucleotide mixture containing 128 different 17-mers was synthesized and used as a probe to screen adult and fetal human liver cDNA libraries. A recombinant clone with a 1540-nucleotide insert was identified from the fetal library, and DNA sequence analysis confirmed that this clone encodes human cathepsin B. The clone, designated pCB-1, has sequences for 81% of the coding region (for amino acid residues 50-252) together with approximately equal to 880 nucleotides of the 3' untranslated region of the mRNA. The DNA sequence also shows that the predicted carboxyl terminus of the coding sequence is longer than the mature protein by 6 amino acid residues. Southern blot analysis of restriction enzyme digests of human placental DNA revealed a simple pattern of hybridizing fragments using the cathepsin B coding sequence as probe. The result suggests that there is a single copy of cathepsin B gene per haploid genome.

Amino Acid Sequence↗

Source-optimized irregular repeat accumulate codes with inherent unequal error protection capabilities and their application to scalable image transmission.

The common practice for achieving unequal error protection (UEP) in scalable multimedia communication systems is to design rate-compatible punctured channel codes before computing the UEP rate assignments. This paper proposes a new approach to designing powerful irregular repeat accumulate (IRA) codes that are optimized for the multimedia source and to exploiting the inherent irregularity in IRA codes for UEP. Using the end-to-end distortion due to the first error bit in channel decoding as the cost function, which is readily given by the operational distortion-rate function of embedded source codes, we incorporate this cost function into the channel code design process via density evolution and obtain IRA codes that minimize the average cost function instead of the usual probability of error. Because the resulting IRA codes have inherent UEP capabilities due to irregularity, the new IRA code design effectively integrates channel code optimization and UEP rate assignments, resulting in source-optimized channel coding or joint source-channel coding. We simulate our source-optimized IRA codes for transporting SPIHT-coded images over a binary symmetric channel with crossover probability p. When p = 0.03 and the channel code length is long (e.g., with one codeword for the whole 512 x 512 image), we are able to operate at only 9.38% away from the channel capacity with code length 132380 bits, achieving the best published results in terms of average peak signal-to-noise ratio (PSNR). Compared to conventional IRA code design (that minimizes the probability of error) with the same code rate, the performance gain in average PSNR from using our proposed source-optimized IRA code design is 0.8759 dB when p = 0.1 and the code length is 12800 bits. As predicted by Shannon's separation principle, we observe that this performance gain diminishes as the code length increases.

Algorithms↗