Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Adaptive predictive coding with applications to radiographs.

Picture archiving and communications systems (PACS) require storage and transmission of vast amounts of data. For design/cost considerations, it is desirable to reduce the size of these data without sacrificing the integrity of the stored information. The major considerations in designing a data-compression scheme for a PACS system are discussed: fidelity of the reconstructed image, bit rate, hardware complexity, and processing time. The basic principles of conventional nonadaptive differential pulse-code modulation (DPCM) are reviewed and compared with adaptive techniques. The effect of adaptive quantization on radiographic images is examined. Special consideration is given to block-adaptive DPCM or the "switched quantizer," which greatly enhances the system performance as compared with nonadaptive techniques, and conservatively has a 5:1 compression ratio. Sample radiographs substantiate the results.

Bone and Bones

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Information preserving image compression for archiving NMR images.

This paper presents a result on information preserving compression of NMR images for the archiving purpose. Both Lynch-Davisson coding and linear predictive coding have been studied. For NMR images of 256 x 256 x 12 resolution, the Lynch-Davisson coding with a block size of 64 as applied to prediction error sequences in the Gray code bit planes of each image gave an average compression ratio of 2.3:1 for 14 testing images. The predictive coding with a third order linear predictor and the Huffman encoding of the prediction error gave an average compression ratio of 3.1:1 for 54 images under test, while the maximum compression ratio achieved was 3.8:1. This result is one step further toward the improvement, albeit small, of the information preserving image compression for medical applications.

Algorithms

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms

Prediction of gene structure.

We have developed a hierarchical rule base system for identifying genes in DNA sequences. Atomic sites (such as initiation codons, stop codons, acceptor sites and donor sites) are identified by a number of different methods and evaluated by a set of filters and rules chosen to maximize sensitivity; these are combined into higher-order gene elements (such as exons), evaluated, filtered and combined as equivalence classes into probable genes, which are evaluated and ranked. The system has been tested on an extensive collection of vertebrate genes smaller than 15,000 bases. Results obtained show that, on average, 88% of the predicted coding region for a transcription unit is actually coding, and 80% of the actual coding is correctly predicted. This will, in most applications, be sufficient for a search against protein sequence databases for the identification of probable gene function. In addition, the system provides a general test platform for both gene atomic site identification and the rules for their evaluation and assembly.

Algorithms

Hepatitis B virus large surface protein is not secreted but is immunogenic when selectively expressed by recombinant vaccinia virus.

The envelope region of the hepatitis B virus (HBV) genome contains an open reading frame that begins upstream of the major surface protein gene. The two minor proteins that are initiated within this pre-s segment are immunogenic and may be involved in virus attachment to hepatocytes. We have constructed a recombinant vaccinia virus that contains the predicted coding segment for the large surface protein (LS) under control of a vaccinia virus that contains the predicted coding segment for the large surface protein (LS) under control of a vaccinia virus promoter. Cells infected with the recombinant virus synthesized HBV polypeptides of 39 and 42 kilodaltons, corresponding to the unglycosylated and glycosylated forms of LS, respectively. The presence of pre-s epitopes in the 39- and 42-kilodalton polypeptides was demonstrated by binding of antibody prepared against a synthetic peptide. Synthesis of the 42-kilodalton species was specifically inhibited by tunicamycin, suggesting that it is N-glycosylated. Despite apparent glycosylation, LS was not secreted into the medium of infected cells. Nevertheless, rabbits vaccinated with the purified recombinant virus made antibodies that recognized s and pre-s epitopes. Antibody to the NH2 terminus of LS appeared before or simultaneously with antibody that bound to the major surface protein. The additional immunogenicity provided by expression of LS may be advantageous for the development of an HBV vaccine.

Animals

Cloning and mapping of a testis-specific gene with sequence similarity to a sperm-coating glycoprotein gene.

A testis-specific gene Tpx-1, located between Pgk-2 and Mep-1 on mouse chromosome 17, was isolated from a cosmid clone, and its cDNA sequences were determined. The predicted coding sequence of Tpx-1 isolated from BALB/c mice showed 64.2% nucleotide and 55.1% amino acid sequence similarity with that of a rat sperm-coating glycoprotein gene, the protein product of which is secreted by the epididymis. To examine the evolutionary relationship between Tpx-1 and a sperm-coating glycoprotein gene, the cDNA sequence of TPX1, the human counterpart of Tpx-1, was determined. The comparison of the predicted coding sequences of Tpx-1 and TPX1 showed 77.8% nucleotide and 70% amino acid sequence similarity. Since Tpx-1 (from mouse) is more similar to TPX1 (from man) than it is to a rat sperm-coating glycoprotein gene, we conclude that Tpx-1 (TPX1) and a sperm-coating glycoprotein gene are closely related, but distinct, genes belonging to the same gene family. The predicted Tpx-1 protein of a t mutant mouse CRO437 differs from that of BALB/c mice by one amino acid insertion in the putative signal peptide. TPX1 was mapped to 6p21-qter by Southern blot analysis of interspecies somatic hybrid cell lines.

Amino Acid Sequence

Oligopeptide biases in protein sequences and their use in predicting protein coding regions in nucleotide sequences.

We have examined oligopeptides with lengths ranging from 2 to 11 residues in protein sequences that show no obvious evolutionary relationship. All sequences in the Protein Identification Resource database were carefully classified by sensitive homology searches into superfamilies to obtain unbiased oligopeptide counts. The results, contrary to previous studies, show clear prejudices in protein sequences. The oligopeptide preferences were used to help decide the significance of sequence homologies and to improve the more general methods for detecting protein coding regions within nucleotide sequences.

Amino Acid Sequence

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames

A suboptimum variable-length encoding procedure for discrete quantized data.

A method has been devised which allows to encode the quantized output of a discrete-time memoryless Gaussian source nearly as efficiently as Huffman's optimum variable-length encoding procedure. With respect to the mean code-word length the performance of the two methods typically differs only by 0.1-0.2 bit. The basic idea of the new method is that each code word is made up of two components: the prefix and the kernel. The prefix specifies the length of the kernel and is encoded by means of Huffman's method. Despite the fact that the kernel length may vary from one code word to the other, the code used for the kernel is basically a fixed-length code, because once the prefix has been decoded, the begin of the next code word is known without decoding the kernel. The advantage of the new method as compared to Huffman's method is that only few variable-length codes have to be distinguished so that both encoding and decoding can be accomplished by means of quite simple algorithms requiring only few compare and branch operations and one single addition or subtraction per code word. Areas of application, especially when combining the method with predictive coding, are time series (e.g. the electroencephalogram) and other kind of quantized data.

Algorithms

Characterization of the major capsid protein and cloning of its gene from algal virus PBCV-1.

The major capsid protein (Vp54) from Chlorella virus PBCV-1 is a glycoprotein and the most abundant viral structural protein. The gene encoding Vp54 has been cloned and sequenced. Initially, a region of the gene was amplified using the polymerase chain reaction (PCR) primed with oligonucleotides derived from the N-terminal amino acid sequences of purified protein and cyanogen bromide cleavage fragments. The PCR product was used as a probe to map the location of the gene to PBCV-1 genomic Pstl restriction fragment P8. A 1314-bp open reading frame (ORF) was identified which contained the predicted coding regions from the derived amino acid sequences. The peptide encoded by this ORF had a predicted molecular weight of 48.2 kDa and contained six putative N-linked and 63 putative O-linked glycosylation sites. Primer extension analysis indicated that transcription started 14 bp 5' to the ATG. The gene for Vp54 was transcribed late in infection and this transcript was the most abundant viral RNA present in infected cells.

Amino Acid Sequence

The period clock locus of D. melanogaster codes for a proteoglycan.

The period (per) gene of D. melanogaster is involved in the generation of biological rhythms. The most striking feature of the predicted coding sequence, corresponding to the key 4.5 kb transcript from this locus, is an extensive run of alternating Gly-Thr residues. This is homologous to a series of Gly-Ser repeats in a chondroitin sulfate proteoglycan. To determine whether the per transcript codes for a proteoglycan, a region of its coding sequence was expressed (in bacteria) as part of a fusion protein, which was used to immunize rabbits. When the resultant immune sera were used to probe fly protein preparations, they detected an antigen that is present in wild-type flies and absent in a per- mutant. Biochemical characterization of this antigen indicated that it is indeed a proteoglycan.

Amino Acid Sequence

Complete amino acid sequence of rat L-type pyruvate kinase deduced from the cDNA sequence.

cDNA clones, containing the entire coding region of rat L-type pyruvate kinase, were isolated and their nucleotide sequences were determined by the dideoxy-chain-termination method. The predicted coding region, which spans 543 amino acids, established the complete amino acid sequence of the L-type isozyme of pyruvate kinase for the first time. The deduced amino acid sequence of the L type has one phosphorylation site in its amino terminus and shows about 68% and 48% homologies with M1-type pyruvate kinase of chicken and yeast pyruvate kinase respectively. Domain A exhibits higher homology than domains B and C. The residues in the active site of the L-type enzyme of rats, lying between domains B and A2, are rather different from those of the M1-type enzyme of chickens, but other residues constituting the active site are identical with those of the chicken M1 type except for one amino acid substitution.

Amino Acid Sequence

Comorbidities, complications, and coding bias. Does the number of diagnosis codes matter in predicting in-hospital mortality?

OBJECTIVE: Incomplete coding of secondary diagnoses may bias assessments of patient risks of poor outcomes using administrative health care databases, most of which allow only five diagnoses. The Medicare program is expanding the number of possible diagnoses from five to nine, aiming to improve coding completeness. We examined the impact of having more diagnosis codes available on assessments of risk of death. DESIGN: We used 1988 computerized hospital discharge abstract data from California, which allow up to 25 diagnoses per discharge, to select a sample of hospitalized patients and assessed the relationship between the presence of 29 specific secondary diagnoses and the risk of in-hospital death. SETTING: Nonfederal acute-care hospitals in California. STUDY POPULATION: All patients at least 65 years of age who were hospitalized for stroke, pneumonia, acute myocardial infarction, or congestive heart failure in California in 1988 (N = 162,790). MAIN OUTCOME MEASURES: Relative risk of death for each specific secondary diagnosis. RESULTS: Many conditions that on a clinical basis would be expected to increase the risk of death, such as adult-onset diabetes mellitus, previous myocardial infarction, angina, and ventricular premature beats, were associated with a lower risk of in-hospital death. CONCLUSIONS: Bias against coding of chronic or comorbid conditions on the computerized discharge abstracts of patients who die best explains these results. Efforts to improve diagnosis coding completeness solely by increasing the number of available coding spaces may not succeed.

Aged

Human heme oxygenase-2: characterization and expression of a full-length cDNA and evidence suggesting that the two HO-2 transcripts may differ by choice of polyadenylation signal.

We show by Northern blot analysis that human HO-2 is encoded by two transcripts (1.3 and 1.7 kb) and is a single-copy gene as judged by Southern blot analysis. We further provide evidence based on Northern blot and sequence analysis of a cDNA representing the larger transcript that the transcripts differ in the 3' untranslated region. A 274-base-pair DNA fragment from the rat heme oxygenase-2 gene (I. Cruse and M.D. Maines, 1988, J. Biol. Chem. 263, 3348-3353) was used to isolate a human HO-2 cDNA from a fetal kidney library in lambda gt11. The clone, designated hK-1, was sequenced and the cDNA insert was determined to be 1625 base pairs in length, encoding a protein of 313 amino acids. Two consensus polyadenylation signals separated by 440 nucleotides were identified in the 3' untranslated region. The size of the cDNA insert closely approximated the larger of two mRNAs. The nucleotide sequence was 88% identical to the rat HO-2 gene within the predicted coding region and the putative translation product was also estimated to be 88% identical to the rat gene product (M. O. Rotenberg and D. Maines, 1990, J. Biol. Chem. 265, 7501). The predicted size, 36 kDa, corresponded well with HO-2 detected in human testis microsomes by Western blot analysis. Further, the fusion protein expressed in Escherichia coli displayed significant heme oxygenase activity, which was inhibited by Zn- and Sn-protoporphyrins, known inhibitors of eukaryotic heme oxygenase, but not by sulfhydryl reagents.

Amino Acid Sequence

Cloning and sequencing of the medium-chain S-acyl fatty acid synthetase thioester hydrolase cDNA from rat mammary gland.

cDNA clones coding for the medium-chain S-acyl fatty acid synthetase thioester hydrolase (thioesterase II) from rat mammary gland were identified in a bacteriophage lambda gt11 library and their nucleotide sequences were determined. The predicted coding region spans 263 amino acid residues and includes a sequence identical with that of a peptide derived from the enzyme active site. The rat thioesterase II cDNA sequence exhibits homology with that of a thioesterase found in duck uropygial glands.

Amino Acid Sequence

Nucleotide and predicted amino acid sequences of the Marek's disease virus and turkey herpesvirus thymidine kinase genes; comparison with thymidine kinase genes of other herpesviruses.

In this paper we present the nucleotide sequences of the thymidine kinase (TK) genes of two avian herpesviruses: a highly oncogenic strain of Marek's disease virus (MDV strain RB1B) and its serologically related vaccine virus, the herpesvirus of turkeys (HVT strain Fc-126). The predicted coding regions of the two genes are 1029 and 1050 nucleotides respectively, corresponding to polypeptides of 343 and 350 amino acids in length. Putative nucleotide- and nucleoside-binding sites have been identified within the two predicted amino acid sequences. The MDV and HVT TK amino acid sequences exhibit 58.2% amino acid identity. Comparison with other available herpesvirus TK sequences reveals a greater homology to those of the alphaherpesviruses than to those of the gammaherpesviruses. No overall homology was found when compared with the chicken cytoplasmic TK sequence.

Amino Acid Sequence