Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

NAVIP: Unraveling the influence of neighboring small sequence variants on functional impact prediction.

Once a suitable reference sequence has been generated, intra-species variation is often assessed by re-sequencing. Variant calling processes can reveal all differences between strains, accessions, genotypes, or individuals. These variants can be enriched with predictions about their functional implications based on available structural annotations, i.e., gene models. Although these functional impact predictions on a per-variant basis are often accurate, some challenging cases require the simultaneous incorporation of multiple adjacent variants into this prediction process. Examples include neighboring variants which modify each other's functional impact. The Neighborhood-Aware Variant Impact Predictor (NAVIP) considers all variants within a given protein coding sequence when predicting the effect. As a proof of concept, variants between the Arabidopsis thaliana accessions Columbia-0 and Niederzenz-1 were annotated. NAVIP is freely available on GitHub (https://github.com/bpucker/NAVIP) and accessible through a web server (https://pbb-tools.de).

Arabidopsis↗

[Prediction of protein conformation using a doublet code method].

It is suggested that regions of irregular structure, beta-structure, and alpha-helix are composed of 2, 3, and 5 amino acid residue long elements (structurons), respectively, and that the structurons are encoded solely by residue pairs (doublet codons) (i, i + 1), (i, i + 2), (i, i + 4), respectively. Tables of codons are obtained by statistical analysis of the data on the distribution of these pairs in available secondary structures of 62 proteins. These tables are used to obtain distributions of t-, beta- and alpha-codons for an amino acid sequence of protein. When codons of different structures superpose, that is, include the same sequence regions, selection is performed, the selection being performed so to obtain as much as possible number of the non-superposed codons of different structures. The distributions of structurons obtained after this selection are used for localization of structurons in the sequence and prediction of secondary structure on the basis of this localization. The prediction method is illustrated. An accuracy of the method has been tested on the basis an casual selection of fifteen proteins and found equal 64% for secondary structure on the whole and 79%, 53%, 61% for alpha-helix, beta-structure and coil respectively. This result is similar or better than that communicated for contemporary methods.

Amino Acid Sequence↗

The complete DNA sequence of the mitochondrial genome of Physarum polycephalum.

The complete sequence of the mitochondrial DNA (mtDNA) of the true slime mold Physarun polycephalum has been determined. The mtDNA is a circular 62,862-bp molecule with an A+T content of 74.1%. A search with the program BLAST X identified the protein-coding regions. The mitochondrial genome of P. polycephalum was predicted to contain genes coding for 12 known proteins [for three cytochrome c oxidase subunits, apocytochrome b, two F1Fo-ATPase subunits, five NADH dehydrogenase (nad) subunits, and one ribosomal protein], two rRNA genes, and five tRNA genes. However, the predicted ORFs are not all in the same frame, because mitochondrial RNA in P. polycephalum undergoes RNA editing to produce functional RNAs. The nucleotide sequence of an nad7 cDNA showed that 51 nucleotides were inserted at 46 sites in the mRNA. No guide RNA-like sequences were observed in the mtDNA of P. polycephalum. Comparison with reported Physarum mtDNA sequences suggested that sites of RNA editing vary among strains. In the Physarum mtDNA, 20 ORFs of over 300 nucleotides were found and ORFs 14 19 are transcribed.

Amino Acid Sequence↗

Sequence and analysis of chromosome 4 of the plant Arabidopsis thaliana.

The higher plant Arabidopsis thaliana (Arabidopsis) is an important model for identifying plant genes and determining their function. To assist biological investigations and to define chromosome structure, a coordinated effort to sequence the Arabidopsis genome was initiated in late 1996. Here we report one of the first milestones of this project, the sequence of chromosome 4. Analysis of 17.38 megabases of unique sequence, representing about 17% of the genome, reveals 3,744 protein coding genes, 81 transfer RNAs and numerous repeat elements. Heterochromatic regions surrounding the putative centromere, which has not yet been completely sequenced, are characterized by an increased frequency of a variety of repeats, new repeats, reduced recombination, lowered gene density and lowered gene expression. Roughly 60% of the predicted protein-coding genes have been functionally characterized on the basis of their homology to known genes. Many genes encode predicted proteins that are homologous to human and Caenorhabditis elegans proteins.

Animals↗

PLEXIN-D1, a novel plexin family member, is expressed in vascular endothelium and the central nervous system during mouse embryogenesis.

The genetic defect in Möbius syndrome 2 (MBS2, MIM 601471), a dominantly inherited disorder characterised by paralysis of the facial nerve, is situated at chromosome 3q21-q22. We characterised the cDNA and predicted protein, and examined the expression pattern during mouse embryogenesis of a positional candidate gene, PLEXIN-D1 (PLXND1). The cDNA for PLXND1 is 7095 base pairs in length, coding for a predicted protein of 1925 amino acids. The protein features all known domains of plexin family members, with the exception of the third Met-related sequence. Northern analysis revealed a very low expression of PLXND1 in adult mouse and adult human tissues. To investigate the expression of PlxnD1 during embryogenesis, RNA in situ hybridisation was performed on mouse embryos from various stages. This investigation revealed expression of PlxnD1 in cells from the central nervous system (CNS) and in vascular endothelium. Early expression in the CNS is located in the ganglia, cortical plate of the cortex, and striatum. At later embryologic stages, neural expression was also seen in the external granular layer of the cerebellum and several nerve nuclei. The expression in the vascular system resides solely in the endothelial cells of developing blood vessels. Based on our results, we suggest that this expression of a member of the plexin family in vascular endothelium could point toward a role in embryonic vasculogenesis.

Amino Acid Sequence↗

Importing statistical measures into Artemis enhances gene identification in the Leishmania genome project.

BACKGROUND: Seattle Biomedical Research Institute (SBRI) as part of the Leishmania Genome Network (LGN) is sequencing chromosomes of the trypanosomatid protozoan species Leishmania major. At SBRI, chromosomal sequence is annotated using a combination of trained and untrained non-consensus gene-prediction algorithms with ARTEMIS, an annotation platform with rich and user-friendly interfaces. RESULTS: Here we describe a methodology used to import results from three different protein-coding gene-prediction algorithms (GLIMMER, TESTCODE and GENESCAN) into the ARTEMIS sequence viewer and annotation tool. Comparison of these methods, along with the CODONUSAGE algorithm built into ARTEMIS, shows the importance of combining methods to more accurately annotate the L. major genomic sequence. CONCLUSION: An improvised and powerful tool for gene prediction has been developed by importing data from widely-used algorithms into an existing annotation platform. This approach is especially fruitful in the Leishmania genome project where there is large proportion of novel genes requiring manual annotation.

Algorithms↗

Analysis and recognition of 5' UTR intron splice sites in human pre-mRNA.

Prediction of splice sites in non-coding regions of genes is one of the most challenging aspects of gene structure recognition. We perform a rigorous analysis of such splice sites embedded in human 5' untranslated regions (UTRs), and investigate correlations between this class of splice sites and other features found in the adjacent exons and introns. By restricting the training of neural network algorithms to 'pure' UTRs (not extending partially into protein coding regions), we for the first time investigate the predictive power of the splicing signal proper, in contrast to conventional splice site prediction, which typically relies on the change in sequence at the transition from protein coding to non-coding. By doing so, the algorithms were able to pick up subtler splicing signals that were otherwise masked by 'coding' noise, thus enhancing significantly the prediction of 5' UTR splice sites. For example, the non-coding splice site predicting networks pick up compositional and positional bias in the 3' ends of non-coding exons and 5' non-coding intron ends, where cytosine and guanine are over-represented. This compositional bias at the true UTR donor sites is also visible in the synaptic weights of the neural networks trained to identify UTR donor sites. Conventional splice site prediction methods perform poorly in UTRs because the reading frame pattern is absent. The NetUTR method presented here performs 2-3-fold better compared with NetGene2 and GenScan in 5' UTRs. We also tested the 5' UTR trained method on protein coding regions, and discovered, surprisingly, that it works quite well (although it cannot compete with NetGene2). This indicates that the local splicing pattern in UTRs and coding regions is largely the same. The NetUTR method is made publicly available at www.cbs.dtu.dk/services/NetUTR.

5' Untranslated Regions↗

Expression of the protease gene of equine infectious anemia virus in Escherichia coli: formation of the mature processed enzyme and specific cleavage of the gag precursor.

A 620-bp Bg/II restriction fragment containing the putative protease coding sequence from equine infectious anemia virus (EIAV) proviral DNA was cloned and expressed in E. coli as a Pol precursor protein. In contrast to the 25-kDa fusion protein predicted from the expressed pol sequence, a protein of approximately 10 kDa was generated by apparent autocatalytic processing of the Pol precursor. This mature processed protein was detected in transformed cells using an antisera raised against synthetic peptide from the conserved carboxyl-terminal segment of the predicted EIAV protease coding sequence. Coexpression of this protein with a 35-kDa EIAV Gag-precursor fusion protein resulted in the specific proteolytic processing of the precursor as shown by formation of p26, the major capsid protein of EIAV.

Cloning, Molecular↗

Modeling and predicting transcriptional units of Escherichia coli genes using hidden Markov models.

MOTIVATION: The hidden Markov model (HMM) is a valuable technique for gene-finding, especially because its flexibility enables the inclusion of various sequence features. Recent programs for bacterial gene-finding include the information of ribosomal binding site (RBS) to improve the recognition accuracy of the start codon, using this feature. We report here our attempt to extend the model into the total transcriptional unit, enabling the prediction of operon structures. RESULTS: First, we improved the prediction accuracy of coding sequences (CDSs) by employing the models of 'typical', 'atypical' and 'negative (false-positive)' classes as well as the models of RBS and its downstream spacer. The sensitivity of exactly predicting the 204 experimentally confirmed CDSs reached 90.2% in an objective test. Based on the prediction result of CDSs, the positions of the promoters and terminators were predicted. Our model could exactly recognize 60% of 390 known transcriptional units. Thus, the accuracy and significance of this prediction problem is far from trivial. We would like to propose this problem as an open theme in bioinformatics because the ongoing or planned post-sequencing projects will produce much data for future improvements.

Algorithms↗

GenoMiner: a tool for genome-wide search of coding and non-coding conserved sequence tags.

GenoMiner is a software tool that searches for regions of similarity between user-submitted genome or transcript sequences and user-specified whole genome assemblies. The program then identifies conserved sequence tags (CSTs) in these homologous regions and provides a prediction of their coding or non-coding nature. The analysis is carried out through three steps: (1) definition of sequence regions homologous to the query sequence in the selected target genomes by a fast BLAT alignment; (2) identification of CSTs by a more sensitive BLAST-like alignment between the query and the homologous regions in the target genomes and (3) assessment of the coding or non-coding nature of detected CSTs through the computation of a suitable coding potential score. GenoMiner allows the user to search the query sequence against a number of vertebrate genome assemblies in a single run providing a user-friendly graphical output.

Algorithms↗

A study to determine the sensitivity and specificity of hospital discharge diagnosis data used in the MICA study.

AIMS: To determine the sensitivity and specificity of each ICD9 code for a diagnosis of definite or possible myocardial infarction (MI) from the perspective of the Myocardial Infarction Causality Study (MICA) and to use these data to estimate the likely number of MICA cases in Scotland that would be undetected were these codes omitted from the study. SETTING: Women resident and registered with general practitioners in the Tayside region of Scotland between October 1993 and October 1995. METHOD: All SMR1 records of Tayside hospitalizations containing ICD9 (International Classification of Diseases, ninth revision) codes for myocardial infarction (410) or possible myocardial infarction (411, 412, 413, 414, 427.4, 427.5, 786.5) were identified for women aged between 16 and 44 years between 1 October 1993 and 15 October 1995. Original case records were sought and each episode abstracted using a predefined form. Records were independently scrutinized by two consultant cardiologists blinded to the SMR1 code. Cases were categorized as definite MI, possible MI or unlikely MI. Where there was disagreement between the two cardiologists, the profiles for such events were examined by a third cardiologist who acted as the final adjudicator. The adjudicator's verdict was, in this study, considered dominant. The sensitivity, specificity and positive predictive value of each ICD9 code was determined. RESULTS: Two hundred and fifty-three women fulfilled the SMR1 search criteria. Case records of 204 (81%) were retrieved but four case records contained no data on the admission of interest and were classified as invalid. Forty-six of the 200 remaining patients were ineligible for the MICA study leaving 154 records for evaluation. There were 12 patients who had a discharge code for MI (ICD9 410). Of these, 11 were judged as a definite MI by both cardiologists. One event (discharge code ICD9 410) was judged as 'possible' by one cardiologist and 'unlikely' by the other. The adjudicator subsequently judged this event as 'definite'. Another six events were subsequently judged as 'possible'. Thus, after adjudication, 12 cases of definite MI and six cases of 'possible' MI were identified. The sensitivity and specificity of ICD9 code 410 was 67% and 100% respectively. The positive predictive value was 100%. The sensitivity of code 411 was 5.6%. The specificity was 99% and the positive predictive value was 50%. Code 413 had a sensitivity of 5.6% with a specificity of 94% and a positive predictive value of 9.1%. Code 414 also had a sensitivity of 5.6%. The specificity was 86% and the positive predictive value was 4.5%. Code 786.5 had a sensitivity of 17%, a specificity of 23% and a positive predictive value of 2.5%. Code 427.5 failed to identify any definite or possible cases. CONCLUSIONS: In the MICA Study, ICD9 code 410 was found to be the most robust. All 12 patients judged to have had a definite MI had the appropriate discharge code (ICD9 410). The six patients judged to have had a possible MI all had discharge codes other than that for MI (410). However, identifying these six patients required the validation of a further 160 events-giving a combined sensitivity of 33%, a specificity of 0% and a positive predictive value of only 3.8%. The use of ICD9 codes 411, 413, 414, 427.5 and 786.5 must, therefore, only be employed when circumstances fully justify the additional workload.

Journal Article↗

The genome of the heartwater agent Ehrlichia ruminantium contains multiple tandem repeats of actively variable copy number.

Heartwater, a tick-borne disease of domestic and wild ruminants, is caused by the intracellular rickettsia Ehrlichia ruminantium (previously known as Cowdria ruminantium). It is a major constraint to livestock production throughout subSaharan Africa, and it threatens to invade the Americas, yet there is no immediate prospect of an effective vaccine. A shotgun genome sequencing project was undertaken in the expectation that access to the complete protein coding repertoire of the organism will facilitate the search for vaccine candidate genes. We report here the complete 1,516,355-bp sequence of the type strain, the stock derived from the South African Welgevonden isolate. Only 62% of the genome is predicted to be coding sequence, encoding 888 proteins and 41 stable RNA species. The most striking feature is the large number of tandemly repeated and duplicated sequences, some of continuously variable copy number, which contributes to the low proportion of coding sequence. These repeats have mediated numerous translocation and inversion events that have resulted in the duplication and truncation of some genes and have also given rise to new genes. There are 32 predicted pseudogenes, most of which are truncated fragments of genes associated with repeats. Rather then being the result of the reductive evolution seen in other intracellular bacteria, these pseudogenes appear to be the product of ongoing sequence duplication events.

Base Sequence↗

Nonresponse error in injury-risk surveys.

BACKGROUND: Nonresponse is a potentially serious source of error in epidemiologic surveys concerned with injury control and risk. This study presents the findings of a records-matching approach to investigating the degree to which survey nonresponse may bias indicators of violence-related and unintentional injuries in a random-digit-dialed (RDD) telephone survey. METHODS: Data from a statewide RDD survey of 4155 individuals aged 16 years and older conducted in Illinois in 2003 were merged with ZIP code-level data from the 2000 Census. Using hierarchical linear models, ZIP code-level indicators were used to predict survey response propensity at the individual level. Additional models used the same ZIP code measures to predict a set of injury-risk indicators. RESULTS: Several ZIP code measures were found to be predictive of both response propensity and the likelihood of reporting partner violence. For example, people residing in high-income areas were less likely to participate in the survey and less likely to report forced sex by partner, processes that suggest an over-estimation of this form of violence. In contrast, estimates of partner isolation may be under-estimated, as those residing in geographic areas with smaller-sized housing were less likely to participate in the survey but more likely to report partner isolation. No ZIP code-level correlates of survey response propensity, however, were found also to be associated with driving-under-the-influence (DUI) indicators. CONCLUSIONS: There is evidence of a linkage between survey response propensity and one variety of injury prevention measure (partner violence) but not another (DUI). The approach described in this paper provides an effective and inexpensive tool for evaluating nonresponse error in surveys of injury prevention and other health-related conditions.

Accidents↗

Chicken protamine genes are intronless. The complete genomic sequence and organization of the two loci.

A positive cosmid clone obtained from a pwe15-rooster DNA library using a chicken protamine cDNA probe reveals the complete sequence of the two loci for the rooster protamine genes. The organization of these two loci within the cosmid clone matches that of genomic DNA. The copy number per haploid genome is two. The sequence for the rooster protamine predicted from the coding region shows differences from that previously determined at the protein level (Nakano, M., Tobita, T., and Ando, T. (1976) Int. J. Peptide Protein Res. 8, 565-578). A recent re-determination of the rooster protamine amino acid sequence (28 residues from the N terminus) matches that predicted from the genome rather than the sequence of Nakano et al. (1976). Both loci are intronless and the gene is extremely GC-rich (88% in the coding region). The 5' region of the gene contains a typical TATAAA box, several CG boxes, as well as other characteristic motifs. The 3' region of the gene contains the polyadenylation signal and several GT repeats of known Z-DNA forming potential. A correlation between the functional map of the gene and the tendency of the DNA to bend or to adopt the Z-conformation is presented and possible roles for these conformations in the transcription of this gene are discussed.

Amino Acid Sequence↗

Identification of a gene encoding the predicted ribosomal protein L7b divergently transcribed from POL1 in fission yeast Schizosaccharomyces pombe.

A 0.85 Kb RNA molecule is transcribed in the region upstream from the 5'-end of the S. pombe POL1 gene encoding the catalytic subunit of DNA polymerase alpha. The nucleotide sequence of the DNA region hybridizing with the 0.85 Kb transcript allowed us to identify an open reading frame coding for a predicted peptide which shows 50% identity with the rat ribosomal protein L7 and which is transcribed divergently from POL1. We have named this gene RPL7b because of the existence in S. pombe of a different sequence, named RPL7, which also codes for a putative protein showing homology with the rat ribosomal protein L7. The RPL7b gene includes a 291 bp-long intron containing the sequences necessary for intron excision and RNA splicing in S. pombe. The precise location of the intron was established by amplification and sequencing of a partial cDNA copy of the mRNA, whereas the initiation site of transcription was determined by reverse transcription of the 5' region of the mRNA. The 320 bp separating the starting methionine codons of RPL7b and POL1 genes should contain the signals necessary for their divergent transcription and regulation. The sequence 5'-AAGACAGTCACA-3', whose primary structure is homologous to a conserved block present in the 5'-untranscribed regions of other S. pombe genes of ribosomal proteins, is located about 50 bp upstream the transcription initiation site of RPL7b.

Amino Acid Sequence↗

Precompression quality-control algorithm for JPEG 2000.

In this paper, a precompression quality-control algorithm is proposed. It can greatly reduce computational power of the embedded block coding (EBC) and memory requirement to buffer bit streams. By using the propagation property and the randomness property of the EBC algorithm, rate and distortion of coding passes is approximately predicted. Thus, the truncation points are chosen before actual coding by the entropy coder. Therefore, the computational power, which is measured with the number of contexts to be processed, is greatly reduced since most of the computations are skipped. The memory requirement, which is measured with the amount required to buffer bit streams, is also reduced since the skipped contexts do not generate bit streams. Experimental results show that the proposed algorithm reduces the computational power of the EBC by 80% on average at 0.8 bpp compared with the conventional postcompression rate-distortion optimization algorithm. Moreover, the memory requirement is also reduced by 90%. The average PSNR degrades only about 0.1-0.3 dB, on average.

Algorithms↗

Rate coding model for discrimination of simple tones in the presence of noise.

The predictions of a rate coding model for frequency and amplitude jnd's in the presence of noise are presented for a 1-kHz, 100-ms tone. The model for the neural response incorporates physiological data on dynamic range distribution and rate suppression. A central processor is assumed to estimate the tone frequency, or amplitude, from the tone-evoked rate increment profile. This central processor acts like an ideal detector with respect to the neural noise. The effects of the neural noise as well as the signal variability on the discrimination performance level are evaluated, and the signal variability is found to be significant. The combined effect of threshold distribution, rate suppression, and signal variability make the jnd's practically invariant with noise level, in accordance with published psychophysical data. The values of the frequency jnd at high signal-to-noise ratio, however, are borderline in their consistency with the data. A more obvious discrepancy exists between the model and the psychophysical data regarding the ratio of frequency to amplitude Weber fractions, which can be resolved only by modifying the model auditory filters to be five times sharper than those measured in cats.

Acoustic Stimulation↗

The nucleotide sequence of adenovirus type 5 early region E1: the region between map positions 8.0 (HindIII site) and 11.8 (SmaI site).

The nucleotide sequence of the region between map positions 8.0 (HindIII site) and 11.8 (SmaI site) of adenovirus type 5 (Ad5) has been determined. Together with the sequences reported earlier (Van Ormondt et al., 1978; Maat and Van Ormondt, 1979) it encompasses the entire leftmost early region E1 of Ad5 DNA (4126 base pairs). The total sequence revealed a number of potential regulatory signals (promoter sites, ribosome binding sites, 3'-poly(A)-associated sequences), which confirm that region E1 is divided into subregions, E1a and E1b, and a region coding for semi-late viral protein IX. By taking into account the adenovirus 2 (Ad2) RNA-splicing data of Perricaudet et al. (1979; 1980) and the Ad2 RNA mapping data of Chow et al. (1979) we predict that E1a codes for polypeptides of 32, 26 and ca. 13 kd, and subregion E1b for polypeptides of 67 kd and 20 kd; the expected molecular weight of protein IX is 14.4 kd.

Adenoviruses, Human↗