Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Structural analysis of Arabidopsis thaliana chromosome 5. IX. Sequence features of the regions of 1,011,550 bp covered by seventeen P1 and TAC clones.

In this series of projects sequencing the entire genome of Arabidopsis thaliana chromosome 5, non-redundant P1 and TAC clones have been sequenced according to the fine physical map, and as of May 7, 1999, the sequences of 16.2 Mb representing approximately 60% of chromosome 5 have been accumulated and released at our web site. In parallel, structural features of the sequenced regions have been analyzed by applying a variety of computer programs, and to date we have predicted a total of 2380 potential protein-coding genes in the 10,154,580 bp regions, which are covered by 142 P1 and TAC clones. In this paper, we newly analyzed the structural features of the 1,011,550 bp regions covered by additional 17 P1 and TAC clones, and predicted 298 protein-coding genes. The average density of the genes identified was 1 gene per 3394 bp. Introns were observed in 67% of the genes, and the average number per gene and the average length of the introns were 3.2 and 159 bp, respectively. The gene density became higher than the value estimated in the previously analyzed regions (1 gene per 4,267 bp), as the data in this paper were compiled based on a new standard of gene assignment including the computer-predicted hypothetical genes. The regions also contained 8 tRNA genes when searched by similarity to reported tRNA genes and the tRNA scan-SE program. The sequence data and information on the potential genes are available on the database KAOS (Kazusa Arabidopsis data Opening Site) at http://www.kazusa.or.jp/arabi/.

Arabidopsis↗

Integrating behavior and cardiovascular responses: the code.

The next revolution in biology is predicted to be in the integrative domain, and the need to involve physiologists in this kind of research has been recognized. This paper represents an approach to providing some of the tools required for dealing with integrative physiology at the behavioral level. Video tape recordings are made of the activities of a group of five baboons (Papio hamadryas) while simultaneous recordings of arterial blood pressure, heart rate, renal blood flow, and mesenteric or iliac blood flow are telemetered from two of the members of the group. The telemetered cardiovascular information is recorded on the two audio channels of the videotape. Subsequently the videotape is viewed, and a two-dimensional code is used to record the behavior of the two animals with the telemetry equipment. The first dimension of the code categorizes the behavior changes precisely regarding those aspects of behavior that are related to cardiovascular dynamics and does so with an accuracy of 16 ms. The second dimension codes relevant environmental changes. The paper describes the code and presents illustrations of how the code reflects the cardiovascular dynamics associated with the behavioral changes.

Animals↗

Automated SNP detection from a large collection of white spruce expressed sequences: contributing factors and approaches for the categorization of SNPs.

BACKGROUND: High-throughput genotyping technologies represent a highly efficient way to accelerate genetic mapping and enable association studies. As a first step toward this goal, we aimed to develop a resource of candidate Single Nucleotide Polymorphisms (SNP) in white spruce (Picea glauca [Moench] Voss), a softwood tree of major economic importance. RESULTS: A white spruce SNP resource encompassing 12,264 SNPs was constructed from a set of 6,459 contigs derived from Expressed Sequence Tags (EST) and by using the bayesian-based statistical software PolyBayes. Several parameters influencing the SNP prediction were analysed including the a priori expected polymorphism, the probability score (PSNP), and the contig depth and length. SNP detection in 3' and 5' reads from the same clones revealed a level of inconsistency between overlapping sequences as low as 1%. A subset of 245 predicted SNPs were verified through the independent resequencing of genomic DNA of a genotype also used to prepare cDNA libraries. The validation rate reached a maximum of 85% for SNPs predicted with either PSNP > or = 0.95 or > or = 0.99. A total of 9,310 SNPs were detected by using PSNP > or = 0.95 as a criterion. The SNPs were distributed among 3,590 contigs encompassing an array of broad functional categories, with an overall frequency of 1 SNP per 700 nucleotide sites. Experimental and statistical approaches were used to evaluate the proportion of paralogous SNPs, with estimates in the range of 8 to 12%. The 3,789 coding SNPs identified through coding region annotation and ORF prediction, were distributed into 39% nonsynonymous and 61% synonymous substitutions. Overall, there were 0.9 SNP per 1,000 nonsynonymous sites and 5.2 SNPs per 1,000 synonymous sites, for a genome-wide nonsynonymous to synonymous substitution rate ratio (Ka/Ks) of 0.17. CONCLUSION: We integrated the SNP data in the ForestTreeDB database along with functional annotations to provide a tool facilitating the choice of candidate genes for mapping purposes or association studies.

Algorithms↗

Accuracy of ICD-9 codes in identifying ischemic stroke in the General Hospital of Lugo di Romagna (Italy).

We assessed the sensitivity and the positive predictive value (PPV) of the ICD-9 codes in identifying ischemic strokes. The study involved the cross-sectional comparison between patients with an ischemic stroke diagnosis made by neurologists and patients with the 434 or 436 discharge codes. Sensitivity of the codes (all diagnostic levels and first level respectively) was 82% and 76%; PPV: 71% and 76%. The annual crude incidence of ischemic stroke was 2.62 per 1000 based on verified strokes and 3.03 per 1000 based on 434 or 436 coded medical records (at all diagnostic levels). Thirty-day case fatality ratio was 22.3% in verified strokes and 36.8% among patients diagnosed with codes 434 or 436 but without stroke (all levels). Our results disclosed inaccuracy in use of the ICD-9 codes in the diagnosis of ischemic stroke in the general hospital of Lugo di Romagna, Ravenna Province, Italy. The misdiagnosis of patients could be influenced by the degree of severity of clinical features. Epidemiological data and cost-analysis forecasts based only on the ICD-9 system must be considered with caution.

Cerebrovascular Disorders↗

Hot spot mutations in morphological transforming region II (ORF 79) of cytomegalovirus strains causing disease from bone marrow transplant recipients.

Nested polymerase chain reaction (PCR) amplifying the morphological transforming region II (mtrII) of cytomegalovirus (CMV) has been shown to be useful in the detection of CMV DNA in bone marrow transplant (BMT) recipients. However, there has never been any report on mutation hot spots and subtypes of this open reading frame. Using primers derived from sequences upstream and downstream of mtrII (ORF 79), CMV DNA from peripheral blood leukocytes (PBL) and conventional CMV culture of 16 BMT recipients were amplified by PCR, cloned into pUC118, and sequenced. The amino acid sequences were predicted using the standard triplet code. The DNA sequences obtained from direct amplification of CMV in PBL obtained from the 16 patients were 100% identical to the corresponding ones obtained by amplification of CMV DNA extracted from conventional CMV culture. Within mtrII (ORF 79), hot spot single base mutations were observed at positions +40 (G-->A), +123 (A-->G), +213 (T-->C), and +219 (T-->C). However, because of third base degeneracy, only amino acid 14 was changed from valine to isoleucine in the predicted protein of 13 patients. This corresponded to the hot spot mutation at position +40 (GTC-->ATC), while the rest were silent mutations. An insertion of 3 bases (ACG) was observed in the CMV DNA of 10 patients at positions +91 to +93, leading to a threonine insertion at amino acid 31 in these patients. For patient no. 147 there was a 65 bp deletion in the CMV DNA amplified later in the course of BMT as compared with that early in the course. This gave rise to a frame shift mutation and a change of more than 70% in the predicted amino acid sequence of the protein.

Amino Acid Sequence↗

Evaluation of gene structure prediction programs.

We evaluate a number of computer programs designed to predict the structure of protein coding genes in genomic DNA sequences. Computational gene identification is set to play an increasingly important role in the development of the genome projects, as emphasis turns from mapping to large-scale sequencing. The evaluation presented here serves both to assess the current status of the problem and to identify the most promising approaches to ensure further progress. The programs analyzed were uniformly tested on a large set of vertebrate sequences with simple gene structure, and several measures of predictive accuracy were computed at the nucleotide, exon, and protein product levels. The results indicated that the predictive accuracy of the programs analyzed was lower than originally found. The accuracy was even lower when considering only those sequences that had recently been entered and that did not show any similarity to previously entered sequences. This indicates that the programs are overly dependent on the particularities of the examples they learn from. For most of the programs, accuracy in this test set ranged from 0.60 to 0.70 as measured by the Correlation Coefficient (where 1.0 corresponds to a perfect prediction and 0.0 is the value expected for a random prediction), and the average percentage of exons exactly identified was less than 50%. Only those programs including protein sequence database searches showed substantially greater accuracy. The accuracy of the programs was severely affected by relatively high rates of sequence errors. Since the set on which the programs were tested included only relatively short sequences with simple gene structure, the accuracy of the programs is likely to be even lower when used for large uncharacterized genomic sequences with complex structure. While in such cases, programs currently available may still be of great use in pinpointing the regions likely to contain exons, they are far from being powerful enough to elucidate its genomic structure completely.

Alternative Splicing↗

A tomato cDNA inducible by salt stress and abscisic acid: nucleotide sequence and expression pattern.

We have characterized a new tomato cDNA, TAS14, inducible by salt stress and abscisic acid (ABA). Its nucleotide sequence predicts an open reading frame coding for a highly hydrophilic and glycine-rich (23.8%) protein of 130 amino acids. Southern blot analysis of tomato DNA suggests that there is one TAS14 structural gene per haploid genome. TAS14 mRNA accumulates in tomato seedlings upon treatment with NaCl, ABA or mannitol. It is also induced in roots, stems and leaves of hydroponically grown tomato plants treated with NaCl or ABA. TAS14 mRNA is not induced by other stress conditions such as cold and wounding. The sequence of the predicted TAS14 protein shows four structural domains similar to the rice RAB21, cotton LEA D11 and barley and maize dehydrin genes.

Abscisic Acid↗

Induction of ermC requires translation of the leader peptide.

ermC confers resistance to macrolide-lincosamide streptogramin B antibiotics by specifying a ribosomal RNA methylase, which results in decreased ribosomal affinity for these antibiotics. ermC expression is induced by exposure to erythromycin. We have previously proposed a translational regulation model in which erythromycin causes stalling of a ribosome, which is translating a leader peptide. Stalling causes a conformation shift in the ermC mRNA which in turn unmasks the methylase ribosomal binding site. A prediction of this translational attenuation model for ermC induction was tested by replacing the second codon of the putative ermC leader peptide coding region by TAA. As expected, the introduction of this mutation resulted in an uninducible phenotype which was suppressible by two ochre suppressor mutations in Bacillus subtilis. It is concluded that translation through the leader peptide coding region, in frame with the predicted leader peptide, is required for ermC induction.

Bacterial Proteins↗

Evolution of intron/exon structure of DEAD helicase family genes in Arabidopsis, Caenorhabditis, and Drosophila.

The DEAD box RNA helicase (RH) proteins are homologs involved in diverse cellular functions in all of the organisms from prokaryotes to eukaryotes. Nevertheless, there is a lack of conservation in the splicing pattern in the 53 Arabidopsis thaliana (AtRHs), the 32 Caenorhabditis elegans (CeRHs) and the 29 Drosophila melanogaster (DmRHs) genes. Of the 153 different observed intron positions, 4 are conserved between AtRHs, CeRHs, and DmRHs, and one position is also found in RHs from yeast and human. Of the 27 different AtRH structures with introns, 20 have at least one predicted ancient intron in the regions coding for the catalytic domain. In all of the organisms examined, we found at least one gene with most of its intron predicted to be ancient. In A. thaliana, the large diversity in RH structures suggests that duplications of the ancestral RH were followed by a high number of intron deletions and additions. The very high bias toward phase 0 introns is in favor of intron addition, preferentially in phase 0. Results from this comparative study of the same gene family in a plant and in two animals are discussed in terms of the general mechanisms of gene family evolution.

Animals↗

cpbAg1 encodes an active carboxypeptidase B expressed in the midgut of Anopheles gambiae.

We previously used differential display to identify several Anopheles gambiae genes, whose expression in the mosquito midgut was regulated upon ingestion of Plasmodium falciparum. Here, we report the characterization of one of these genes, cpbAg1, which codes for the first zinc-carboxypeptidase B identified in An. gambiae and in any insect. Expression of cpbAg1 in baculovirus gave rise to an active enzyme, and determination of the N-terminal amino acids confirmed that CPBAg1 contains a signal peptide and a pro-peptide, typical features of digestive zinc carboxypeptidases. cpbAg1 mRNA was mainly produced in the mosquito midgut, where it accumulated in unfed females and was rapidly down-regulated upon blood feeding. Annotation of the An. gambiae genome predicts twenty-three sequences coding for zinc-carboxypeptidases of which only two (cpbAg1 and cpbAg2) are expressed at a significant level in the mosquito midgut.

Amino Acid Sequence↗

Adaptive predictive multiplicative autoregressive model for medical image compression.

In this paper, an adaptive predictive multiplicative autoregressive (APMAR) method is proposed for lossless medical image coding. The adaptive predictor is used for improving the prediction accuracy of encoded image blocks in our proposed method. Each block is first adaptively predicted by one of the seven predictors of the JPEG lossless mode and a local mean predictor. It is clear that the prediction accuracy of an adaptive predictor is better than that of a fixed predictor. Then the residual values are processed by the MAR model with Huffman coding. Comparisons with other methods [MAR, SMAR, adaptive JPEG (AJPEG)] on a series of test images show that our method is suitable for reversible medical image compression.

Angiography↗

Referential coding and attention-shifting accounts of the Simon effect.

Stoffer (1991) and Umiltà and Nicoletti (1992) have proposed an attention-shifting account of the Simon effect. However, Hommel (1993) has presented evidence suggesting that the effect can be explained in terms of referential coding, without invoking attentional shifts. Five experiments are reported here, whose primary purpose is to test implications of the referential-coding account. All of the experiments compared conditions in which a noise stimulus was presented in the position opposite the target stimulus with conditions in which it was not. Contrary to the referential-coding account, (a) the basic Simon effect was larger without a fixation point to serve as a referent than with one; (b) the noise stimulus increased the magnitude of the Simon effect when a fixation point was used, but not when there was no fixation point; and (c) the magnitude of the Simon effect obtained in the presence of a noise stimulus was reduced substantially when the noise and the target (and, if present, the fixation point) were in different colors. The results, although counter to predictions of the referential-coding account, can be accommodated by the attention-shifting account if it is assumed that a fixation point provides an anchor that minimizes attention shifts.

Adult↗

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens↗

The M RNA of impatiens necrotic spot Tospovirus (Bunyaviridae) has an ambisense genomic organization.

The nucleotide sequence of Impatiens necrotic spot virus (INSV) M RNA was determined from cDNA clones. The INSV M RNA was 4972 nucleotides in length with two open reading frames (ORFs) in an ambisense genomic organization. The larger ORF near the 3' end of the viral RNA, coding for a protein with a predicted molecular weight of 124.9 kDa, was in the viral complementary sense and produced the G2 and G1 proteins. A smaller ORF in the viral sense was capable of coding for a 34.1-kDa polypeptide, designated the NSm protein. Two subgenomic RNA species were detected in INSV-infected tissue that corresponded to the predicted sizes (3.3 and 1.0 kb) of the G2-G1 and NSm mRNAs. The ORFs were separated by a 478 nucleotide A-U-rich intergenic region similar to the regions found in other viral RNAs with ambisense ORFs. The intergenic region was predicted to form a stable stem-loop structure (-81.2 kcal/mole). The ambisense genomic organization is characteristic of the S RNA for members of the Phlebovirus, Uukuvirus, and Tospovirus genera in the Bunyaviridae family. This is the first report of an ambisense Bunyaviridae M RNA.

Amino Acid Sequence↗

Molecular cloning and chromosomal localization of the murine homolog of the human helix-loop-helix gene SCL.

The human SCL gene is a member of the family of genes that encode the helix-loop-helix (HLH) class of DNA-binding proteins. A murine SCL cDNA was isolated from a normal macrophage cDNA library by using HLH-specific oligonucleotides as hybridization probes. The coding region is 987 base pairs and encodes a predicted protein of 34 kDa. The nucleotide sequence of the coding region shows 88% identity to the human SCL gene, and the amino acid sequence is 94% identical. The HLH motif and upstream hydrophilic region are entirely conserved in the murine and human proteins. The identity between the mouse and human sequences was less marked in the 5' and 3' untranslated regions. Two murine SCL transcripts that differ in the 3' noncoding region have been detected in fetal liver and various cell lines. Variation was also observed in the 5' untranslated region. Interestingly, immediately downstream of the protein-termination codon, both the human SCL sequence and the murine homolog share an E-box element--the suggested target site for DNA binding of HLH proteins. The murine SCL homolog was mapped to the central part of chromosome 4.

Alleles↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.

The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.

Bacterial Proteins↗

The structural genes for three Drosophila glue proteins reside at a single polytene chromosome puff locus.

The polytene chromosome puff at 68C on the Drosophila melanogaster third chromosome is thought from genetic experiments to contain the structural gene for one of the secreted salivary gland glue polypeptides, sgs-3. Previous work has demonstrated that the DNA included in this puff contains sequences that are transcribed to give three different polyadenylated RNAs that are abundant in third-larval-instar salivary glands. These have been called the group II, group III, and group IV RNAs. In the experiments reported here, we used the nucleotide sequence of the DNA coding for these RNAs to predict some of the physical and chemical properties expected of their protein products, including molecular weight, amino acid composition, and amino acid sequence. Salivary gland polypeptides with molecular weights similar to those expected for the 68C RNA translation products, and with the expected degree of incorporation of different radioactive amino acids, were purified. These proteins were shown by amino acid sequencing to correspond to the protein products of the 68C RNAs. It was further shown that each of these proteins is a part of the secreted salivary gland glue: the group IV RNA codes for the previously described sgs-3, whereas the group II and III RNAs code for the newly identified glue polypeptides sgs-8 and sgs-7.

Amino Acid Sequence↗

SGP-1: prediction and validation of homologous genes based on sequence alignments.

Conventional methods of gene prediction rely on the recognition of DNA-sequence signals, the coding potential or the comparison of a genomic sequence with a cDNA, EST, or protein database. Reasons for limited accuracy in many circumstances are species-specific training and the incompleteness of reference databases. Lately, comparative genome analysis has attracted increasing attention. Several analysis tools that are based on human/mouse comparisons are already available. Here, we present a program for the prediction of protein-coding genes, termed SGP-1 (Syntenic Gene Prediction), which is based on the similarity of homologous genomic sequences. In contrast to most existing tools, the accuracy of depends little on species-specific properties such as codon usage or the nucleotide distribution. may therefore be applied to nonstandard model organisms in vertebrates as well as in plants, without the need for extensive parameter training. In addition to predicting genes in large-scale genomic sequences, the program may be useful to validate gene structure annotations from databases. To this end, SGP-1 output also contains comparisons between predicted and annotated gene structures in HTML format. The program can be accessed via a Web server at http://soft.ice.mpg.de/sgp-1. The source code, written in ANSI C, is available on request from the authors.

Algorithms↗