Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Uncertainties in radiation effect predictions for the natural radiation environments of space.

Future manned missions beyond low earth orbit require accurate predictions of the risk to astronauts and to critical systems from exposure to ionizing radiation. For low-level exposures, the hazards are dominated by rare single-event phenomena where individual cosmic-ray particles or spallation reactions result in potentially catastrophic changes in critical components. Examples might be a biological lesion leading to cancer in an astronaut or a memory upset leading to an undesired rocket firing. The risks of such events appears to depend on the amount of energy deposited within critical sensitive volumes of biological cells and microelectronic components. The critical environmental information needed to estimate the risks posed by the natural space environments, including solar flares, is the number of times more than a threshold amount of energy for an event will be deposited in the critical microvolumes. These predictions are complicated by uncertainties in the natural environments, particularly the composition of flares, and by the effects of shielding. Microdosimetric data for large numbers of orbits are needed to improve the environmental models and to test the transport codes used to predict event rates.

Cosmic Radiation↗

Identification of proteins associated with murine cytomegalovirus virions.

Proteins associated with the murine cytomegalovirus (MCMV) viral particle were identified by a combined approach of proteomic and genomic methods. Purified MCMV virions were dissociated by complete denaturation and subjected to either separation by sodium dodecyl sulfate-polyacrylamide gel electrophoresis and in-gel digestion or treated directly by in-solution tryptic digestion. Peptides were separated by nanoflow liquid chromatography and analyzed by tandem mass spectrometry (LC-MS/MS). The MS/MS spectra obtained were searched against a database of MCMV open reading frames (ORFs) predicted to be protein coding by an MCMV-specific version of the gene prediction algorithm GeneMarkS. We identified 38 proteins from the capsid, tegument, glycoprotein, replication, and immunomodulatory protein families, as well as 20 genes of unknown function. Observed irregularities in coding potential suggested possible sequence errors in the 3'-proximal ends of m20 and M31. These errors were experimentally confirmed by sequencing analysis. The MS data further indicated the presence of peptides derived from the unannotated ORFs ORF(c225441-226898) (m166.5) and ORF(105932-106072). Immunoblot experiments confirmed expression of m166.5 during viral infection.

Amino Acid Sequence↗

Re-evaluation and in silico annotation of the Tupaia herpesvirus proteins.

Herpesviruses represent an exceptionally suitable model to analyze evolutionary old pathogens, their competency to adapt to existing and changing molecular niches in host species, and the modulation of the gene content and function to comply with the requirements of life. The basis for numerous studies dealing with these questions are reliable statements about the gene content of herpesviral genomes and the functions of viral proteins. The recent determination of the coding strategy of the chimpanzee cytomegalovirus genome and the re-evaluation of the gene content of the human cytomegalovirus genome made it also necessary to restructure the putative transcription map of the Tupaia herpesvirus (THV) genome. Twenty-three THV-specific ORFs formerly predicted to be coding for viral proteins were deleted from the THV transcription map resulting in a gene layout that is now characterized by the presence of conserved genes in the genome center, that probably reflect the genome structure of common herpesviral ancestors, and species-specific genes at the termini. The conserved regions in the THV genome are characterized by high G + C contents between 60% and 80%, a high CpG dinucleotide frequency, and the presence of densely packed putative CpG islands. The genome termini seem to provide the requirements of large scale rearrangements and complements of the gene content to adapt to new environmental demands. With the help of the recently designed method of dictionary-driven, pattern-based protein annotation it was possible to assign putative functions to almost all potential THV proteins, e.g. 123 were found to be putative membrane or secreted proteins, putative signal domains were identified in 69, and 29 proteins were predicted to be glycosylated. The present study adds new aspects to the knowledge about the precise gene composition of herpesvirus genomes and viral protein functions that are of exceptional importance for studies dealing with the phylogeny, the evolution, vaccine vector development, virus-host interactions, pathogenesis and the determination of protein functions of herpesviruses.

Animals↗

Complete structure of the chloroplast genome of a legume, Lotus japonicus.

The nucleotide sequence of the entire chloroplast genome (150,519 bp) of a legume, Lotus japonicus, has been determined. The circular double-stranded DNA contains a pair of inverted repeats of 25,156 bp which are separated by a small and a large single copy region of 18,271 bp and 81,936 bp, respectively. A total of 84 predicted protein-coding genes including 7 genes duplicated in the inverted repeat regions, 4 ribosomal RNA genes and 37 tRNA genes (30 gene species) representing 20 amino acids species were assigned on the genome based on similarity to genes previously identified in other chloroplasts. All the predicted genes were conserved among dicot plants except that rpl22, a gene encoding chloroplast ribosomal protein CL22, was missing in L. japonicus. Inversion of a 51-kb segment spanning rbcL to rpsl6 (positions 5161-56,176) in the large single copy region was observed in the chloroplast genome of L. japonicus. The sequence data and gene information are available on our World Wide Web database at http://www.kazusa.or.jp/en/plant/database.html.

Amino Acids↗

Bacterial features in the genome of Methanococcus jannaschii in terms of gene composition and biased base composition in ORFs and their surrounding regions.

As a result of genome projects, the complete nucleotide sequence of the entire genome of an archaeon, Methanococcus jannaschii, was recently determined as well as other complete sequences of bacterial and eucaryal genomes. When all the 1680 predicted protein-coding [corrected] genes of M. jannaschii were classified on the basis of sequence similarity, it was found that this archaeon had a chimeric set of 1016 bacterial-type, 471 eucaryal-type and 193 species- or archaebacteria-specific genes. However, most of the genes predicted to be involved in translation and transcription pathways including RNA genes were of the eucaryal-type with only a few exceptions such as 16S ribosomal RNA and some translation factor-like genes. This appeared curious since previous studies indicated that methanogens have bacterial features in gene organization and expression. To understand the apparent inconsistency between physiological observations and the result of the classification of genes for transcription and translation, we examined the structural relatedness of the genome of M. jannaschii to those of other species. In practice, we compared base compositional patterns in ORFs and their surrounding regions. This made it possible to reveal the relationships among the translation- and transcription-related structures in genomes. In this study, we conducted a statistical test called 'G-test' to evaluate the base biases around the boundaries of ORFs. We then found that M. jannaschii possesses more bacterial features in base biases than eucaryal ones, e.g. strong G biases at the positions corresponding to the Shine-Dalgarno site. This indicates that the few exceptional bacterial genes for translation, such as 16S ribosomal RNA and translation factor-like genes, play crucial roles in the translation pathway in M. jannaschii. The possibility that the genome structure in the last common ancestor of all present species was bacterial is discussed.

Base Composition↗

Color vision in honeybees.

Theoretical and experimental investigations of the color vision system in honeybees are reviewed. Grassmann's model and receptor models of color vision are discussed with respect to the problem of color difference. A recent analysis of the bee's color opponent coding system is presented in brief. Predictions for the spectral sensitivity of color opponent coding neurons derived directly from the color opponent coding (COC) model and predictions for the Bezold-Brücke color shift and the spectral discrimination function derived from the model via color difference formula are presented. The predictions are compared with electrophysiological data and with choice proportions of behavioral experiments, respectively.

Animals↗

Characterization of dull1, a maize gene coding for a novel starch synthase.

The maize dull1 (du1) gene is a determinant of the structure of endosperm starch, and du1- mutations affect the activity of two enzymes involved in starch biosynthesis, starch synthase II (SSII) and starch branching enzyme IIa (SBEIIa). Six novel du1- mutations generated in Mutator-active plants were identified. A portion of the du1 locus was cloned by transposon tagging, and a nearly full-length Du1 cDNA sequence was determined. Du1 codes for a predicted 1674-residue protein, comprising one portion that is similar to SSIII of potato, as well as a large unique region. Du1 transcripts are present in the endosperm during the time of starch biosynthesis, but the mRNA was undetectable in leaf or root tissue. The predicted size of the Du1 gene product and its expression pattern are consistent with those of maize SSII. The Du1 gene product contains two repeated regions in its unique N terminus. One of these contains a sequence identical to a conserved segment of SBEs. We conclude that Du1 codes for a starch synthase, most likely SSII, and that secondary effects of du1- mutations, such as reduction of SBEIIa, result from the primary deficiency in this starch synthase.

Amino Acid Sequence↗

Recognition of protein coding regions in DNA sequences.

We give a test for protein coding regions which is based on simple and universal differences between protein-coding and noncoding DNA. The test is simple enough to use without a computer and is completely objective. The test has been thoroughly proven on 400,000 bases of sequence data: it misclassifies 5% of the regions tested and gives an answer of "No Opinion" one fifth of the time. We predict some new coding and noncoding regions in published sequences.

Computers↗

Cloning of a complementary DNA coding for the 100-kD antigenic protein of the PM-Scl autoantigen.

Anti-PM-Scl antibodies are associated with polymyositis-scleroderma overlap or either disease alone. Among sera from 39 patients with anti-PM-Scl, 23 recognized the 100-kD band in immunoblot against HeLa cell extract, 16 of which also stained the 70-kD band. A human thymocyte lambda gt11 cDNA expression library was screened with anti-PM-Scl serum, and two clones were identified whose products reacted with 33 and 37 of 39 anti-PM-Scl sera, respectively, but none of 26 negative control sera. Affinity-purified antibody reacting specifically with plaques of the clone stained the 100-kD band on immunoblot, reacted with nucleoli of HEp-2 cells, and immunoprecipitated the PM-Scl protein complex. Partial sequences of both inserts were identical. One insert was fully sequenced, and additional 5' and 3' sequence was obtained using a gene-specific primer to form a cDNA with HeLa cell RNA as template followed by PCR. The complete nucleotide sequence included 2,739-bp coding for a predicted full-length protein of 98,088 D. There was no homology with the PM-Scl 75-kD protein and no significant homology with other proteins. A mixed-charge cluster was identified, with 22 charged amino acids of 37. In conclusion, the full-length cDNA sequence was determined coding for the PM-Scl 100-kD protein, the most commonly antigenic protein of the PM-Scl complex.

Amino Acid Sequence↗

Generation of diversity in nonerythroid spectrins. Multiple polypeptides are predicted by sequence analysis of cDNAs encompassing the coding region of human nonerythroid alpha-spectrin.

Nonerythroid alpha-spectrin (alpha-fodrin) is a major component of the membrane skeleton in diverse cell types. Overlapping cDNAs have been isolated which encompass the coding region of human lung fibroblast nonerythroid alpha-spectrin. The composite sequence of 7,787 nucleotides encodes a polypeptide of 2,472 amino acids (predicted Mr of 283,964). This sequence has 58% amino acid identity with human erythroid alpha-spectrin, which is encoded on a different gene, and 96% amino acid identity with the full-length sequence of chicken brain alpha-spectrin. We previously reported the variable expression in human fibroblast alpha-spectrin of 20 amino acids between repeats 10 and 11 (McMahon, A. P., Giebelhaus, D. H., Champion, J. E., Bailes, J. A., Lacey, S., Carritt, B., Henchman, S. K., and Moon, R. T. (1987) Differentiation 34, 68-78). In this study, we report additional heterogeneity in fibroblast alpha-spectrin near the carboxyl-terminal end. One of the fibroblast cDNAs (clone 3D) has an in-frame deletion of 18 nucleotides within spectrin repeat 21 when compared to an overlapping fibroblast cDNA (clone 7). As this heterogeneity in amino acid sequence occurs near domains of nonerythroid alpha-spectrin suggested to bind calcium or actin, it is possible that fibroblasts express functionally distinct isoforms of nonerythroid alpha-spectrin.

Amino Acid Sequence↗

Structure and expression of the murine L-myc gene.

We have isolated a 12 kb clone from the murine genome which we show by DNA transfection studies to contain an entire functional L-myc gene and the transcriptional promoter sequences necessary for its expression. We have also isolated a 3.1 kb cDNA sequence from a murine brain cDNA library which corresponds to most of the L-myc mRNA. We have identified the L-myc coding region within the genomic clone by a combination of S1 nuclease analyses. Northern blotting analyses and comparative nucleotide sequence analyses with the cDNA clone. The L-myc gene appears to be organized similarly to the other well-characterized myc-family genes, c-myc and N-myc. The predicted amino acid coding sequence of the L-myc gene indicates that the L-myc protein is significantly smaller than c- and N-myc, but is highly related. In particular, comparison of the N- and c-myc protein sequences reveals seven relatively conserved regions interspersed among non-conserved regions; the L-myc gene retains five of these conserved regions but lacks two others. In addition, a portion of one highly conserved region is encoded within a different region of the L-myc gene but, due to changes in the size of L-myc exons relative to those of N- and c-myc, maintains its overall position in the peptide backbone with respect to other conserved regions. We discuss these findings in the context of potential functional domains and the possibility of overlapping and distinct activities of myc-family proteins.

Amino Acid Sequence↗

Macronuclear genome sequence of the ciliate Tetrahymena thermophila, a model eukaryote.

The ciliate Tetrahymena thermophila is a model organism for molecular and cellular biology. Like other ciliates, this species has separate germline and soma functions that are embodied by distinct nuclei within a single cell. The germline-like micronucleus (MIC) has its genome held in reserve for sexual reproduction. The soma-like macronucleus (MAC), which possesses a genome processed from that of the MIC, is the center of gene expression and does not directly contribute DNA to sexual progeny. We report here the shotgun sequencing, assembly, and analysis of the MAC genome of T. thermophila, which is approximately 104 Mb in length and composed of approximately 225 chromosomes. Overall, the gene set is robust, with more than 27,000 predicted protein-coding genes, 15,000 of which have strong matches to genes in other organisms. The functional diversity encoded by these genes is substantial and reflects the complexity of processes required for a free-living, predatory, single-celled organism. This is highlighted by the abundance of lineage-specific duplications of genes with predicted roles in sensing and responding to environmental conditions (e.g., kinases), using diverse resources (e.g., proteases and transporters), and generating structural complexity (e.g., kinesins and dyneins). In contrast to the other lineages of alveolates (apicomplexans and dinoflagellates), no compelling evidence could be found for plastid-derived genes in the genome. UGA, the only T. thermophila stop codon, is used in some genes to encode selenocysteine, thus making this organism the first known with the potential to translate all 64 codons in nuclear genes into amino acids. We present genomic evidence supporting the hypothesis that the excision of DNA from the MIC to generate the MAC specifically targets foreign DNA as a form of genome self-defense. The combination of the genome sequence, the functional diversity encoded therein, and the presence of some pathways missing from other model organisms makes T. thermophila an ideal model for functional genomic studies to address biological, biomedical, and biotechnological questions of fundamental importance.

Animals↗

Therapist alliance-building behavior within a cognitive-behavioral treatment for anxiety in youth.

Explored the specific behavior of therapists contributing to a child client's perception of a therapeutic alliance with youth (n = 56) who received a manualized cognitive-behavioral treatment for anxiety disorders. The first 3 sessions were coded for 11 therapist behaviors hypothesized to predict ratings of alliance. Child, therapist, and observer alliance ratings were gathered after the 3rd and 7th therapy sessions. "Collaboration" positively predicted early child ratings of alliance, and "finding common ground" and "pushing the child to talk" negatively predicted early child ratings of alliance. Although no coded therapist behaviors predicted early therapist ratings of alliance, "collaboration" and "not being overly formal" positively predicted therapist alliance ratings by Session 7. Child, observer, and therapist ratings of alliance were significantly correlated. Results are discussed with regard to the identified behavior of the therapist as a step toward the identification of empirically supported strategies for building a stronger child-therapist alliance.

Adolescent↗

Partial nucleotide sequence of the Murray Valley encephalitis virus genome. Comparison of the encoded polypeptides with yellow fever virus structural and non-structural proteins.

The sequence of 5400 bases corresponding to the 5'-terminal half of the Murray Valley encephalitis virus genome has been determined. The genome contains a 5' non-coding region of about 97 nucleotides, followed by a single continuous open reading frame that encodes the structural proteins followed by the non-structural proteins. Amino acid sequence homology between the Murray Valley encephalitis and yellow fever (Rice et al., 1985) polyproteins is 42% over the region sequenced. The start points of the various Murray Valley encephalitis virus-coded proteins have been assigned on the basis of this homology and a consistent set of potential proteolytic cleavage sites identified, the sequences of which are similar in Murray Valley encephalitis and yellow fever. The deduced Murray Valley encephalitis gene order is 5'-C-prM (M)-E-NS1-ns2a-ns2b-NS3-3'. The genome organization of Murray Valley encephalitis and yellow fever appears to be identical and the sizes of the predicted virus-coded proteins similar between the two viruses. Both viruses encode a basic capsid protein followed by three glycoproteins; the glycoproteins appear to have the conventional topology of N terminus outside with a C-terminal membrane-spanning domain. There are conserved glycosylation sites in prM, the precursor to the M protein of the virion, and in NS1, a non-structural protein of uncertain function. The glycosylation sites in E, the major envelope protein of the virion, are not conserved as to position. We predict the existence, in flavivirus-infected cells, of two small, hydrophobic peptides, ns2a and ns2b, which show only limited amino acid sequence homology. Finally, about half of the amino acid sequence of NS3 has been obtained; NS3 is a hydrophilic non-structural protein that shows 55% amino acid sequence similarity between Murray Valley encephalitis and yellow fever over the region sequenced and is probably involved in RNA replication.

Amino Acid Sequence↗

Concerted evolution at a multicopy locus in the protozoan parasite Theileria parva: extreme divergence of potential protein-coding sequences.

Concerted evolution of multicopy gene families in vertebrates is recognized as an important force in the generation of biological novelty but has not been documented for the multicopy genes of protozoa. A multicopy locus, Tpr, which consists of tandemly arrayed open reading frames (ORFs) containing several repeated elements has been described for Theileria parva. Herein we show that probes derived from the 5'/N-terminal ends of ORFs in the genomic DNAs of T. parva Uganda (1,108 codons) and Boleni (699 codons) hybridized with multicopy sequences in homologous DNA but did not detect similar sequences in the DNA of 14 heterologous T. parva stocks and clones. The probe sequences were, however, protein coding according to predictive algorithms and codon usage. The 3'/C-terminal ends of the Uganda and Boleni ORFs exhibited 75% similarity and identity, respectively, to the previously identified Tpr1 and Tpr2 repetitive elements of T. parva Muguga. Tpr1-homologous sequences were detected in two additional species of Theileria. Eight different Tpr1-homologous transcripts were present in piroplasm mRNA from a single T. parva Muguga-infected animal. The Tpr1 and Tpr2 amino acid sequences contained six predicted membrane-associated segments. The ratio of synonymous to nonsynonymous substitutions indicates that Tpr1 evolves like protein-encoding DNA. The previously determined nucleotide sequence of the gene encoding the p67 antigen is completely identical in T. parva Muguga, Boleni, and Uganda, including the third base in codons. The data suggest that concerted evolution can lead to the radical divergence of coding sequences and that this can be a mechanism for the generation of novel genes.

Amino Acid Sequence↗

Complementary DNA cloning of a protein highly homologous to mammalian sarcoplasmic reticulum Ca-ATPase from the crustacean Artemia.

Complementary DNA clones coding for an Artemia ATPase have been isolated using an oligonucleotide probe for a region highly conserved between P-type ATPases. The nucleotide sequence of three overlapping clones, 3309 base-pairs, has been established. This sequence includes 78 nucleotides of 5' untranslated sequence, an open reading frame of 3009 nucleotides and 222 nucleotides of 3' untranslated sequences. The amino acid sequence predicted for the coding region is 71% similar to that of slow and fast twitch rabbit muscle sarcoplasmic reticulum Ca-ATPases. The homology is specially high in some regions of the protein that include the previously described regions that are similar between all known P-type ATPases, as well as transmembrane domains and intra- and extracellular domains adjacent to the membrane that are not conserved in P-type ATPases but have been proposed to be involved in calcium binding and transport in rabbit sarcoplasmic reticulum Ca-ATPases. Probes of this likely sarcoplasmic reticulum Ca-ATPase hybridize to two mRNAs of 5200 and 4500 bases. Although both mRNAs are already present in cryptobiotic embryos, the levels of the 5200 base mRNA decrease after development is reassumed, being undetectable after hatching of the nauplii. The levels of the 4500 base mRNA increase during development; maximal levels are reached by ten hours and are maintained at later stages of development.

Amino Acid Sequence↗

From first base: the sequence of the tip of the X chromosome of Drosophila melanogaster, a comparison of two sequencing strategies.

We present the sequence of a contiguous 2.63 Mb of DNA extending from the tip of the X chromosome of Drosophila melanogaster. Within this sequence, we predict 277 protein coding genes, of which 94 had been sequenced already in the course of studying the biology of their gene products, and examples of 12 different transposable elements. We show that an interval between bands 3A2 and 3C2, believed in the 1970s to show a correlation between the number of bands on the polytene chromosomes and the 20 genes identified by conventional genetics, is predicted to contain 45 genes from its DNA sequence. We have determined the insertion sites of P-elements from 111 mutant lines, about half of which are in a position likely to affect the expression of novel predicted genes, thus representing a resource for subsequent functional genomic analysis. We compare the European Drosophila Genome Project sequence with the corresponding part of the independently assembled and annotated Joint Sequence determined through "shotgun" sequencing. Discounting differences in the distribution of known transposable elements between the strains sequenced in the two projects, we detected three major sequence differences, two of which are probably explained by errors in assembly; the origin of the third major difference is unclear. In addition there are eight sequence gaps within the Joint Sequence. At least six of these eight gaps are likely to be sites of transposable elements; the other two are complex. Of the 275 genes in common to both projects, 60% are identical within 1% of their predicted amino-acid sequence and 31% show minor differences such as in choice of translation initiation or termination codons; the remaining 9% show major differences in interpretation.

Animals↗

Prediction of protein subcellular locations by support vector machines using compositions of amino acids and amino acid pairs.

MOTIVATION: The subcellular location of a protein is closely correlated to its function. Thus, computational prediction of subcellular locations from the amino acid sequence information would help annotation and functional prediction of protein coding genes in complete genomes. We have developed a method based on support vector machines (SVMs). RESULTS: We considered 12 subcellular locations in eukaryotic cells: chloroplast, cytoplasm, cytoskeleton, endoplasmic reticulum, extracellular medium, Golgi apparatus, lysosome, mitochondrion, nucleus, peroxisome, plasma membrane, and vacuole. We constructed a data set of proteins with known locations from the SWISS-PROT database. A set of SVMs was trained to predict the subcellular location of a given protein based on its amino acid, amino acid pair, and gapped amino acid pair compositions. The predictors based on these different compositions were then combined using a voting scheme. Results obtained through 5-fold cross-validation tests showed an improvement in prediction accuracy over the algorithm based on the amino acid composition only. This prediction method is available via the Internet.

Algorithms↗