Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Maps, codes, and sequence elements: can we predict the protein output from an alternatively spliced locus?

Alternative splicing choices are governed by splicing regulatory protein interactions with splicing silencer and enhancer elements present in the pre-mRNA. However, the prediction of these choices from genomic sequence is difficult, in part because the regulators can act as either enhancers or silencers. A recent study describes how for a particular neuronal splicing regulatory protein, Nova, the location of its binding sites is highly predictive of the protein's effect on an exon's splicing.

Alternative Splicing↗

Stage-specific expression of a Plasmodium falciparum protein related to the eukaryotic mitogen-activated protein kinases.

We have identified a putative protein kinase gene from both Plasmodium falciparum cDNA and genomic DNA libraries. The nucleotide sequence contains an open-reading frame of 2646 bp, which codes for a predicted protein of 882 amino acid residues. Comparison of the predicted amino acid sequence with those in GenBank suggests that this gene codes for a protein similar to the mitogen-activated protein (MAP) kinase of other organisms. This MAP kinase-related protein, named PfMRP, contains the TDY dual phosphorylation site upstream of the highly conserved VATRWYRAPE sequence in subdomain VIII. PfMRP contains an unusually large and highly charged domain within its carboxyl-terminal segment, which includes two repetitive sequences of either a tetrapeptide or octapeptide motif. PfMRP gene is located on chromosome 14. Northern blot analysis of total RNA reveals the presence of a single mRNA transcript approximately 4.2 kb in length, which is predominantly expressed in gametocytes and gametes/zygotes.

Amino Acid Sequence↗

Exploration and grading of possible genes from 183 bacterial strains by a common protocol to identification of new genes: Gene Trek in Prokaryote Space (GTPS).

A large number of complete microorganism genomes has been sequenced and submitted to the public database and then incorporated into our complete genome database, Genome Information Broker (GIB, http://gib.genes.nig.ac.jp/). However, when comparative genomics is carried out, researchers must be aware that there are protein-coding genes not confirmed by homology or motif search and that reliable protein-coding genes are missing. Therefore, we developed a protocol (Gene Trek in Prokaryote Space, GTPS) for finding possible protein-coding genes in bacterial genomes. GTPS assigns a degree of reliability to predicted protein-coding genes. We first systematically applied the protocol to the complete genomes of all 123 bacterial species and strains that were publicly available as of July 2003, and then to those of 183 species and strains available as of September 2004. We found a number of incorrect genes and several new ones in the genome data in question. We also found a way to estimate the total number of orthologous genes in the bacterial world.

Bacteria↗

Nucleotide sequence of the pilin gene of Bacteroides nodosus 340 (serogroup D) and implications for the relatedness of serogroups.

The gene encoding pilin of Bacteroides nodosus 340 has been isolated and the nucleotide sequence determined. The gene is present as a single copy within the B. nodosus genome and a protein of Mr 16683 can be predicted from the proposed coding region. A comparison of the predicted amino acid sequence with pilin from other strains of B. nodosus indicated that the protein of strain 340 (serogroup D) has a high degree of similarity with pilin of strain 265 (serogroup H). The degree of similarity between pilins from these strains and from other B. nodosus serogroups is no greater than that between B. nodosus pilins and the homologous proteins of several different bacterial species. These findings suggest that serogroups D and H may form a subset of B. nodosus serogroups.

Amino Acid Sequence↗

Independent coding of movement direction and reward prediction by single pallidal neurons.

Associating action with its reward value is a basic ability needed by adaptive organisms and requires the convergence of limbic, motor, and associative information. To chart the basal ganglia (BG) involvement in this association, we recorded the activity of 61 well isolated neurons in the external segment of the globus pallidus (GPe) of two monkeys performing a probabilistic visuomotor task. Our results indicate that most (96%) neurons responded to multiple phases of the task. The activity of many (34%) pallidal neurons was modulated solely by direction of movement, and the activity of only a few (3%) pallidal neurons was modulated exclusively by reward prediction. However, the activity of a large number (41%) of single pallidal neurons was comodulated by both expected trial outcome and direction of arm movement. The information carried by the neuronal activity of single pallidal neurons dynamically changed as the trial progressed. The activity was predominantly modulated by both outcome prediction and future movement direction at the beginning of trials and became modulated mainly by movement-direction toward the end of trials. GPe neurons can either increase or decrease their discharge rate in response to predicted future reward. The effects of movement-direction and reward probability on neural activity are linearly summed and thus reflect two independent modulations of pallidal activity. We propose that GPe neurons are uniquely suited for independent processing of a multitude of parameters. This is enabled by the funnel-structure characteristic of the BG architecture, as well as by the anatomical and physiological properties of GPe neurons.

Action Potentials↗

Gene prediction by spectral rotation measure: a new method for identifying protein-coding regions.

A new measure for gene prediction in eukaryotes is presented. The measure is based on the Discrete Fourier Transform (DFT) phase at a frequency of 1/3, computed for the four binary sequences for A, T, C, and G. Analysis of all the experimental genes of S. cerevisiae revealed distribution of the phase in a bell-like curve around a central value, in all four nucleotides, whereas the distribution of the phase in the noncoding regions was found to be close to uniform. Similar findings were obtained for other organisms. Several measures based on the phase property are proposed. The measures are computed by clockwise rotation of the vectors, obtained by DFT for each analysis frame, by an angle equal to the corresponding central value. In protein coding regions, this rotation is assumed to closely align all vectors in the complex plane, thereby amplifying the magnitude of the vector sum. In noncoding regions, this operation does not significantly change this magnitude. Computing the measures with one chromosome and applying them on sequences of others reveals improved performance compared with other algorithms that use the 1/3 frequency feature, especially in short exons. The phase property is also used to find the reading frame of the sequence.

Chromosomes, Fungal↗

Reannotation of Shewanella oneidensis genome.

As more and more complete bacterial genome sequences become available, the genome annotation of previously sequenced genomes may become quickly outdated. This is primarily due to the discovery and functional characterization of new genes. We have reannotated the recently published genome of Shewanella oneidensis with the following results: 51 new genes have been identified, and functional annotation has been added to the 97 genes, including 15 new and 82 existing ones with previously unassigned function. The identification of new genes was achieved by predicting the protein coding regions using the HMM-based program GeneMark.hmm. Subsequent comparison of the predicted gene products to the non-redundant protein database using BLAST and the COG (Clusters of Orthologous Groups) database using COGNITOR provided for the functional annotation.

Algorithms↗

Cloning and nucleotide sequence analysis of the dog insulin gene. Coded amino acid sequence of canine preproinsulin predicts an additional C-peptide fragment.

A 4.0-kilobase HindIII/EcoRI-cleaved dog genomic DNA fragment was shown to contain the dog insulin gene by restriction mapping using a human insulin cDNA probe. This fragment was subsequently cloned in a lambda vector, and the nucleotide sequence of the dog insulin gene was determined. As in several other species, the insulin gene of the dog is interrupted by two intervening sequences, one of 151 base pairs located in the 5' untranslated region and the other of 264 base pairs occurring within the codon of the 7th amino acid of the C-peptide. Translation of the nucleotide sequence in one frame revealed the primary structure of canine preproinsulin. An interesting feature of the coded amino acid sequence is that it predicts a C-peptide of 31 amino acids, 8 residues longer than that reported by Peterson et al. (Peterson, J. D., Nehrlich, S., Oyer, P. E., and Steiner, D. F. (1973) J. Biol. Chem. 247, 4866-4871). The additional octapeptide sequence, Glu-Val-Glu-Asp-Leu-Gln-Val-Arg, is located NH2-terminal to the 23-residue C-peptide sequence described in the earlier report. Its coding sequence is interrupted by the second intervening sequence. The arginine at position 8 suggests that a trypsin-like cleavage may separate the NH2-terminal octapeptide from the remainder of the C-peptide during the post-translational processing of dog proinsulin in the pancreas. The revised C-peptide sequence suggests that the proinsulin C-peptide is more highly conserved in length and overall sequence than was previously supposed.

Amino Acid Sequence↗

Cofolga: a genetic algorithm for finding the common folding of two RNAs.

In order to predict non-coding RNA genes and functions on the basis of genome sequences, accurate secondary structure prediction is useful. Although single-sequence folding programs such as mfold have been successful, it is of great importance to develop a novel approach for further improvement of the prediction performance. In the present paper, a secondary structure prediction method based on genetic algorithm, Cofolga, is proposed. The program developed performs folding and alignment of two homologous RNAs simultaneously. Cofolga was tested with a dataset composed of 13 tRNAs, seven 5S rRNAs, five RNase P RNAs, and five SRP RNAs; as a result, it turned out that the average prediction accuracies for the tRNAs, 5S rRNAs, RNase P RNAs, and SRP RNAs obtained by Cofolga with an optimal weight factor and default parameters were 83.6, 81.8, 73.5, and 67.7%, respectively. These results were superior to those obtained by a single-sequence folding based on free-energy minimization in which corresponding average prediction accuracies were 52.4, 47.4, 57.7, and 52.3%, respectively. Cofolga has a post-processing in which a single-sequence folding is performed after fixation of a predicted common structure; this post-processing enables Cofolga to predict a structure that is present in one of two RNAs alone. The executable files of Cofolga (for Windows/Unix/Mac) can be obtained by an e-mail request.

Algorithms↗

The hemoglobin of Urechis caupo. The cDNA-derived amino acid sequence.

The nucleotide sequence of a cDNA transcript containing part of the 5' noncoding region, the entire coding region, and the entire 3' noncoding region has been determined. The protein sequence predicted from the coding region matches almost exactly the aminoterminal sequence and the sequence of several peptides from Urechis caupo F-I globin. Only 11-20% of the amino acid positions are identical with those of other known globins.

Amino Acid Sequence↗

Leveraging functional annotations to map rare variants associated with Alzheimer disease with gruyere.

Increased availability of whole-genome sequencing (WGS) has facilitated the study of rare variants (RVs) in complex diseases. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most do not fully leverage the availability of variant-level functional annotations. We propose genome-wide rare variant enrichment evaluation (gruyere), an empirical Bayesian framework that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimer's Disease Sequencing Project to identify Alzheimer disease (AD)-associated genes and annotations. Growing evidence suggests that the disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not examined rare non-coding effects that incorporate such cell-type-specific information. To address this gap, we (1) define per-gene non-coding RV test sets using predicted enhancer and promoter regions in microglia and other brain cell types (oligodendrocytes, astrocytes, and neurons) and (2) include cell-type-specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 13 significant genetic associations not detected by other RV methods, four of which remain significant in omnibus tests. We find that deep-learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a robust framework incorporating functional annotations, coding RVs, and cell-type-associated non-coding RVs to perform genome-wide association tests, uncovering AD-relevant genes and annotations.

Alzheimer Disease↗

The uvsF gene region in Aspergillus nidulans codes for a protein with homology to DNA replication factor C.

The UV-sensitive mutant uvsF201 of Aspergillus nidulans shows increased spontaneous and UV-induced mutation and generally resembles mutants defective in nucleotide excision repair (NER). Fully-complementing uvsF clones were isolated from cosmid and cDNA libraries for sequencing. The uvsF gene is approximately 3.75 kb long and codes for a predicted polypeptide of 1092 amino acids (aa). Three small introns are clustered early in the coding region of the protein. A major part of the sequence shows homology to human, mouse and yeast RFC1 genes which code for the large subunit of the DNA replication factor C. The uvsF gene product may therefore function primarily in general DNA replication but in addition be required for the replication step of DNA repair. Extended sequencing of the uvsF gene region identified a second closely adjacent gene of unknown function which is divergently transcribed from a small (0.2 kb) intergenic promoter region.

Amino Acid Sequence↗

A 10,400-molecular-weight membrane protein is coded by region E3 of adenovirus.

Previous studies with adenovirus mutants have indicated that a 10,400-molecular-weight (10.4K) protein predicted to be coded by an open reading frame in region E3 of adenovirus functions to down regulate the epidermal growth factor receptor (C. R. Carlin, A. E. Tollefson, H. A. Brady, B. L. Hoffman, and W. S. M. Wold, Cell 57:135-144, 1989). We now demonstrate that the 10.4K protein is in fact synthesized in cells infected by group C adenoviruses. This was done by immunoprecipitation of 10.4K from cells infected by a variety of E3 mutants, using antisera against three different synthetic peptides corresponding to the predicted 10.4K sequence. The 10.4K protein was translated primarily from E3 mRNA f, as indicated by cell-free translation of mRNA purified by hybridization from cells infected with an RNA processing mutant that synthesizes predominantly mRNA f. The 10.4K protein was overproduced or underproduced in vivo, respectively, by mutants that overproduce or underproduce E3 mRNA f, also indicating that the 10.4K protein is translated primarily from mRNA f. The 10.4K protein migrated as two bands with apparent molecular weights of 16,000 and 11,000 (10 to 18% gradient gels); both bands contained 10.4K epitopes, as shown by Western blot (immunoblot). Only the 16K band was obtained by cell-free translation, suggesting that the 16K protein is the precursor to the 11K protein. The 10.4K protein is a membrane protein, as shown by cell fractionation experiments and as predicted from its sequence. The predicted 10.4K sequence as well as a putative N-terminal signal sequence and 30-residue transmembrane domain are conserved in adenovirus types 2 and 5 (group C) and in types 3, 7, and 35 (group B).

Adenoviruses, Human↗

Annotation of the Drosophila melanogaster euchromatic genome: a systematic review.

BACKGROUND: The recent completion of the Drosophila melanogaster genomic sequence to high quality and the availability of a greatly expanded set of Drosophila cDNA sequences, aligning to 78% of the predicted euchromatic genes, afforded FlyBase the opportunity to significantly improve genomic annotations. We made the annotation process more rigorous by inspecting each gene visually, utilizing a comprehensive set of curation rules, requiring traceable evidence for each gene model, and comparing each predicted peptide to SWISS-PROT and TrEMBL sequences. RESULTS: Although the number of predicted protein-coding genes in Drosophila remains essentially unchanged, the revised annotation significantly improves gene models, resulting in structural changes to 85% of the transcripts and 45% of the predicted proteins. We annotated transposable elements and non-protein-coding RNAs as new features, and extended the annotation of untranslated (UTR) sequences and alternative transcripts to include more than 70% and 20% of genes, respectively. Finally, cDNA sequence provided evidence for dicistronic transcripts, neighboring genes with overlapping UTRs on the same DNA sequence strand, alternatively spliced genes that encode distinct, non-overlapping peptides, and numerous nested genes. CONCLUSIONS: Identification of so many unusual gene models not only suggests that some mechanisms for gene regulation are more prevalent than previously believed, but also underscores the complex challenges of eukaryotic gene prediction. At present, experimental data and human curation remain essential to generate high-quality genome annotations.

Animals↗

A 2-D orientation-adaptive prediction filter in lifting structures for image coding.

Lifting-style implementations of wavelets are widely used in image coders. A two-dimensional (2-D) edge adaptive lifting structure, which is similar to Daubechies 5/3 wavelet, is presented. The 2-D prediction filter predicts the value of the next polyphase component according to an edge orientation estimator of the image. Consequently, the prediction domain is allowed to rotate +/-45 degrees in regions with diagonal gradient. The gradient estimator is computationally inexpensive with additional costs of only six subtractions per lifting instruction, and no multiplications are required.

Algorithms↗

Sequence characterization of the membrane protein gene of paramyxovirus simian virus 5.

The complete nucleotide sequence of the membrane (M) protein gene of the paramyxovirus simian virus 5 (SV5) was determined from cDNA clones of viral mRNAs. The M gene boundaries were determined by (i) primer extension sequencing on M mRNA; (ii) nuclease S1 analysis; and (iii) primer extension sequencing on viral genomic RNA. The M gene mRNA consisted of 1371 templated nucleotides. It contains a single large open reading frame that can encode a protein of 377 amino acids with a predicted Mr = 42,253. The authenticity of the predicted M protein coding sequence was confirmed by synthesis of the M protein from mRNA synthesized from cDNA. The predicted M amino acid sequence indicated it is an overall hydrophobic protein carrying a net positive charge. Alignment of the SV5 protein amino acid sequence with the M protein sequences of other paramyxoviruses indicated that these viruses fall into the following two groups: (1) SV5, mumps virus, and Newcastle disease virus; or (2) Sendai, parainfluenza virus type 3, measles virus, and canine distemper virus, with mumps virus M sequence being the most closely related to SV5.

Amino Acid Sequence↗

Molecular cloning of cDNA and analysis of protein secondary structure of Candida albicans enolase, an abundant, immunodominant glycolytic enzyme.

We isolated and sequenced a clone for Candida albicans enolase from a C. albicans cDNA library by using molecular genetic techniques. The 1.4-kbp cDNA encoded one long open reading frame of 440 amino acids which was 87 and 75% similar to predicted enolases of Saccharomyces cerevisiae and enolases from other organisms, respectively. The cDNA included the entire coding region and predicted a protein of molecular weight 47,178. The codon usage was highly biased and similar to that found for the highly expressed EF-1 alpha proteins of C. albicans. Northern (RNA) blot analysis showed that the enolase cDNA hybridized to an abundant C. albicans mRNA of 1.5 kb present in both yeast and hyphal growth forms. The polypeptide product of the cloned cDNA, which was purified as a recombinant protein fused to glutathione S-transferase, had enolase enzymatic activity and inhibited radioimmunoprecipitation of a single C. albicans protein of molecular weight 47,000. Analysis of the predicted C. albicans enolase showed strong conservation in regions of alpha helices, beta sheets, and beta turns, as determined by comparison with the crystal structure of apo-enolase A of S. cerevisiae. The lack of cysteine residues and a two-amino-acid insertion in the main domain differentiated C. albicans enolase from S. cerevisiae enolase. Immunofluorescence of whole C. albicans cells by using a mouse antiserum generated against the purified fusion protein showed that enolase is not located on the surface of C. albicans. Recombinant C. albicans enolase will be useful in understanding the pathogenesis and host immune response in disseminated candidiasis, since enolase is an immunodominant antigen which circulates during disseminated infections.

Amino Acid Sequence↗

The therapy process observational coding system-alliance scale: measure characteristics and prediction of outcome in usual clinical practice.

The authors describe psychometric characteristics of the new Therapy Process Observational Coding System-Alliance scale (TPOCS-A; B. D. McLeod, 2001) and illustrate its use in the study of treatment as usual. The TPOCS-A uses session observation to assess child-therapist and parent-therapist alliance. Both child and parent forms showed acceptable interrater reliability and internal consistency; when applied to cases treated for internalizing disorders, both forms were associated with youth outcomes. Child-therapist alliance during treatment predicted reduced anxiety symptoms at the end of treatment. Parent-therapist alliance during treatment predicted reduced internalizing, anxiety, and depression symptoms at the end of treatment. The findings held up well after confounding variables were controlled, which suggests that both child-therapist and parent-therapist alliance play key (and potentially different) roles in the outcome of treatment as usual.

Adolescent↗