Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Structural variant of the intergenic internal ribosome entry site elements in dicistroviruses and computational search for their counterparts.

The intergenic region (IGR) located upstream of the capsid protein gene in dicistroviruses contains an internal ribosome entry site (IRES). Translation initiation mediated by the IRES does not require initiator methionine tRNA. Comparison of the IGRs among dicistroviruses suggested that Taura syndrome virus (TSV) and acute bee paralysis virus have an extra side stem loop in the predicted IRES. We examined whether the side stem is responsible for translation activity mediated by the IGR using constructs with compensatory mutations. In vitro translation analysis showed that TSV has an IGR-IRES that is structurally distinct from those previously described. Because IGR-IRES elements determine the translation initiation site by virtue of their own tertiary structure formation, the discovery of this initiation mechanism suggests the possibility that eukaryotic mRNAs might have more extensive coding regions than previously predicted. To test this hypothesis, we searched full-length cDNA databases and whole genome sequences of eukaryotes using the pattern matching program, Scan For Matches, with parameters that can extract sequences containing secondary structure elements resembling those of IGR-IRES. Our search yielded several sequences, but their predicted secondary structures were suggested to be unstable in comparison to those of dicistroviruses. These results suggest that RNAs structurally similar to dicistroviruses are not common. If some eukaryotic mRNAs are translated independently of an initiator methionine tRNA, their structures are likely to be significantly distinct from those of dicistroviruses.

Computational Biology↗

JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.

BACKGROUND: Predicting complete protein-coding genes in human DNA remains a significant challenge. Though a number of promising approaches have been investigated, an ideal suite of tools has yet to emerge that can provide near perfect levels of sensitivity and specificity at the level of whole genes. As an incremental step in this direction, it is hoped that controlled gene finding experiments in the ENCODE regions will provide a more accurate view of the relative benefits of different strategies for modeling and predicting gene structures. RESULTS: Here we describe our general-purpose eukaryotic gene finding pipeline and its major components, as well as the methodological adaptations that we found necessary in accommodating human DNA in our pipeline, noting that a similar level of effort may be necessary by ourselves and others with similar pipelines whenever a new class of genomes is presented to the community for analysis. We also describe a number of controlled experiments involving the differential inclusion of various types of evidence and feature states into our models and the resulting impact these variations have had on predictive accuracy. CONCLUSION: While in the case of the non-comparative gene finders we found that adding model states to represent specific biological features did little to enhance predictive accuracy, for our evidence-based 'combiner' program the incorporation of additional evidence tracks tended to produce significant gains in accuracy for most evidence types, suggesting that improved modeling efforts at the hidden Markov model level are of relatively little value. We relate these findings to our current plans for future research.

Computational Biology↗

Adaptive evolution of HoxA-11 and HoxA-13 at the origin of the uterus in mammals.

The evolution of morphological characters is mediated by the evolution of developmental genes. Evolutionary changes can either affect cis-regulatory elements, leading to differences in their temporal and spatial regulation, or affect the coding region. Although there is ample evidence for the importance of cis-regulatory evolution, it has only recently been shown that transcription factors do not remain functionally equivalent during evolution. These results suggest that the evolution of transcription factors may play an active role in the evolution of development. To test this idea we investigated the molecular evolution of two genes essential for the development and function of the mammalian female reproductive organs, HoxA-11 and HoxA-13. We predicted that if coding-region evolution plays an active role in developmental evolution, then these genes should have experienced adaptive evolution at the origin of the mammalian female reproductive system. We report the sequences of HoxA-11 from basal mammalian and amniote taxa and analyse HoxA-11 and HoxA-13 for signatures of adaptive molecular evolution. The data demonstrate that these genes were under strong positive (directional) selection in the stem lineage of therian and eutherian mammals, coincident with the evolution of the uterus and vagina. These results support the idea that adaptive evolution of transcription factors can be an integral part in the evolution of novel structures.

Adaptation, Biological↗

Possible evolution of splice-junction signals in eukaryotic genes from stop codons.

Splice-junction sequence signals are strongly conserved structural components of eukaryotic genes. These sequences border exon/intron junctions and aid in the process of removing introns by the RNA splicing machinery. Although substantial research has been undertaken to understand the mechanism of splicing, little is known about the origin and evolution of these splice signal sequences. Based on the previously published theory that the primitive genes evolved in pieces from primordial genetic sequences to avoid the interfering stop codons, a "stop-codon walk" mechanism is proposed in this paper to have assisted in the evolution of coding genes. This mechanism predicts the presence of stop codons in splice-junction signals inside the introns. Evidence of the consistent presence of stop codons in the splice-junction signals, in a position where they are expected, is shown by the analysis of codon statistics in these signal sequences in the GenBank databank. The results suggest that the splice-junction signals may have evolved from stop codons as a consequence of a selective pressure to avoid stop codons during the original evolution of coding genes. They also suggest that other splice signals within the introns, such as the branch-point sequence, may have evolved from stop codons for similar reasons.

Biological Evolution↗

Gamma carbonic anhydrases in plant mitochondria.

Three genes from Arabidopsis thaliana with high sequence similarity to gamma carbonic anhydrase (gammaCA), a Zn containing enzyme from Methanosarcina thermophila (CAM), were identified and characterized. Evolutionary and structural analyses predict that these genes code for active forms of gammaCA. Phylogenetic analyses reveal that these Arabidopsis gene products cluster together with CAM and related sequences from alpha and gamma proteobacteria, organisms proposed as the mitochondrial endosymbiont ancestor. Indeed, in vitro and in vivo experiments indicate that these gene products are transported into the mitochondria as occurs with several mitochondrial protein genes transferred, during evolution, from the endosymbiotic bacteria to the host genome. Moreover, putative CAM orthologous genes are detected in other plants and green algae and were predicted to be imported to mitochondria. Structural modeling and sequence analysis performed in more than a hundred homologous sequences show a high conservation of functionally important active site residues. Thus, the three histidine residues involved in Zn coordination (His 81, 117 and 122), Arg 59, Asp 61, Gin 75, and Asp 76 of CAM are conserved and properly arranged in the active site cavity of the models. Two other functionally important residues (Glu 62 and Glu 84 of CAM) are lacking, but alternative amino acids that might serve to their roles are postulated. Accordingly, we propose that photosynthetic eukaryotic organisms (green algae and plants) contain gammaCAs and that these enzymes codified by nuclear genes are imported into mitochondria to accomplish their biological function.

Amino Acid Sequence↗

Parametric dependence of SAR on permittivity values in a man model.

The development and widespread use of advanced three-dimensional digital anatomical models to calculate specific absorption rate (SAR) values in biological material has resulted in the need to understand how model parameters (e.g., permittivity value) affect the predicted whole-body and localized SAR values. The application of the man dosimetry model requires that permittivity values (dielectric value and conductivity) be allocated to the various tissues at all the frequencies to which the model will be exposed. In the 3-mm-resolution man model, the permittivity values for all 39 tissue-types were altered simultaneously for each orientation and applied frequency. In addition, permittivity values for muscle, fat, skin, and bone marrow were manipulated independently. The finite-difference time-domain code was used to predict localized and whole-body normalized SAR values. The model was processed in the far-field conditions at the resonant frequency (70 MHz) and above (200, 400, 918, and 2060 MHz) for E orientation. In addition, other orientations (K, H) of the model to the incident fields were used where no substantial resonant frequency exists. Variability in permittivity values did not substantially influence whole-body SAR values, while localized SAR values for individual tissues were substantially affected by these changes. Changes in permittivity had greatest effect on localized SAR values when they were low compare to the whole-body SAR value or when errors involved tissues that represent a substantial proportion of the body mass (i.e., muscle). Furthermore, we establish the partial derivative of whole-body and localized SAR values with respect to the dielectric value and conductivity for muscle independently. It was shown that uncertainties in dielectric value or conductivity do not substantially influence normalized whole-body SAR. Detailed investigation on localized SAR ratios showed that conductivity presents a more substantial factor in absorption of energy in tissues than dielectric value for almost all applied exposure conditions.

Absorption↗

The beta-tubulin gene of Epichloë typhina from perennial ryegrass (Lolium perenne).

Epichloë typhina is a biotrophic fungal pathogen which causes choke disease of pooid grasses. The anamorphic state, Acremonium typhinum, is placed in the section Albo-lanosa along with related, mutualistic, seed-disseminated endophytes. As an initial study of gene structure and evolution in Epichloë and related endophytes, the beta-tubulin gene, tub2, of the perennial ryegrass choke pathogen (EtPRG) was cloned and sequenced. The coding sequence and the predicted beta-tubulin amino acid sequence were highly homologous to the Neurospora crassa homologs, and to one of the two beta-tubulin genes of Emericella nidulans. However, two introns characteristic of the N. crassa and Em. nidulans genes were absent in the E. typhina gene. Furthermore, one of the remaining introns possessed the uncommon 5' splice junction, GC. In contrast to published observations concerning other Ascomycetes, a mutant of EtPRG, selected for resistance to methyl-2-benzimidazole carbamate (benomyl), possessed no alteration of its beta-tubulin coding sequence.

Amino Acid Sequence↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

How to find small non-coding RNAs in bacteria.

Small non-coding RNAs (sRNAs) have attracted considerable attention as an emerging class of gene expression regulators. In bacteria, a few regulatory RNA molecules have long been known, but the extent of their role in the cell was not fully appreciated until the recent discovery of hundreds of potential sRNA genes in the bacterium Escherichia coli. Orthologs of these E. coli sRNA genes, as well as unrelated sRNAs, were also found in other bacteria. Here we review the disparate experimental approaches used over the years to identify sRNA molecules and their genes in prokaryotes. These include genome-wide searches based on the biocomputational prediction of non-coding RNA genes, global detection of non-coding transcripts using microarrays, and shotgun cloning of small RNAs (RNomics). Other sRNAs were found by either co-purification with RNA-binding proteins, such as Hfq or CsrA/RsmA, or classical cloning of abundant small RNAs after size fractionation in polyacrylamide gels. In addition, bacterial genetics offers powerful tools that aid in the search for sRNAs that may play a critical role in the regulatory circuit of interest, for example, the response to stress or the adaptation to a change in nutrient availability. Many of the techniques discussed here have also been successfully applied to the discovery of eukaryotic and archaeal sRNAs.

Cloning, Molecular↗

Context in verbal short-term memory.

We tested the hypothesis that stimulus-related contextual information that is incidental to task demands-an episodic code-is automatically, obligatorily encoded and stored as a part of short-term memory (STM) representations. Four experiments employed a running span task to investigate the effects of manipulating two types of contextual information: stimulus grouping and color. Three experiments established that grouping context effects are sensitive neither to volitional control nor to task difficulty and that they generalize across testing procedures (yes/no recognition and immediate serial recall). A fourth experiment demonstrated an effect of manipulating the congruity of the color of stimuli between study and test. These demonstrations of the robustness and generality of context effects in STM are consistent with the predictions of the episodic coding model of STM.

Adolescent↗

Pseudoknots in RNA secondary structures: representation, enumeration, and prevalence.

A number of non-coding RNA are known to contain functionally important or conserved pseudoknots. However, pseudoknotted structures are more complex than orthodox, and most methods for analyzing secondary structures do not handle them. I present here a way to decompose and represent general secondary structures which extends the tree representation of the stem-loop structure, and use this to analyze the frequency of pseudoknots in known and in random secondary structures. This comparison shows that, though a number of pseudoknots exist, they are still relatively rare and mostly of the simpler kinds. In contrast, random secondary structures tend to be heavily knotted, and the number of available structures increases dramatically when allowing pseudoknots. Therefore, methods for structure prediction and non-coding RNA identification that allow pseudoknots are likely to be much less powerful than those that do not, unless they penalize pseudoknots appropriately.

Algorithms↗

Tomato LeAGP-1 arabinogalactan-protein purified from transgenic tobacco corroborates the Hyp contiguity hypothesis.

Functional analysis of the hyperglycosylated arabinogalactan-proteins (AGPs) attempts to relate biological roles to the molecular properties that result largely from O-Hyp glycosylation putatively coded by the primary sequence. The Hyp contiguity hypothesis predicts contiguous Hyp residues as attachment sites for arabino-oligosaccharides (arabinosides) and clustered, non-contiguous Hyp residues as arabinogalactan polysaccharide sites. Although earlier tests of naturally occurring hydroxyproline-rich glycoproteins (HRGPs) and HRGPs designed by synthetic genes were consistent with a sequence-driven code, the predictive value of the hypothesis starting from the DNA sequences of known AGPs remained untested due to difficulties in purifying a single AGP for analysis. However, expression in tobacco (Nicotiana tabacum) of the major tomato (Lycopersicon esculentum) AGP, LeAGP-1, as an enhanced green fluorescent protein fusion glycoprotein (EGFP)-LeAGP-1, increased its hydrophobicity sufficiently for chromatographic purification from other closely related endogenous AGPs. We also designed and purified two variants of LeAGP-1 for future functional analysis: one lacking the putative glycosylphosphatidylinositol (GPI)-anchor signal sequence; the other lacking a 12-residue internal lysine-rich region. Fluorescence microscopy of plasmolysed cells confirmed the location of LeAGP-1 at the plasma membrane outer surface and in Hechtian threads. Hyp glycoside profiles of the fusion glycoproteins gave ratios of Hyp-polysaccharides to Hyp-arabinosides plus non-glycosylated Hyp consistent with those predicted from DNA sequences by the Hyp contiguity hypothesis. These results demonstrate a route to the purification of AGPs and the use of the Hyp contiguity hypothesis for predicting the Hyp O-glycosylation profile of an HRGP from its DNA sequence.

Amino Acid Sequence↗

Positive predictive value of the diagnosis of acute myocardial infarction in an administrative database.

OBJECTIVE: To determine the positive predictive value of ICD-9-CM coding of acute myocardial infarction and cardiac procedures. METHODS: Using chart-abstracted data as the standard, we examined administrative data from the Veterans Health Administration for a national random sample of 5,151 discharges. MAIN RESULTS: The positive predictive value of acute myocardial infarction coding in the primary position was 96.9%. The sensitivity and specificity of coding were, respectively, 96% and 99% for catheterization, 95.7% and 100% for coronary artery bypass graft surgery, and 90.3% and 99. 7% for percutaneous transluminal coronary angioplasty. CONCLUSIONS: The positive predictive value of acute myocardial infarction and related procedure coding is comparable to or better than previously reported observations of administrative databases.

Aged↗

Coding of stimulus frequency by latency in thalamic networks through the interplay of GABAB-mediated feedback and stimulus shape.

A temporal sensory code occurs in posterior medial (POm) thalamus of the rat vibrissa system, where the latency for the spike rate to peak is observed to increase with increasing frequency of stimulation between 2 and 11 Hz. In contrast, the latency of the spike rate in the ventroposterior medial (VPm) thalamus is constant in this frequency range. We consider the hypothesis that two factors are essential for latency coding in the POm. The first is GABAB-mediated feedback inhibition from the reticular thalamic (Rt) nucleus, which provides delayed and prolonged input to thalamic structures. The second is sensory input that leads to an accelerating spike rate in brain stem nuclei. Essential aspects of the experimental observations are replicated by the analytical solution of a rate-based model with a minimal architecture that includes only the POm and Rt nuclei, i.e., an increase in stimulus frequency will increase the level of inhibitory output from Rt thalamus and lead to a longer latency in the activation of POm thalamus. This architecture, however, admits period-doubling at high levels of GABAB-mediated conductance. A full architecture that incorporates the VPm nucleus suppresses period-doubling. A clear match between the experimentally measured spike rates and the numerically calculated rates for the full model occurs when VPm thalamus receives stronger brain stem input and weaker GABAB-mediated inhibition than POm thalamus. Our analysis leads to the prediction that the latency code will disappear if GABAB-mediated transmission is blocked in POm thalamus or if the onset of sensory input is too abrupt. We suggest that GABAB-mediated inhibition is a substrate of temporal coding in normal brain function.

Action Potentials↗

Single-Cell Splicing Isoform Atlas of the Adult Human Heart and Heart Failure.

BACKGROUND: Alternative splicing plays crucial roles in normal heart development and cardiac disease by influencing protein-coding sequences, functional domains, and molecular networks. However, a detailed characterization of the human heart isoform landscape remains incomplete. METHODS: Leveraging long-read single-nucleus RNA sequencing and computational analysis, we dissected full-length isoform heterogeneities, expression patterns, and usage shifts across cell types, cell states, and cardiac conditions of the adult left ventricle. We applied in silico approaches to assess the functional relevance of identified isoforms; validated isoform compositions of representative cardiac genes using reverse transcription quantitative polymerase chain reaction and targeted amplicon sequencing; and developed a web server for interactive navigation of our results. RESULTS: The data revealed that isoform heterogeneity is widespread in the cardiac cellular system, serving as a posttranscriptional buffer mechanism that calibrates the molecule reservoirs in human hearts. In healthy left ventricles, ≈30% of cell type-specific genes were polyform, using multiple isoforms tailored to cell type-specific programs. Among ubiquitously expressed genes, >300 showed differential isoform usage with cell type specificity in normal hearts. Comparisons of cardiomyocytes across conditions uncovered 379 genes with marked isoform usage shifts, most of which are predicted to change protein coding outcomes through direct changes in protein coding sequences and switches between intron retention and non-protein-coding biotypes. In contrast, cell state-specific programs tend to operate on monoform genes associated with changes among cell states. In addition, our data revealed heart failure-associated differential isoform usage events in stromal and immune cell types in the cardiac microenvironment. CONCLUSIONS: We present a comprehensive atlas of splicing isoforms in the normal adult heart and heart failure through long-read single-nucleus RNA sequencing and computational analyses. The results suggest crucial roles of isoforms in buffering core cellular programs and contributing to disease-associated cell states. The full-length details of these cell-specific isoforms serve as an important reference for downstream translational and mechanistic studies and are available on our online data portal at https://github.com/gaolabtools/heart-isoform-atlas.

Humans↗

Automatic annotation of eukaryotic genes, pseudogenes and promoters.

BACKGROUND: The ENCODE gene prediction workshop (EGASP) has been organized to evaluate how well state-of-the-art automatic gene finding methods are able to reproduce the manual and experimental gene annotation of the human genome. We have used Softberry gene finding software to predict genes, pseudogenes and promoters in 44 selected ENCODE sequences representing approximately 1% (30 Mb) of the human genome. Predictions of gene finding programs were evaluated in terms of their ability to reproduce the ENCODE-HAVANA annotation. RESULTS: The Fgenesh++ gene prediction pipeline can identify 91% of coding nucleotides with a specificity of 90%. Our automatic pseudogene finder (PSF program) found 90% of the manually annotated pseudogenes and some new ones. The Fprom promoter prediction program identifies 80% of TATA promoters sequences with one false positive prediction per 2,000 base-pairs (bp) and 50% of TATA-less promoters with one false positive prediction per 650 bp. It can be used to identify transcription start sites upstream of annotated coding parts of genes found by gene prediction software. CONCLUSION: We review our software and underlying methods for identifying these three important structural and functional genome components and discuss the accuracy of predictions, recent advances and open problems in annotating genomic sequences. We have demonstrated that our methods can be effectively used for initial automatic annotation of the eukaryotic genome.

Animals↗

Sequence of the cDNA coding for the lethal neurotoxin Tx1 from the Brazilian "armed" spider Phoneutria nigriventer predicts the synthesis and processing of a preprotoxin.

A cDNA library was constructed from the venom glands of the Brazilian "armed" spider, Phoneutria nigriventer, and a clone coding for Tx1, a lethal toxin, was identified and sequenced. The sequence data derived from this cDNA clone combined with the previously determined amino acid sequence predict that Tx1 is initially synthesized as a preprotoxin. Four segments (comprising the signal sequence, a short, 15-amino acid, glutamate-rich sequence, the functional toxin, and 2 glycine residues) can be distinguished. The structure of the preprotoxin and the proposed processing steps required to form the mature Tx1 toxin show similarities with the synthesis and processing of omega-agatoxin IA.

Amino Acid Sequence↗

Accuracy of temporal coding: auditory-visual comparisons.

Three experiments were designed to decide whether temporal information is coded more accurately for intervals defined by auditory events or for those defined by visual events. In the first experiment, the irregular-list technique was used, in which a short list of items was presented, the items all separated by different interstimulus intervals. Following presentation, the subject was given three items from the list, in their correct serial order, and was asked to judge the relative interstimulus intervals. Performance was indistinguishable whether the items were presented auditorily or visually. In the second experiment, two unfilled intervals were defined by three nonverbal signals in either the auditory or the visual modality. After delays of 0, 9, or 18 sec (the latter two filled with distractor activity), the subjects were directed to make a verbal estimate of the length of one of the two intervals, which ranged from 1 to 4 sec and from 10 to 13 sec. Again, performance was not dependent on the modality of the time markers. The results of Experiment 3, which was procedurally similar to Experiment 2 but with filled rather than empty intervals, showed significant modality differences in one measure only. Within the range of intervals employed in the present study, our results provide, at best, only modest support for theories that predict more accurate temporal coding in memory for auditory, rather than visual, stimulus presentation.

Adult↗