Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Tracing ancient mRNA hairpins.

From recent developments of the early evolution theory it follows that the earliest mRNAs were short ( approximately 20 nt) (G+C)-rich polynucleotides. These short sequences could form hairpins, which would be of high evolutionary advantage because of stability and uniqueness of their conformations. Due to mutations accumulated during billions of years of evolution, the speculated earliest hairpins would largely lose the initial complementarities. Some of the original complementary base-to-base contacts, however, may have survived. Computational analysis of modern prokaryotic mRNA sequences reveals excess population of the expected short range complementarities. The derived earliest mRNA hairpin size fully corresponds to the predicted size of ancient coding duplexes. The repertoire of the surviving hairpins traced in modern mRNA confirms duplex structure of the earliest mRNA, suggested by the early molecular evolution theory.

Bacteria↗

Molecular cloning, expression, and chromosomal localization of the human earliest lymphocyte activation antigen AIM/CD69, a new member of the C-type animal lectin superfamily of signal-transmitting receptors.

The activation of T lymphocytes, both in vivo and in vitro, induces the expression of CD69. This molecule, which appears to be the earliest inducible cell surface glycoprotein acquired during lymphoid activation, is involved in lymphocyte proliferation and functions as a signal transmitting receptor in lymphocytes, natural killer (NK) cells, and platelets. To determine the structural basis for CD69 function, the cDNA coding for CD69 was isolated by a polymerase chain reaction-based strategy using oligonucleotides deduced from peptide sequences of the purified protein. The isolated cDNA exhibited a single open reading frame of 597 bp coding for CD69, and predicted a 199-amino acid protein of type II membrane topology, with extracellular (COOH-terminal), transmembrane, and intracellular domains. The CD69 clone hybridized to a 1.7-kb mRNA species, which was rapidly induced and degraded after lymphocyte stimulation, consistent with the presence of rapid degradation signals at the 3' untranslated region. Transient expression of the polypeptide encoded by CD69 cDNA in COS-7 cells demonstrated that it presented properties comparable to native CD69 protein. The CD69 gene was regionally mapped to chromosome 12 p13-p12 by both somatic cell hybrid DNA analysis and fluorescence in situ hybridization coupled with GTG banding (G bands by trypsin using Giemsa). Protein sequence homology search revealed that CD69 is a new member of the Ca(2+)-dependent (C-type) lectin superfamily of type II transmembrane receptors, which includes the human NKG2, the rat NKR-P1, and the mouse NKR-P1 families of NK cell-specific genes. CD69 also has a structural homology with other type II lectin cell surface receptors, such as the T cell antigen Ly49, the low avidity immunoglobulin E receptor (CD23), and the hepatic asialoglycoprotein receptors. The CD69 protein also shares functional characteristics with most members of this superfamily, which act as transmembrane signaling receptors in early phases of cellular activation.

Amino Acid Sequence↗

Human T-cell lymphotropic virus type III (HTLV-III) core antigens: synthesis in Escherichia coli and immunoreactivity with human sera.

Fragments of human T-cell lymphotropic virus type III (HTLV-III) proviral DNA carrying the gene for the core antigen (gag) was cloned in the plasmid REV. Several of the recombinants direct high levels of synthesis of the antigens. One clone, pG1, produced a hybrid protein containing 13 amino acid residues of the carboxyl terminus of the 17 kD virion protein, the entire p24, the major core protein of HTLV-III, and 74 amino acid residues of the amino terminal of the 15 kD core ribonucleoprotein. A second clone, pG2, was similar to pG1 except that it contained no p17 sequences and was missing the amino-terminal 77 amino acid residues of the p24. A third clone, pG3, was similar to pG2, except that all but 56 amino acids of the carboxyl terminus of p24 were removed. All three proteins were found to be strongly immunoreactive with anti-HTLV-III antibodies present in sera from patients with acquired immune deficiency syndrome (AIDS) or AIDS-related complex (ARC). In addition, pG1 and pG2, but not pG3, reacted with a monoclonal antibody (M26) specific for the p24 virion core protein. Whereas all three reacted with an anti-p15 monoclonal antibody, none of the clones reacted with an anti-p17 monoclonal antibody. These results provide direct evidence to support the predicted assignment of the coding region of the gag gene of HTLV-III. The product from pG2 was purified and was found to be potentially useful for the detection of anti-p24 antibodies in sera from patients with AIDS or ARC and from individuals at risk from AIDS.

Acquired Immunodeficiency Syndrome↗

Interventions to enhance communication among patients, providers, and families.

Whether patient suffering is caused by physical symptoms, unwanted medical intervention, or spiritual crisis, the common pathway to relief is through a provider who is able to elicit these concerns and is equipped to help the patient and family address them. This paper reviews the current state of knowledge in communication at the end of life, organized according to a framework of information gathering, information giving, and relationship building; and then focuses on interventions to enhance communication among patients, providers, and families. Several observations emerge from the existing literature. Patients have highly individualized desires for information and we cannot predict patient preferences. Communication coding methodology has advanced significantly yet the current systems remain poorly understood and largely inaccessible. Physicians and other health care providers do not discuss sufficiently treatment options, quality of life or respond to emotional cues from patients, and there is plenty of room for improvement. On the positive side, we have also learned that physicians and other health care providers can be taught to communicate better through intensive communication courses, and that communication interventions can improve some patient outcomes. Finally, huge gaps remain in our current knowledge, particularly with regard to understanding the relationship between communication style and outcomes. These findings suggest several recommendations. We should create larger and more diverse datasets; improve upon the analysis of recorded communication data; increase our knowledge about patient preferences for information; establish a stronger link between specific communication behaviors and outcomes; and identify more efficient ways to teach providers communication skills.

Family Relations↗

Predicting protein structure classes from function predictions.

MOTIVATION: We introduce a new approach to using the information contained in sequence-to-function prediction data in order to recognize protein template classes, a critical step in predicting protein structure. The data on which our method is based comprise probabilities of functional categories; for given query sequences these probabilities are obtained by a neural net that has previously been trained on a variety of functionally important features. On a training set of sequences we assess the relevance of individual functional categories for identifying a given structural family. Using a combination of the most relevant categories, the likelihood of a query sequence to belong to a specific family can be estimated. RESULTS: The performance of the method is evaluated using cross-validation. For a fixed structural family and for every sequence, a score is calculated that measures the evidence for family membership. Even for structural families of small size, family members receive significantly higher scores. For some examples, we show that the relevant functional features identified by this method are biologically meaningful. The proposed approach can be used to improve existing sequence-to-structure prediction methods. AVAILABILITY: Matlab code is available on request from the authors. The data are available at http://www.mpisb.mpg.de/~sommer/Fun2Struc/

Algorithms↗

TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders.

UNLABELLED: We describe two new Generalized Hidden Markov Model implementations for ab initio eukaryotic gene prediction. The C/C++ source code for both is available as open source and is highly reusable due to their modular and extensible architectures. Unlike most of the currently available gene-finders, the programs are re-trainable by the end user. They are also re-configurable and include several types of probabilistic submodels which can be independently combined, such as Maximal Dependence Decomposition trees and interpolated Markov models. Both programs have been used at TIGR for the annotation of the Aspergillus fumigatus and Toxoplasma gondii genomes. AVAILABILITY: Source code and documentation are available under the open source Artistic License from http://www.tigr.org/software/pirate

Algorithms↗

Intramolecular surface contacts contain information about protein-protein interface regions.

MOTIVATION: Some amino acids clearly show preferences over others in protein-protein interfaces. These preferences, or so-called interface propensities can be used for a priori interface prediction. We investigated whether the prediction accuracy could be improved by considering not single but pairs of residues in an interface. Here we present the first systematic analysis of intramolecular surface contacts in interface prediction. RESULTS: We show that preferences do exist for contacts within and around an interface region within one molecule: specific pairs of amino acids are more often occurring than others. Using intramolecular contact propensities in a blind test, higher average scores were assigned to interface residues than to non-interface residues. This effect persisted as small but significant when the contact propensities were corrected to eliminate the influence of single amino acid interface propensity. This indicates that intramolecular contact propensities may replace interface propensities in protein-protein interface prediction. AVAILABILITY: The source code is available on request from the authors.

Algorithms↗

Genome-wide identification and molecular characterization of Ole_e_I, Allerg_1 and Allerg_2 domain-containing pollen-allergen-like genes in Oryza sativa.

Pollen allergens play important roles in plant development in addition to their allergenic nature for human. More than 10 groups of pollen allergens have been reported. Among them, Pollen_Ole_e_I (Ole), Pollen_allerg_1 (Allerg1) and Pollen_allerg_2 (Allerg2) domain-containing proteins are the majority of allergens. We have identified 114 pollen-allergen-like genes in rice genome by bioinformatics using public databases. Among them, 45 genes encode Ole domain-containing proteins, 62 with Allerg1 and 7 with Allerg2. They are distributed on 11 of 12 rice chromosomes excluding chromosome 11. Comparison analysis of coding regions from both predicted genes and isolated full-length cDNAs showed that most of predicted genes were correct in the splicing of exons and introns, and only 7 exhibited wrong predictions. The fact suggested the applicability of the prediction programs to identify pollen-allergen genes. Phylogenetic analysis revealed the high diversity within OsOle genes and recent evolutionary event in OsAllerg1 genes, and suggested that some of OsOle genes were new members of the family. Expression analysis by RT-PCR showed that most of the genes were expressed in all tested tissues and only eight genes exhibited panicle-specific expression, suggesting that pollen-allergen genes play roles in not only productive but also vegetative development.

Allergens↗

Cha4p of Saccharomyces cerevisiae activates transcription via serine/threonine response elements.

The CHA1 gene of Saccharomyces cerevisiae encodes the catabolic L-serine (L-threonine) deaminase responsible for the utilization of serine/threonine as nitrogen sources. Previously, we identified two serine/threonine response elements in the CHA1 promoter, UASCHA. We report isolation of a mutation, cha4-1, that impairs serine/threonine induction of CHA1 transcription. The cha4-1 allele causes noninducibility of a CHA1 p-lacZ translational gene fusion, indicating that Cha4p exerts its action through the CHA1 promoter. Molecular and genetic mapping positioned the cha4 locus 17 cM centromere proximal to put1 on chromosome XII. The coding region of CHA4 predicts a 648-amino acid protein with a DNA-binding motif (residues 43-70) belonging to the Cys6 zinc cluster class. Gel retardation employing a recombinant peptide, Cha4p1-174, demonstrated that the peptide in vitro specifically binds UASCHA. Binding is abolished by a G-C to T-A mutation in the middle bases of the two CEZ-elements in UASCHA. The transcriptional activating ability of UASCHA derivatives in vivo correlates with their ability to bind Cha4p1-174 in vitro. We conclude that Cha4p is a positive regulator of CHA1 transcription and that Cha4p alone, or as part of a complex, is binding UASCHA.

Alleles↗

A hypervariable segment in the human dopamine receptor D4 (DRD4) gene.

The human dopamine D4 receptor contains a novel polymorphism within the putative third cytoplasmic loop of the protein. The polymorphism is characterized by a varying number of direct imperfect 48-bp repeats in the gene. Pharmacological characterization has suggested that this receptor is the site through which the atypical neuroleptic clozapine exerts its antipsychotic action and that some polymorphic variants display different pharmacological properties. Further analysis of the repeat region using innovative technologies indicates that the alleles vary not only in the number of repeats (2-8 or 10 repeat units) but also in the sequence of the repeats and the order in which they appear. In 178 unrelated chromosomes we have identified 19 different repeats in 25 different haplotypes coding for 18 different predicted amino acid sequences, making this one of the most variable functional proteins currently described.

Alleles↗

SLAM web server for comparative gene finding and alignment.

SLAM is a program that simultaneously aligns and annotates pairs of homologous sequences. The SLAM web server integrates SLAM with repeat masking tools and the AVID alignment program to allow for rapid alignment and gene prediction in user submitted sequences. Along with annotations and alignments for the submitted sequences, users obtain a list of predicted conserved non-coding sequences (and their associated alignments). The web site also links to whole genome annotations of the human, mouse and rat genomes produced with the SLAM program. The server can be accessed at http://bio.math.berkeley.edu/slam.

Algorithms↗

Fungal elicitation of signal transduction-related plant genes precedes mycorrhiza establishment and requires the dmi3 gene in Medicago truncatula.

Suppressive subtractive hybridization and expressed sequence tag sequencing identified 29 plant genes which are upregulated during the appressorium stage of mycorrhiza establishment between Medicago truncatula J5 (Myc+) and Glomus mosseae. Eleven genes coding plant proteins with predicted functions in signal transduction, transcription, and translation were investigated in more detail for their relation to early events of symbiotic interactions. Expression profiling showed that the genes are activated not only from the appressorium stage up to the fully established symbiosis in the Myc+ genotype of M. truncatula, but also when the symbionts are not in direct cell contact, suggesting that diffusible fungal molecules (Myc factors) play a, role in the induction of a signal-transduction pathway. Transcript accumulation in roots of a mycorrhiza-defective Myc- dmi3 mutant of M. truncatula is not modified by appressorium formation or diffusible fungal molecules, indicating that the signal transduction pathway is required for a successful G. mosseae-M. truncatula interaction leading to symbiosis development. The symbiotic nodulating bacterium Sinorhizobium meliloti does not activate the 11 genes, which supposes early discrimination by plant roots between the microbial symbionts.

Gene Expression Regulation, Plant↗

Characterization of CTG/CAG repeats on chromosome 18: a study of bipolar disorder.

Anticipation has been frequently found in bipolar families ascertained for linkage studies. An association of polymorphic triplet repeats with the bipolar phenotype in some pedigrees has been proposed. We have previously found linkage to chromosome 18 in a set of families with evidence of anticipation. As part of a search for CAG/CTG motifs on chromosome 18, we screened a genomic chromosome 18 cosmid library and identified 65 loci with trinucleotide repeats. Eleven of 33 genotyped loci were polymorphic, though none of these showed any evidence of instability. We performed genetic analysis of six loci in the Hopkins/Dana bipolar pedigrees ascertained for a genetic linkage study of bipolar disorder and found that the CAG repeat within the AD4D2 clone on 18q21.1 showed nominally significant over-transmission of the rare CAG23 allele (P=0.034). We have characterized all 65 trinucleotide repeats and flanking sequences with GENSCAN analysis and find that 29 were predicted to be in coding regions. These 29 trinucleotide-repeat-containing genes may be involved in functional modulation of their respective proteins, and may be candidates for other diseases or disease mechanisms that map to this region.

Animals↗

Linkage on chromosome 10 of several murine retroviral integration loci associated with leukaemia.

Mml loci have been identified as provirus integration sites among a subset of monocytic tumours induced by murine leukaemia virus (MuLV) infection of BALB/c and DBA/2 mice. These myeloid leukaemias contain a retrovirus integrated on chromosome 10 in proximity to the c-myb locus; however, c-myb expression was not altered. Detailed physical mapping enabled placement of the retroviral integration sites approximately 25 kb (Mml1), approximately 51 kb (Mml2), and approximately 70 kb (Mml3) upstream of the c-myb locus. Furthermore, the Fti1 (fit-1) locus, a common integration site in feline leukaemia virus-induced T cell lymphomas, was mapped upstream of Mml3. Sequence analysis of Mml1, Mml2 and Mml3 loci (39.6, 16.4 and 5.9 kb, respectively) in conjunction with the BLAST (basic local alignment search tool) homology searches against the expressed sequence tag (EST) database and the use of gene/exon prediction programs revealed potential coding sequences that were not confirmed by Northern analysis or RT-PCR. The sequences between c-myb and Fti1, which were shown to include two potential scaffold/matrix attachment regions (S/MARs), are most likely regulatory in nature. An extended search for transcribed sequences far upstream of Mml3 revealed five genes, four of which were expressed in multiple tissues in mice. These genes could not be linked to tumour formation by the virus but their homologous sequences were found on human chromosome 6, thus allowing extension of the syntenic region on mouse chromosome 10 to approximately 250 kb.

Amino Acid Sequence↗

Molecular typing of enteroviruses associated with viral meningitis in Cyprus, 2000-2002.

Human enteroviruses are responsible for a wide spectrum of clinical diseases affecting many different organ systems. Although infection is usually asymptomatic, infections of the central nervous system manifested as meningitis or encephalitis can pose a serious public health problem, especially during outbreaks. In this study, samples from 218 patients diagnosed with enteroviral meningitis between January 2000 and December 2002 were analysed in order to assess the epidemiology of human enteroviruses as a cause of viral meningitis in Cyprus. A new typing strategy, based on partial sequencing of the 5' non-coding region (5'NCR), prediction of type, and selection of type-specific primers for sensitive VP1 PCR amplification, was developed. As clustering in the 5'NCR was concordant with clustering in the VP1 region, quick and reliable typing by VP1 sequencing was achieved without virus isolation in cell culture. The most frequent enterovirus serotypes identified were Human echovirus 30 (55.5%), Human echovirus 13 (15.1%), Human echovirus 6 (13.8%) and Human echovirus 9 (8.3%). Human coxsackieviruses B2, B1 and B5, Human echovirus 4, Human enterovirus 71 and Human coxsackievirus A6 represented rather rare serotypes. This is the first molecular epidemiological study of enterovirus meningitis in Cyprus. Serotype distribution corresponded basically with observations in other European countries, suggesting the spread of enteroviruses by tourism.

5' Untranslated Regions↗

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗

Population coding in neuronal systems with correlated noise.

Neuronal representations of external events are often distributed across large populations of cells. We study the effect of correlated noise on the accuracy of these neuronal population codes. Our main question is whether the inherent error in the population code can be suppressed by increasing the size of the population N in the presence of correlated noise. We address this issue using a model of a population of neurons that are broadly tuned to an angular variable in two dimensions. The fluctuations in the neuronal activities are modeled as Gaussian noises with pairwise correlations that decay exponentially with the difference between the preferred angles of the correlated cells. We assume that the system is broadly tuned, which means that both the correlation length and the width of the tuning curves of the mean responses span a substantial fraction of the entire system length. The performance of the system is measured by the Fisher information (FI), which bounds its estimation error. By calculating the FI in the limit of a large N, we show that positive correlations decrease the estimation capability of the network, relative to the uncorrelated population. The information capacity saturates to a finite value as the number of cells in the population grows. In contrast, negative correlations substantially increase the information capacity of the neuronal population. These results are supplemented by the effect of correlations on the mutual information of the system. Our analysis provides an estimate of the effective number of statistically independent degrees of freedom, denoted N(eff), that a large correlated system can have. According to our theory N(eff) remains finite in the limit of a large N. Estimating the parameters of the correlations and tuning curves from experimental data in some cortical areas that code for angles, we predict that the number of effective degrees of freedom embedded in localized populations in these areas is less than or of the order of approximately 10(2).

Animals↗

Initial state randomness improves sequence learning in a model hippocampal network.

Randomness can be a useful component of computation. Using a computationally minimal, but still biologically based model of the hippocampus, we evaluate the effects of initial state randomization on learning a cognitive problem that requires this brain structure. Greater randomness of initial states leads to more robust performance in simulations of the cognitive task called transverse patterning, a context-dependent discrimination task that we code as a sequence prediction problem. At the conclusion of training, greater initial randomness during training trials also correlates with increased, repetitive firing of select individual neurons, previously named local context neurons. In essence, such repetitively firing neurons recognize subsequences, and previously their presence has been correlated with solving the transverse patterning problem. A more detailed analysis of the simulations across training trials reveals more about initial state randomization. The beneficial effects of initial state randomization derive from enhanced variation, across training trials, of the sequential states of a network. This greater variation is not uniformly present during training; it is largely restricted to the beginning of training and when novel sequences are introduced. Little such variation occurs after extensive or even moderate amounts of training. We explain why variation is high early in training, but not later. This automatic modulation of the initial-state-driven random variation through state space is reminiscent of simulated annealing where modulated randomization encourages a selectively broad search through state space. In contrast to an annealing schedule, the selective occurrence of such a random search here is an emergent property, and the critical randomization occurs during training rather than testing.

Animals↗