Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Defining an epigenetic code.

The nucleosome surface is decorated with an array of enzyme-catalysed modifications on histone tails. These modifications have well-defined roles in a variety of ongoing chromatin functions, often by acting as receptors for non-histone proteins, but their longer-term effects are less clear. Here, an attempt is made to define how histone modifications operate as part of a predictive and heritable epigenetic code that specifies patterns of gene expression through differentiation and development.

Animals↗

A cDNA clone for a novel nuclear protein with DNA binding activity.

In an effort to identify trans-acting factors regulating specific genes, we cloned a novel human gene, DBP-5. The cDNA clone contains a predicted open reading frame coding for a potential 1,179 amino acid protein. The mRNA corresponding to DBP-5 is ubiquitously distributed, and the gene is phylogenetically conserved. Immunofluorescence analyses with several cell lines indicate that the protein is localized to the nucleus. Sequence analysis revealed unusual features of the predicted protein structure, including four completely conserved repeats. The phylogenetic conservation of DBP-5, the ubiquity of its expression, its nuclear localization, and its ability to bind DNA sequences, raise the possibility that DBP-5 may play a role in the organization of interphase chromatin and/or in transcriptional regulation.

Amino Acid Sequence↗

Prediction of the site and severity of obstruction in hypertrophic cardiomyopathy by color flow mapping and continuous wave Doppler echocardiography.

OBJECTIVE: We investigated whether the site and severity of an obstruction in hypertrophic cardiomyopathy can be accurately predicted by the combined use of color-coded and continuous wave Doppler echocardiography. BACKGROUND: Predicting the site of obstruction by end-systolic cavity shape is not reliable. Therefore, hemodynamic localization of the obstruction is required before surgery is performed. Such localization should be possible with color flow imaging, which provides two-dimensional velocity mapping reflecting the distribution of pressures within the left ventricle. Discrepancies in assessment of the pressure gradient by Doppler echocardiography and cardiac catheterization (which are usually not performed simultaneously) may be due to spontaneous variation of the dynamic obstruction in addition to technical factors related to both methods. METHODS: Twenty consecutive patients with hypertrophic cardiomyopathy were examined 1 day before transseptal left heart catheterization. The obstruction site was defined by color flow mapping. The pressure gradient was determined by continuous wave Doppler echocardiography. Measurements were also performed simultaneously in 10 patients during cardiac catheterization. RESULTS: Midventricular obstruction was correctly identified in 4 patients and subvalvular obstruction in 15 patients. One patient had no obstruction at rest. Invasively and noninvasively determined pressure gradients correlated well (r = 0.89, SEE = 16.3 mm Hg). Multiple single-beat analysis in 10 patients, also simultaneously examined with Doppler echocardiography and catheterization, yielded an excellent correlation (r = 0.97, SEE = 13.1 mm Hg). Comparing the simultaneous (r = 0.96, SEE = 12.5 mm Hg) and nonsimultaneous (r = 0.81, SEE = 23.8 mm Hg) recordings in these patients, we found that the spontaneous variation of the dynamic obstruction mainly accounted for discrepancies (p less than 0.05). CONCLUSION: The combined use of color-coded and continuous wave Doppler echocardiography provides the relevant hemodynamic information required for decision-making in patients with hypertrophic cardiomyopathy who are considered for transaortic myectomy.

Adult↗

Accuracy of ICD-9 coding for Clostridium difficile infections: a retrospective cohort.

Clostridium difficile (C. diff) is a major nosocomial problem. Epidemiological surveillance of the disease can be accomplished by microbiological or administrative data. Microbiological tracking is problematic since it does not always translate into clinical disease, and it is not always available. Tracking by administrative data is attractive, but ICD-9 code accuracy for C. diff is unknown. By using a large administrative database of hospitalized patients with C. diff (by ICD-9 code or cytotoxic assay), this study found that the sensitivity, specificity, positive, and negative predictive values of ICD-9 coding were 71%, 99%, 87%, and 96% respectively (using micro data as the gold standard). When only using symptomatic patients the sensitivity increased to 82% and when only using symptomatic patients whose test results were available at discharge, the sensitivity increased to 88%. C. diff ICD-9 codes closely approximate true C. diff infection, especially in symptomatic patients whose test results are available at the time of discharge, and can therefore be used as a reasonable alternative to microbiological data for tracking purposes.

Boston↗

Genome-based analysis of virulence genes in a non-biofilm-forming Staphylococcus epidermidis strain (ATCC 12228).

Staphylococcus epidermidis strains are diverse in their pathogenicity; some are invasive and cause serious nosocomial infections, whereas others are non-pathogenic commensal organisms. To analyse the implications of different virulence factors in Staphylococcus epidermidis infections, the complete genome of Staphylococcus epidermidis strain ATCC 12228, a non-biofilm forming, non-infection associated strain used for detection of residual antibiotics in food products, was sequenced. This strain showed low virulence by mouse and rat experimental infections. The genome consists of a single 2499 279 bp chromosome and six plasmids. The chromosomal G + C content is 32.1% and 2419 protein coding sequences (CDS) are predicted, among which 230 are putative novel genes. Compared to the virulence factors in Staphylococcus aureus, aside from delta-haemolysin and beta-haemolysin, other toxin genes were not found. In contrast, the majority of adhesin genes are intact in ATCC 12228. Most strikingly, the ica operon coding for the enzymes synthesizing interbacterial cellular polysaccharide is missing in ATCC 12228 and rearrangements of adjacent genes are shown. No mec genes, IS256, IS257, were found in ATCC 12228. It is suggested that the absence of the ica operon is a genetic marker in commensal Staphylococcus epidermidis strains which are less likely to become invasive.

Adhesins, Bacterial↗

An exploration of the sequence of a 2.9-Mb region of the genome of Drosophila melanogaster: the Adh region.

A contiguous sequence of nearly 3 Mb from the genome of Drosophila melanogaster has been sequenced from a series of overlapping P1 and BAC clones. This region covers 69 chromosome polytene bands on chromosome arm 2L, including the genetically well-characterized "Adh region." A computational analysis of the sequence predicts 218 protein-coding genes, 11 tRNAs, and 17 transposable element sequences. At least 38 of the protein-coding genes are arranged in clusters of from 2 to 6 closely related genes, suggesting extensive tandem duplication. The gene density is one protein-coding gene every 13 kb; the transposable element density is one element every 171 kb. Of 73 genes in this region identified by genetic analysis, 49 have been located on the sequence; P-element insertions have been mapped to 43 genes. Ninety-five (44%) of the known and predicted genes match a Drosophila EST, and 144 (66%) have clear similarities to proteins in other organisms. Genes known to have mutant phenotypes are more likely to be represented in cDNA libraries, and far more likely to have products similar to proteins of other organisms, than are genes with no known mutant phenotype. Over 650 chromosome aberration breakpoints map to this chromosome region, and their nonrandom distribution on the genetic map reflects variation in gene spacing on the DNA. This is the first large-scale analysis of the genome of D. melanogaster at the sequence level. In addition to the direct results obtained, this analysis has allowed us to develop and test methods that will be needed to interpret the complete sequence of the genome of this species. Before beginning a Hunt, it is wise to ask someone what you are looking for before you begin looking for it. Milne 1926

Alcohol Dehydrogenase↗

Theatre: A software tool for detailed comparative analysis and visualization of genomic sequence.

Theatre is a web-based computing system designed for the comparative analysis of genomic sequences, especially with respect to motifs likely to be involved in the regulation of gene expression. Theatre is an interface to commonly used sequence analysis tools and biological sequence databases to determine or predict the positions of coding regions, repetitive sequences and transcription factor binding sites in families of DNA sequences. The information is displayed in a manner that can be easily understood and can reveal patterns that might not otherwise have been noticed. In addition to web-based output, Theatre can produce publication quality colour hardcopies showing predicted features in aligned genomic sequences. A case study using the p53 promoter region of four mammalian species and two fish species is described. Unlike the mammalian sequences the promoter regions in fish have not been previously predicted or characterized and we report the differences in the p53 promoter region of four mammals and that predicted for two fish species. Theatre can be accessed at http://www.hgmp.mrc.ac.uk/Registered/Webapp/theatre/.

Animals↗

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO↗

Role of experience and oscillations in transforming a rate code into a temporal code.

In the vast majority of brain areas, the firing rates of neurons, averaged over several hundred milliseconds to several seconds, can be strongly modulated by, and provide accurate information about, properties of their inputs. This is referred to as the rate code. However, the biophysical laws of synaptic plasticity require precise timing of spikes over short timescales (<10 ms). Hence it is critical to understand the physiological mechanisms that can generate precise spike timing in vivo, and the relationship between such a temporal code and a rate code. Here we propose a mechanism by which a temporal code can be generated through an interaction between an asymmetric rate code and oscillatory inhibition. Consistent with the predictions of our model, the rate and temporal codes of hippocampal pyramidal neurons are highly correlated. Furthermore, the temporal code becomes more robust with experience. The resulting spike timing satisfies the temporal order constraints of hebbian learning. Thus, oscillations and receptive field asymmetry may have a critical role in temporal sequence learning.

Action Potentials↗

Predicted versus measured tritium oxide concentrations at the Savannah River Site.

Measured tritium oxide concentrations in air at various offsite locations are compared with concentrations predicted by three computer codes that are utilized at the Savannah River Site to estimate doses to maximally exposed offsite individuals. Annual average concentrations calculated by the computer models were compared with measured average concentrations taken from monitoring data collected over the last 10 y. The computer programs used for the comparison are AXAIRQ, MAXIGASP, and CAP88. The 10-y averaged ratios of predicted-to-measured tritium oxide air concentrations using AXAIRQ, MAXIGASP, and CAP88 are 1.89+/-0.56, 1.70+/-0.48, and 1.40+/-0.39, respectively. The difference in ratios is primarily due to different wind speed averages used within each of the models. These results show exceptional agreement, considering Gaussian plume models typically over predict annual average air concentrations by a factor of two to four.

Air Pollutants, Radioactive↗

Signal sequence analysis of expressed sequence tags from the nematode Nippostrongylus brasiliensis and the evolution of secreted proteins in parasites.

BACKGROUND: Parasitism is a highly successful mode of life and one that requires suites of gene adaptations to permit survival within a potentially hostile host. Among such adaptations is the secretion of proteins capable of modifying or manipulating the host environment. Nippostrongylus brasiliensis is a well-studied model nematode parasite of rodents, which secretes products known to modulate host immunity. RESULTS: Taking a genomic approach to characterize potential secreted products, we analyzed expressed sequence tag (EST) sequences for putative amino-terminal secretory signals. We sequenced ESTs from a cDNA library constructed by oligo-capping to select full-length cDNAs, as well as from conventional cDNA libraries. SignalP analysis was applied to predicted open reading frames, to identify potential signal peptides and anchors. Among 1,234 ESTs, 197 (~16%) contain predicted 5' signal sequences, with 176 classified as conventional signal peptides and 21 as signal anchors. ESTs cluster into 742 distinct genes, of which 135 (18%) bear predicted signal-sequence coding regions. Comparisons of clusters with homologs from Caenorhabditis elegans and more distantly related organisms reveal that the majority (65% at P < e-10) of signal peptide-bearing sequences from N. brasiliensis show no similarity to previously reported genes, and less than 10% align to conserved genes recorded outside the phylum Nematoda. Of all novel sequences identified, 32% contained predicted signal peptides, whereas this was the case for only 3.4% of conserved genes with sequence homologies beyond the Nematoda. CONCLUSIONS: These results indicate that secreted proteins may be undergoing accelerated evolution, either because of relaxed functional constraints, or in response to stronger selective pressure from host immunity.

Animals↗

Identification of open reading frames unique to a select agent: Ralstonia solanacearum race 3 biovar 2.

An 8x draft genome was obtained and annotated for Ralstonia solanacearum race 3 biovar 2 (R3B2) strain UW551, a United States Department of Agriculture Select Agent isolated from geranium. The draft UW551 genome consisted of 80,169 reads resulting in 582 contigs containing 5,925,491 base pairs, with an average 64.5% GC content. Annotation revealed a predicted 4,454 protein coding open reading frames (ORFs), 43 tRNAs, and 5 rRNAs; 2,793 (or 62%) of the ORFs had a functional assignment. The UW551 genome was compared with the published genome of R. solanacearum race 1 biovar 3 tropical tomato strain GMI1000. The two phylogenetically distinct strains were at least 71% syntenic in gene organization. Most genes encoding known pathogenicity determinants, including predicted type III secreted effectors, appeared to be common to both strains. A total of 402 unique UW551 ORFs were identified, none of which had a best hit or >45% amino acid sequence identity with any R. solanacearum predicted protein; 16 had strong (E < 10(-13)) best hits to ORFs found in other bacterial plant pathogens. Many of the 402 unique genes were clustered, including 5 found in the hrp region and 38 contiguous, potential prophage genes. Conservation of some UW551 unique genes among R3B2 strains was examined by polymerase chain reaction among a group of 58 strains from different races and biovars, resulting in the identification of genes that may be potentially useful for diagnostic detection and identification of R3B2 strains. One 22-kb region that appears to be present in GMI1000 as a result of horizontal gene transfer is absent from UW551 and encodes enzymes that likely are essential for utilization of the three sugar alcohols that distinguish biovars 3 and 4 from biovars 1 and 2.

Arginine↗

Discovery of estrogen receptor alpha target genes and response elements in breast tumor cells.

BACKGROUND: Estrogens and their receptors are important in human development, physiology and disease. In this study, we utilized an integrated genome-wide molecular and computational approach to characterize the interaction between the activated estrogen receptor (ER) and the regulatory elements of candidate target genes. RESULTS: Of around 19,000 genes surveyed in this study, we observed 137 ER-regulated genes in T-47D cells, of which only 89 were direct target genes. Meta-analysis of heterogeneous in vitro and in vivo datasets showed that the expression profiles in T-47D and MCF-7 cells are remarkably similar and overlap with genes differentially expressed between ER-positive and ER-negative tumors. Computational analysis revealed a significant enrichment of putative estrogen response elements (EREs) in the cis-regulatory regions of direct target genes. Chromatin immunoprecipitation confirmed ligand-dependent ER binding at the computationally predicted EREs in our highest ranked ER direct target genes, NRIP1, GREB1 and ABCA3. Wider examination of the cis-regulatory regions flanking the transcriptional start sites showed species conservation in mouse-human comparisons in only 6% of predicted EREs. CONCLUSIONS: Only a small core set of human genes, validated across experimental systems and closely associated with ER status in breast tumors, appear to be sufficient to induce ER effects in breast cancer cells. That cis-regulatory regions of these core ER target genes are poorly conserved suggests that different evolutionary mechanisms are operative at transcriptional control elements than at coding regions. These results predict that certain biological effects of estrogen signaling will differ between mouse and human to a larger extent than previously thought.

Binding Sites↗

Dose uncertainties for large solar particle events: input spectra variability and human geometry approximations.

The true uncertainties in estimates of body organ absorbed dose and dose equivalent, from exposures of interplanetary astronauts to large solar particle events (SPEs), are essentially unknown. Variations in models used to parameterize SPE proton spectra for input into space radiation transport and shielding computer codes can result in uncertainty about the reliability of dose predictions for these events. Also, different radiation transport codes and their input databases can yield significant differences in dose predictions, even for the same input spectra. Different results may also be obtained for the same input spectra and transport codes if different spacecraft and body self-shielding distributions are assumed. Heretofore there have been no systematic investigations of the variations in dose and dose equivalent resulting from these assumptions and models. In this work we present a study of the variability in predictions of organ dose and dose equivalent arising from the use of different parameters to represent the same incident SPE proton data and from the use of equivalent sphere approximations to represent human body geometry. The study uses the BRYNTRN space radiation transport code to calculate dose and dose equivalent for the skin, ocular lens and bone marrow using the October 1989 SPE as a model event. Comparisons of organ dose and dose equivalent, obtained with a realistic human geometry model and with the oft-used equivalent sphere approximation, are also made. It is demonstrated that variations of 30-40% in organ dose and dose equivalent are obtained for slight variations in spectral fitting parameters obtained when various data points are included or excluded from the fitting procedure. It is further demonstrated that extrapolating spectra from low energy (< or = 30 MeV) proton fluence measurements, rather than using fluence data extending out to 100 MeV results in dose and dose equivalent predictions that are underestimated by factors as large as 2-3. Finally, it is also demonstrated that the use of equivalent sphere approximations to represent body organ self-shielding distributions results in organ doses and dose equivalent predictions that are 2-3 times larger than values obtained with anthropomorphic shielding configurations.

Aluminum↗

Linear programming optimization and a double statistical filter for protein threading protocols.

The design of scoring functions (or potentials) for threading, differentiating native-like from non-native structures with a limited computational cost, is an active field of research. We revisit two widely used families of threading potentials: the pairwise and profile models. To design optimal scoring functions we use linear programming (LP). The LP protocol makes it possible to measure the difficulty of a particular training set in conjunction with a specific form of the scoring function. Gapless threading demonstrates that pair potentials have larger prediction capacity compared with profile energies. However, alignments with gaps are easier to compute with profile potentials. We therefore search and propose a new profile model with comparable prediction capacity to contact potentials. A protocol to determine optimal energy parameters for gaps, using LP, is also presented. A statistical test, based on a combination of local and global Z-scores, is employed to filter out false-positives. Extensive tests of the new protocol are presented. The new model provides an efficient alternative for threading with pair energies, maintaining comparable accuracy. The code, databases, and a prediction server are available at http://www.tc.cornell.edu/CBIO/loopp.

Amino Acid Sequence↗

Identification of novel transcribed sequences on human chromosome 22 by expressed sequence tag mapping.

To identify sequences on the human genome that are actually transcribed, we mapped expressed sequence tags (ESTs) of long cDNAs ranging from 4 kb to 7 kb along a 33.4-Mb sequence of human chromosome 22, the first human chromosome entirely sequenced. By the EST mapping of 30,683 long cDNAs in silico, 603 cDNA sequences were found to locate on chromosome 22 and classified into 169 clusters. Comparison of the genomic loci of these cDNA sequences with 679 genes already annotated on chromosome 22q revealed that 46 clusters represented newly identified transcribed sequences. To further characterize these sequences, we sequenced 12 cDNAs in their entirety out of 46 clusters. Of these 12 cDNAs, 6 were predicted to include a protein-coding region while the remaining 6 were unlikely to encode proteins. Interestingly, 3 out of the 12 cDNAs had the nucleotide sequences of the opposite strands of the genes previously annotated, which suggested that these genomic regions were transcribed bi-directionally. In addition to these newly identified 12 cDNAs, another 12 cDNAs were entirely sequenced since these cDNAs were likely to contain new information about the predicted protein-coding sequences previously annotated. In the cases of KIAA1670 and KIAA1672, these single cDNA sequences covered two separately annotated transcribed regions. For example, the sequence of a clone for KIAA1670 indicated that the CHKL and CPT1B genes were co-transcribed as a contiguous transcript without making both the protein-coding regions fused. In conclusion, the mapping of ESTs derived from long cDNAs followed by sequencing of the entire cDNAs provided indispensable information for the precise annotation of genes on the genome together with ESTs derived from short cDNAs.

Brain Chemistry↗

Overproduction and crystallization of FokI restriction endonuclease.

To overproduce FokI endonuclease (R.FokI) in an Escherichia coli system, the coding region of R.FokI predicted from the nucleotide sequence was generated from the FokI operon and joined to the tac promoter of an expression vector, pKK223-3. By introduction of the plasmid into E. coli UT481 cells expressing the FokI methylase gene, the R.FokI activity was overproduced about 30-fold, from which R.FokI was purified in amounts sufficient for crystallization. The removal of a stem-loop structure immediately upstream of the R.FokI coding region was essential for overproduction.

Amino Acid Sequence↗

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae↗