Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Alteration of amino-terminal codons of human granulocyte-colony-stimulating factor increases expression levels and allows efficient processing by methionine aminopeptidase in Escherichia coli.

We have improved the expression of recombinant human granulocyte-colony-stimulating factor (G-CSF), produced by either pL or trpP expression vectors in Escherichia coli, by altering the sequence at the 5' end of the G-CSF-coding region. Initial attempts to express G-CSF resulted in neither detectable G-CSF mRNA nor protein in the trpP system, and only G-CSF mRNA was detectable in the pL system. We modified both expression vectors to decrease the G + C content of the 5' end of the coding region without altering the predicted amino acid sequence. This resulted in expression of detectable G-CSF mRNA and protein in both systems. Expression reached 17% and 6.5% of the total soluble cellular protein in the pL and trpP expression systems, respectively. The N-terminal sequence of the recombinant G-CSF from the pL system was Met-Thr-Pro-Leu-Gly-Pro-. G-CSF isolated from several human cell lines (including the LD-1 cell line reported here), does not have an N-terminal methionyl residue. Deletion of the threonine codon at the beginning of the coding region for the mature G-CSF resulted in efficient removal of the N-terminal methionine residue during expression in E. coli.

Aminopeptidases↗

Nucleolar localization of rRNA coding sequences in Prorocentrum micans Ehr. (dinomastigote, kingdom Protoctist) by in situ hybridization.

To define the molecular mechanisms of ribosome biogenesis and to find out in which nucleolar compartment transcription of rDNA occurs, we have performed in situ hybridization (ISH) of RNase-treated cryosections using biotinylated rRNA coding sequences as a probe and the eukaryotic dinoflagellate nucleolar system as a model. Recent data from ISH of eukaryotic ribosomal genes by electron microscopy (EM) has so far failed to establish a consensus which clearly defines the function of the three compartments of the nucleolus. Dinomastigote protoctists are the only known eukaryotes whose chromatin is totally devoid of nucleosomes. Their chromosomes remain permanently condensed during the entire cell cycle and active nucleoli arise from an unwound part of some of the otherwise compact chromosomes. In this work, DNA-DNA hybrids were detected either by fluorescent avidin or by indirect immunogold staining procedures in EM; this is the first use of cryosections to detect hybrids in EM not only in the nucleolus sensu lato but also in a dinomastigote cell. Coding sequences of ribosomal genes were detected both in the periphery of the nucleolar organizer region (NOR), which corresponds to the unwound part of the nucleolar chromosome, and in the proximal part of the fibrillo-granular (FG) region. These results suggest that the rRNA gene transcription predominantly occurs at the periphery of the NOR where the coding sequences are located. A predictive model summarizes and allows discussions and comparisons with other eukaryotes in which nucleolar mechanisms were previously studied. This leads to the conclusion that dinoflagellate cells constitute an excellent model for the study of the functional structure of the eukaryotic nucleolus.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Vertebrate gene predictions and the problem of large genes.

To find unknown protein-coding genes, annotation pipelines use a combination of ab initio gene prediction and similarity to experimentally confirmed genes or proteins. Here, we show that although the ab initio predictions have an intrinsically high false-positive rate, they also have a consistently low false-negative rate. The incorporation of similarity information is meant to reduce the false-positive rate, but in doing so it increases the false-negative rate. The crucial variable is gene size (including introns)--genes of the most extreme sizes, especially very large genes, are most likely to be incorrectly predicted.

Animals↗

Effects of response eccentricity and relative position on orthogonal stimulus-response compatibility with joystick and keypress responses.

When unimanual left-right movement responses are made to up-down stimuli, performance is better with the up-right/down-left mapping when responding in the right hemispace and with the up-left/down-right mapping when responding in the left hemispace. We evaluated whether this response eccentricity effect is explained best in terms of rotational properties of the hand (the end-state comfort hypothesis) or asymmetric coding of the stimulus and response alternatives (the salient features coding hypothesis). Experiment 1 showed that bimanual keypresses yield a response eccentricity effect similar to that obtained with unimanual movement responses. In Experiment 2, an inactive response apparatus was placed to the left or right of the active response apparatus to provide a referent. For half of the participants, the active and inactive apparatuses were joysticks, and for half they were response boxes with keys. For both response types, an up-right/down-left advantage was evident when the relative position of the active response apparatus was right but not when it was left. That bimanual keypresses yield similar eccentricity and relative location effects to those for unimanual movements is predicted by the salient features coding perspective but not by the end-state comfort hypothesis.

Functional Laterality↗

Specific binding of sso II DNA methyltransferase to its promoter region provides the regulation of sso II restriction-modification gene expression.

The regulation of the Sso II restriction-modification system from Shigella sonnei was studied in vivo and in vitro . In lacZ fusion experiments, Sso II methyltransferase (M. Sso II) was found to repress its own synthesis but stimulate expression of the cognate restriction endonuclease (ENase). The N-terminal 72 amino acids of M. Sso II, predicted to form a helix-turn-helix (HTH) motif, was found to be responsible for the specific DNA-binding and regulatory function of M. Sso II. Similar HTH motifs are predicted in the N-terminus of a number of 5-methylcytosine methyltransferases, particularly M. Eco RII, M.dcm and M. Msp I, of which the ability to regulate autogenously has been proposed. In vitro, the binding of M. Sso II to its target DNA was investigated using a mobility shift assay. M. Sso II forms a specific and stable complex with a 140 bp DNA fragment containing the promoter region of Sso II R-M system. The dissociation constant (Kd) was determined to be 1.5x10(-8) M. DNaseI footprinting experiments demonstrated that M. Sso II protects a 48-52 bp region immediately upstream of the M. Sso II coding sequence which includes the predicted -10 promoter sequence of M. Sso II and the -10 and -35 sequences of R. Sso II.

Amino Acid Sequence↗

Predicting in-hospital deaths from coronary artery bypass graft surgery. Do different severity measures give different predictions?

OBJECTIVES: Severity-adjusted death rates for coronary artery bypass graft (CABG) surgery by provider are published throughout the country. Whether five severity measures rated severity differently for identical patients was examined in this study. METHODS: Two severity measures rate patients using clinical data taken from the first two hospital days (MedisGroups, physiology scores); three use diagnoses and other information coded on standard, computerized hospital discharge abstracts (Disease Staging, Patient Management Categories, all patient refined diagnosis related groups). The database contained 7,764 coronary artery bypass graft patients from 38 hospitals with 3.2% in-hospital deaths. Logistic regression was performed to predict deaths from age, age squared, sex, and severity scores, and c statistics from these regressions were used to indicate model discrimination. Odds ratios of death predicted by different severity measures were compared. RESULTS: Code-based measures had better c statistics than clinical measures: all patient refined diagnosis related groups, c = 0.83 (95% C.I. 0.81, 0.86) versus MedisGroups, c = 0.73 (95% C.I. 0.70, 0.76). Code-based measures predicted very different odds of dying than clinical measures for more than 30% of patients. Diagnosis codes indicting postoperative, life-threatening conditions may contribute to the superior predictive power of code-based measures. CONCLUSIONS: Clinical and code-based severity measures predicted different odds of dying for many coronary artery bypass graft patients. Although code-based measures had better statistical performance, this may reflect their reliance on diagnosis codes for life-threatening conditions occurring late in the hospitalization, possibly as complications of care. This compromises their utility for drawing inferences about quality of care based on severity-adjusted coronary artery bypass graft death rates.

Coronary Artery Bypass↗

Computational detection of genomic cis-regulatory modules applied to body patterning in the early Drosophila embryo.

BACKGROUND: Regulation of gene transcription is crucial for the function and development of all organisms. While gene prediction programs that identify protein coding sequence are used with remarkable success in the annotation of genomes, the development of computational methods to analyze noncoding regions and to delineate transcriptional control elements is still in its infancy. RESULTS: Here we present novel algorithms to detect cis-regulatory modules through genome wide scans for clusters of transcription factor binding sites using three levels of prior information. When binding sites for the factors are known, our statistical segmentation algorithm, Ahab, yields about 150 putative gap gene regulated modules, with no adjustable parameters other than a window size. If one or more related modules are known, but no binding sites, repeated motifs can be found by a customized Gibbs sampler and input to Ahab, to predict genes with similar regulation. Finally using only the genome, we developed a third algorithm, Argos, that counts and scores clusters of overrepresented motifs in a window of sequence. Argos recovers many of the known modules, upstream of the segmentation genes, with no training data. CONCLUSIONS: We have demonstrated, in the case of body patterning in the Drosophila embryo, that our algorithms allow the genome-wide identification of regulatory modules. We believe that Ahab overcomes many problems of recent approaches and we estimated the false positive rate to be about 50%. Argos is the first successful attempt to predict regulatory modules using only the genome without training data. Complete results and module predictions across the Drosophila genome are available at http://uqbar.rockefeller.edu/~siggia/.

Algorithms↗

A truncated isoform of Ca2+/calmodulin-dependent protein kinase II expressed in human islets of Langerhans may result from trans-splicing.

Calcium/calmodulin-dependent protein kinase II (CaM kinase II) has been proposed to play a key role in glucose stimulated insulin secretion. Using the rapid amplification of cDNA ends technique we amplified the 3' end of the CaM kinase II gamma gene from human islet RNA. A novel cDNA was detected composed of 5' sequence from the human CaM kinase II gamma gene joined to the 3' end of the human signal recognition particle 72 (SRP72) gene. We predict that this mRNA species will code for a truncated form of CaM kinase II, designated gammaSRP, comprising the entire catalytic and regulatory domains of the protein and with a predicted molecular weight of 37 kDa. We mapped the human SRP72 gene to chromosome 18 and, as the CaM kinase II gamma gene was previously mapped to human chromosome 10q22, we suggest this novel cDNA may have resulted from trans-splicing.

Alternative Splicing↗

Gene structure conservation aids similarity based gene prediction.

One of the primary tasks in deciphering the functional contents of a newly sequenced genome is the identification of its protein coding genes. Existing computational methods for gene prediction include ab initio methods which use the DNA sequence itself as the only source of information, comparative methods using multiple genomic sequences, and similarity based methods which employ the cDNA or protein sequences of related genes to aid the gene prediction. We present here an algorithm implemented in a computer program called Projector which combines comparative and similarity approaches. Projector employs similarity information at the genomic DNA level by directly using known genes annotated on one DNA sequence to predict the corresponding related genes on another DNA sequence. It therefore makes explicit use of the conservation of the exon-intron structure between two related genes in addition to the similarity of their encoded amino acid sequences. We evaluate the performance of Projector by comparing it with the program Genewise on a test set of 491 pairs of independently confirmed mouse and human genes. It is more accurate than Genewise for genes whose proteins are <80% identical, and is suitable for use in a combined gene prediction system where other methods identify well conserved and non-conserved genes, and pseudogenes.

Algorithms↗

Coding of perineal lacerations and other complications of obstetric care in hospital discharge data.

OBJECTIVE: To assess the validity of obstetric complications, including the Joint Commission on Accreditation of Healthcare Organizations (JCAHO) Core Measure on perineal lacerations, in the California Patient Discharge Data Set. METHODS: We randomly sampled 1,611 deliveries from 52 of the 267 hospitals that performed more than 678 eligible deliveries in California in 1992-1993. We compared hospital-reported complications against our recoding of the same records. RESULTS: Third- and fourth-degree perineal lacerations were reported accurately, with estimated sensitivities exceeding 90% and positive predictive values exceeding 65% (weighted to account for the stratified sampling design) or 85% (unweighted). Based on in-depth review of discrepant cases, we estimate the actual positive predictive value at over 90%. Most coding discrepancies were between no injury and first degree, or between first and second degree. Most postpartum complications, including urinary tract and wound infections, endometritis, anesthesia complications, and postpartum hemorrhage were reported with less than 70% sensitivity, but at least 80% positive predictive value. Composite measures from HealthGrades and Solucient, which include these complication codes, also suffer from high false-negative rates. CONCLUSION: Third- and fourth-degree perineal lacerations are accurately reported on hospital discharge abstracts, confirming the validity of related quality indicators sponsored by the Agency for Healthcare Research and Quality and JCAHO. Administrative data seem less useful for monitoring other in-hospital postpartum complications.

Adult↗

Construction of an immunogenic cell death-related LncRNA signature to predict the prognosis of patients with lung adenocarcinoma.

BACKGROUND: Lung adenocarcinoma (LUAD) is one of the most common malignant diseases worldwide. This study aimed to construct an immunogenic cell death (ICD)-related long non-coding RNA (lncRNA) signature to effectively predict the prognosis of LUAD. METHODS: The RNA-sequencing and clinical data of LUAD were downloaded from The Cancer Genome Atlas (TCGA). Least absolute shrinkage and selection operator (LASSO) and stepwise multivariate Cox proportional hazard regression analysis were utilized to construct lncRNA signature. Then, the reliability of the signature was evaluated in the training, validation and whole cohorts. The differences in the immune landscape and drug sensitivity between the low- and high-risk groups were analyzed. Finally, the expression level of the selected ICD-related lncRNAs in LUAD cell lines via reverse transcription quantitative PCR (RT-qPCR). CCK-8 and transwell assays were performed to study biological function of AC245014.3. RESULTS: A signature consisting of 5 ICD-related lncRNAs was constructed. Kaplan Meier (K-M) survival analysis showed shorter overall survival (OS) in high-risk group. The receiver operating characteristic (ROC) curves and Multivariate Cox regression analysis showed the signature was good predictive and independent prognostic factor in LUAD. Moreover, the high-risk group had a lower level of antitumor immunity and was less sensitive to some chemotherapeutics and targeted drugs. Finally, the expression level of selected ICD-related lncRNAs was validated in LUAD cell lines by RT-qPCR. Knockdown of AC245014.3 significantly suppressed LUAD proliferation, migration and invasion. CONCLUSIONS: In this study, an ICD-related lncRNA signature was constructed, which could accurately predict the prognosis of LUAD patients and guide clinical treatment.

Humans↗

Measuring potentially avoidable hospital readmissions.

The objectives of this study were to develop a computerized method to screen for potentially avoidable hospital readmissions using routinely collected data and a prediction model to adjust rates for case mix. We studied hospital information system data of a random sample of 3,474 inpatients discharged alive in 1997 from a university hospital and medical records of those (1,115) readmitted within 1 year. The gold standard was set on the basis of the hospital data and medical records: all readmissions were classified as foreseen readmissions, unforeseen readmissions for a new affection, or unforeseen readmissions for a previously known affection. The latter category was submitted to a systematic medical record review to identify the main cause of readmission. Potentially avoidable readmissions were defined as a subgroup of unforeseen readmissions for a previously known affection occurring within an appropriate interval, set to maximize the chance of detecting avoidable readmissions. The computerized screening algorithm was strictly based on routine statistics: diagnosis and procedures coding and admission mode. The prediction was based on a Poisson regression model. There were 454 (13.1%) unforeseen readmissions for a previously known affection within 1 year. Fifty-nine readmissions (1.7%) were judged avoidable, most of them occurring within 1 month, which was the interval used to define potentially avoidable readmissions (n = 174, 5.0%). The intra-sample sensitivity and specificity of the screening algorithm both reached approximately 96%. Higher risk for potentially avoidable readmission was associated with previous hospitalizations, high comorbidity index, and long length of stay; lower risk was associated with surgery and delivery. The model offers satisfactory predictive performance and a good medical plausibility. The proposed measure could be used as an indicator of inpatient care outcome. However, the instrument should be validated using other sets of data from various hospitals.

Adolescent↗

Identification of cDNA clones encoding a precursor of rat liver cathepsin B.

Recent studies have suggested that many lysosomal enzymes, including cathepsin B (EC 3.4.22.1), may be synthesized as larger precursors and proteolytically processed to their mature forms. To determine the structure of the primary translation product of cathepsin B, we have screened a phage cDNA library for clones encoding rat liver cathepsin B. We synthesized two extended DNA oligonucleotides to use as hybridization probes: a 50-mer corresponding to the coding segment for residues 215-231 of mature cathepsin B and a 54-mer corresponding to residues 117-134. After screening 600,000 plaques, five clones were obtained that hybridized to the 32P-labeled 50-mer; of these, two (lambda rCB3 and lambda rCB5) also reacted with the 54-mer. DNA sequence analysis confirmed that lambda rCB3 and lambda rCB5 both encoded rat liver cathepsin B, and the translated sequence is in agreement with the sequence determined [Takio, K., Towatari, T., Katunuma, N., Teller, D. C. & Titani, K. (1983) Proc. Natl. Acad. Sci. USA 80, 3666-3670], except for a tryptophan for glycine substitution at residue 78 and the presence of two amino acids at the junction site of the light and heavy chains. Moreover, the DNA sequence reveals an open reading frame extending beyond the 5' (NH2 terminus), and the predicted COOH terminus of the coding sequence for the mature protein is extended by six amino acids. These results confirm that the biosynthesis of cathepsin B involves a larger precursor form and demonstrate the effectiveness of long oligonucleotide probes for screening to detect rare cloned mRNAs.

Animals↗

Retroviral-mediated transfer and expression of hepatitis B e antigen in human primary skin fibroblasts and Epstein-Barr virus-transformed B lymphocytes.

Previously, an amphotropic retroviral expression system coding for the neomycin resistance gene was developed and used to synthesize hepatitis B e antigen (HBeAg) and hepatitis B core/e antigen (HBc/eAg) in transfected mouse NIH 3T3 fibroblasts (A. McLachlan et al., 1987, J. Virol. 61, 683-692). In the present study, these transfected cell lines were infected with a helper amphotropic murine leukemia virus resulting in the production of infectious recombinant retrovirus. The recombinant retrovirus was examined for its capacity to transmit resistance to the antibiotic, G418, and to express hepatitis B virus antigens in mouse NIH 3T3 fibroblasts, human primary skin fibroblasts, and Epstein-Barr virus (EBV)-transformed B lymphocytes. A mouse NIH 3T3 fibroblast clone was generated which produced recombinant retrovirus with the capacity to transmit HBeAg expression to these murine and human cell lines. In contrast, it was not possible to transmit HBc/eAg synthesis efficiently to these cell lines by recombinant retroviral infection. The difference between the efficiencies of transmission of HBeAg and HBc/eAg expression by recombinant retroviral-mediated infection was not predicted as the expression vector coding for HBc/eAg synthesis differs only by the deletion of approximately 90 nucleotides of HBV DNA sequence from the vector coding for HBeAg synthesis.

Adult↗

Variability and transmission by Aphis glycines of North American and Asian Soybean mosaic virus isolates.

The variability of North American and Asian strains and isolates of Soybean mosaic virus was investigated. First, polymerase chain reaction (PCR) products representing the coat protein (CP)-coding regions of 38 SMVs were analyzed for restriction fragment length polymorphisms (RFLP). Second, the nucleotide and predicted amino acid sequence variability of the P1-coding region of 18 SMVs and the helper component/protease (HC/Pro) and CP-coding regions of 25 SMVs were assessed. The CP nucleotide and predicted amino acid sequences were the most similar and predicted phylogenetic relationships similar to those obtained from RFLP analysis. Neither RFLP nor sequence analyses of the CP-coding regions grouped the SMVs by geographical origin. The P1 and HC/Pro sequences were more variable and separated the North American and Asian SMV isolates into two groups similar to previously reported differences in pathogenic diversity of the two sets of SMV isolates. The P1 region was the most informative of the three regions analyzed. To assess the biological relevance of the sequence differences in the HC/Pro and CP coding regions, the transmissibility of 14 SMV isolates by Aphis glycines was tested. All field isolates of SMV were transmitted efficiently by A. glycines, but the laboratory isolates analyzed were transmitted poorly. The amino acid sequences from most, but not all, of the poorly transmitted isolates contained mutations in the aphid transmission-associated DAG and/or KLSC amino acid sequence motifs of CP and HC/Pro, respectively.

Animals↗

Sequence and expression of hamster prolactin and growth hormone messenger RNAs.

Complementary DNAs encompassing the complete protein-encoding regions for PRL and GH of the Syrian Golden hamster were sequenced and used as probes to examine the expression of hamster PRL and GH messenger RNA (mRNA)s. The complementary DNA (cDNA) for hamster PRL encodes a 226 amino acid preprotein which, by analogy to rat and mouse PRLs, is predicted to be processed to yield a 197 amino acid secreted protein. The hamster GH cDNA codes for a 216 amino acid preprotein predicted to yield a 190 amino acid secreted protein. Both hamster proteins are highly homologous to the corresponding rat and mouse hormones. For the secreted proteins, hamster PRL has 82% amino acid identity with rat PRL and 72% identity with mouse PRL. The rodent GH sequences are more strongly conserved, with 97-98% sequence identity between hamster, rat, and mouse GHs. The hamster hormones contain the highly conserved cysteine residues (six in hamster PRL and four in hamster GH) present in other mammalian PRLs and GHs. Neither hamster PRL nor hamster GH contains cysteine residues corresponding to the unique pair of cysteines present in hamster placental lactogen-II. The hamster PRL and GH cDNAs each hybridized to pituitary mRNAs of approximately 1 kilobase. Expression of hamster PRL and GH mRNAs was compared between 2 days of the estrous cycle (proestrus and estrus) and early, mid, and late pregnancy (days 5, 10, and 15). PRL mRNA levels in cycling hamsters were approximately 50% of those in pregnant hamsters. No other significant differences in PRL or GH mRNA levels were observed, suggesting that differences in circulating PRL and GH protein levels during the estrous cycle and pregnancy in the hamster are the result largely of factors other than changes in mRNA levels.

Amino Acid Sequence↗

Gene prediction using the Self-Organizing Map: automatic generation of multiple gene models.

BACKGROUND: Many current gene prediction methods use only one model to represent protein-coding regions in a genome, and so are less likely to predict the location of genes that have an atypical sequence composition. It is likely that future improvements in gene finding will involve the development of methods that can adequately deal with intra-genomic compositional variation. RESULTS: This work explores a new approach to gene-prediction, based on the Self-Organizing Map, which has the ability to automatically identify multiple gene models within a genome. The current implementation, named RescueNet, uses relative synonymous codon usage as the indicator of protein-coding potential. CONCLUSIONS: While its raw accuracy rate can be less than other methods, RescueNet consistently identifies some genes that other methods do not, and should therefore be of interest to gene-prediction software developers and genome annotation teams alike. RescueNet is recommended for use in conjunction with, or as a complement to, other gene prediction methods.

Chromosome Mapping↗

DNA sequence of the adenovirus type 41 hexon gene and predicted structure of the protein.

The gene for the major capsid protein (hexon) of human adenovirus type 41 (Ad41) has been isolated and the complete DNA sequence determined. Comparison of the predicted amino acid sequence with hexons from human Ad2 and Ad5 and bovine adenovirus type 3 reveals regions of high homology at the N and C termini separated by a central region of low homology. Fitting of the Ad41 hexon sequence to the known three-dimensional structure of the Ad2 hexon demonstrates that both hexons have a common architecture. Regions of the hexon which in the trimer constitute the pseudohexagonal base are highly conserved, with the major amino acid changes concentrated in the domains forming the triangular towers which represent the surface of the capsid. Changes in the Ad41 towers therefore permit the virus to present a unique surface to the environment while conservation of residues in the base maintains the integrity of hexon-hexon contacts. A striking difference is the absence in the Ad41 sequence of 32 amino acids which are present in the Ad2 sequence. In Ad2 this region is highly charged and may be responsible for pH-induced conformational changes within the virus capsid. The DNA sequence in the region surrounding the Ad41 hexon gene was also determined and revealed an open reading frame which appeared to code for the homologue of the Ad2-coded endoprotease. Comparison of the predicted amino acid sequences of the Ad41 and Ad2 proteins revealed a high degree of homology suggesting that this protein may have an important role in the infectious cycle of the virus.

Adenoviruses, Human↗