Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Cloning, sequencing, and expression of a thermostable cellulase gene of Humicola grisea.

The egl2 gene coding a thermostable endoglucanase (EGL2) was cloned from Humicola grisea. The DNA sequence of egl2 predicted two putative introns in the coding region. The deduced amino acid sequence of EGL2 was 388 amino acids in length and showed 99.5% identity with the H. insolens CMC 3. In addition to TATA box and CAAT motifs, putative CREA binding sites were observed in the 5' upstream region of the egl2 gene. The egl2 gene was expressed in Aspergillus oryzae, and EGL2 was purified. EGL2 produced by A. oryzae showed a high activity toward carboxymethyl cellulose. The optimal temperature of EGL2 was 75 degrees C, and EGL2 had more than 80% residual activity after heating up to 75 degrees C for 10 min. This is the first report of enzymatic properties of the EGL2-type thermostable cellulase homologs from Humicola.

Amino Acid Sequence↗

Transcript annotation in FANTOM3: mouse gene catalog based on physical cDNAs.

The international FANTOM consortium aims to produce a comprehensive picture of the mammalian transcriptome, based upon an extensive cDNA collection and functional annotation of full-length enriched cDNAs. The previous dataset, FANTOM2, comprised 60,770 full-length enriched cDNAs. Functional annotation revealed that this cDNA dataset contained only about half of the estimated number of mouse protein-coding genes, indicating that a number of cDNAs still remained to be collected and identified. To pursue the complete gene catalog that covers all predicted mouse genes, cloning and sequencing of full-length enriched cDNAs has been continued since FANTOM2. In FANTOM3, 42,031 newly isolated cDNAs were subjected to functional annotation, and the annotation of 4,347 FANTOM2 cDNAs was updated. To accomplish accurate functional annotation, we improved our automated annotation pipeline by introducing new coding sequence prediction programs and developed a Web-based annotation interface for simplifying the annotation procedures to reduce manual annotation errors. Automated coding sequence and function prediction was followed with manual curation and review by expert curators. A total of 102,801 full-length enriched mouse cDNAs were annotated. Out of 102,801 transcripts, 56,722 were functionally annotated as protein coding (including partial or truncated transcripts), providing to our knowledge the greatest current coverage of the mouse proteome by full-length cDNAs. The total number of distinct non-protein-coding transcripts increased to 34,030. The FANTOM3 annotation system, consisting of automated computational prediction, manual curation, and final expert curation, facilitated the comprehensive characterization of the mouse transcriptome, and could be applied to the transcriptomes of other species.

Animals↗

Accuracy of ICD-9-CM codes in detecting community-acquired pneumococcal pneumonia for incidence and vaccine efficacy studies.

Studies have used medical record discharge data as coded by the International Classification of Diseases, 9th Revision, Clinical Modification (ICD-9-CM) to estimate pneumococcal pneumonia incidence and vaccine efficacy. However, the accuracy of coding data to identify laboratory-confirmed pneumococcal pneumonia is not known. With the use of information collected in Ohio for a community-based pneumonia incidence study, the authors calculated the sensitivities, specificities, positive predictive values (PPV), and negative predictive values (NPV) of specific codes for pneumococcal pneumonia among hospitalized patients with community-acquired pneumonia. Sensitivities of the most common ICD-9-CM codes listed in the first five positions for patients with laboratory-confirmed pneumococcal pneumonia were 58.3% (code 481.0, pneumococcal pneumonia), 20.4% (38.0, streptococcal septicemia), 19.2% (38.2, pneumococcal septicemia), 15.0% (518.81, respiratory failure), 14.2% (486.0, pneumonia, organism unspecified), and 11.3% (482.3, streptococcal pneumonia). Using the first five listed ICD-9-CM codes rather than just the first listed code increased sensitivity without causing substantial change in specificity, PPV, and NPV. Sensitivity, PPV, and NPV of individual and groups of codes varied with different case definitions of pneumococcal pneumonia. Incidence and vaccine efficacy studies with the ability to validate diagnoses by medical chart review can use a combination of many ICD-9-CM codes to maximize sensitivity. However, without the ability to review medical charts, researchers must carefully decide which codes would best suit their studies.

Adolescent↗

Does diagnostic information contribute to predicting functional decline in long-term care?

BACKGROUND: Compared with the acute-care setting, use of risk-adjusted outcomes in long-term care is relatively new. With the recent development of administrative databases in long-term care, such uses are likely to increase. OBJECTIVES: The objective of this study was to determine the contribution of ICD-9-CM diagnosis codes from administrative data in predicting functional decline in long-term care. RESEARCH DESIGN: We used a retrospective sample of 15,693 long-term care residents in VA facilities in 1996. METHODS: We defined functional decline as an increase of > or =2 in the activities of daily living (ADL) summary score from baseline to semiannual assessment. A base regression model was compared to a full model enhanced with ICD-9-CM codes. We calculated validated measures of model performance in an independent cohort. RESULTS: The full model fit the data significantly better than the base model as indicated by the likelihood ratio test (chi2 = 179, df = 11, P <0.001). The full model predicted decline more accurately than the base model (R2 = 0.06 and 0.05, respectively) and discriminated better (c statistics were 0.70 and 0.68). Observed and predicted risks of decline were similar within deciles between the 2 models, suggesting good calibration. Validated R2 statistics were 0.05 and 0.04 for the full and base models; validated c statistics were 0.68 and 0.66. CONCLUSIONS: Adding specific diagnostic variables to administrative data modestly improves the prediction of functional decline in long-term care residents. Diagnostic information from administrative databases may present a cost-effective alternative to chart abstraction in providing the data necessary for accurate risk adjustment.

Activities of Daily Living↗

A chromosome-scale assembly for the genome of southern corn rootworm, Diabrotica undecimpunctata.

Diabrotica undecimpunctata ssp. howardi, the southern corn rootworm or eastern 12-spotted cucumber beetle, is a generalist insect herbivore that causes damage and yield loss to several crops in North America including maize. Unresolved phylogenetic relationships within and among D. undecimpunctata subspecies are impacting current quarantine policies. We report the chromosome-level haploid genome assembly, icDiaUnde3, constructed using HiFi and Hi-C read data from a single male D. undecimpunctata collected and identified as subspecies howardi based on geographic location and morphology. The primary 1.74 Gbp assembly is scaffolded into 11 chromosome-length scaffolds representing 9 autosomes, a single X chromosome and a supernumerary (B) chromosome (scaffold N50&#x2009;=&#x2009;162.8 Mb and L50&#x2009;=&#x2009;5). Ab initio and evidence-based structural reference sequence (RefSeq) annotations predicted 18,959 protein-coding genes, in which 99.2% of the 1,367 Benchmark Universal Single-Copy Orthologs from Insecta were complete. Repeat elements occupy 1.26 Gbp (72.33%) of the icDiaUnde3 assembly, with nearly 36% predicted to be retroelements. Alignment of whole chromosomes from icDiaUnde3 with those previously assembled from Diabrotica spp. predicted 2 and 6 autosomal inversions with D. balteata and D. virgifera virgifera, respectively. The mitochondrial genome had an annotated gene order and orientation conserved among beetles. The icDiaUnde3 reference genome assembly is a vital resource for taxonomic, comparative, and functional studies to enhance sustainable crop production.

agriculture↗

Validation of using EMS dispatch codes to identify low-acuity patients.

OBJECTIVE: To validate the predictive ability of previously derived emergency medical services (EMS) dispatch codes to identify patients with low-acuity illnesses. METHODS: This prospective descriptive study was conducted in Rochester, New York. An expert panel reviewed and modified a previously derived set of low-priority EMS dispatch codes. Patients assigned these 21 codes between July 2002 and June 2003 were included for further analysis. Dispatch data and level of EMS care were recorded for each dispatch code. The proportion of low-acuity patients (i.e., those who received only basic life support (BLS) care or those who were not transported using lights and sirens) was determined using previously established definitions. Codes were defined as associated with low-acuity patients if the lower bound of the 95% confidence interval (CI) exceeded 90%. Medical records for patients identified as high-acuity were reviewed to evaluate whether the advanced life support (ALS) level care that was provided had a clinical impact. RESULTS: Emergency medical services cared for 43,602 patients during the study, and 7,540 were dispatched as low-priority. We found that 7,197 (95%; 95% CI: 95-96%) of these patients met low-acuity criteria and that 11 of the evaluated codes were validated, with low-acuity care provided at least 90% of the time. Of the 343 patients identified as high-acuity, 62 (18%; 95% CI: 14-23%) were determined to have received interventions that had a clinical impact. CONCLUSIONS: This study prospectively validates 11 EMS dispatch codes as being associated with low-acuity patients. These codes could be used to triage EMS patients based on dispatch information.

Acute Disease↗

The characterisation of a cervine immunoregulatory cytokine, interleukin 12.

The cloning, sequencing, and production of cervine interleukin-12 is described. The cervine IL-12 p35 subunit coding sequence is 666 bp long and has highest homology to bovine p35 (94%), followed by human (79%), then murine (57%). The cDNA codes for a 221 aa long protein with predicted molecular weight of 24,902 Da. The cervine p40 subunit has a coding sequence of 984 bp and shows 96% homology to bovine, 85% homology to human, and 65% homology to murine p40 respectively. Cervine p40 cDNA codes for a 327 aa long protein with a predicted molecular weight of 37,461. Both subunits were inserted into a recombinant baculovirus that was then used to produce cervine IL-12 in Trichoplusia ni cells. Interleukin-12 was secreted into the culture medium and was biologically active as measured by proliferation of mitogen sensitised peripheral blood lymphocytes and the induction of interferon-gamma transcription in peripheral blood lymphocytes.

Amino Acid Sequence↗

Characterization of three heat-shock-protein genes and their developmental regulation during somatic embryogenesis in white spruce [Picea glauca (Moench) Voss].

Three cDNAs (PgEMB22, 27 and 29) predicted to encode low-molecular-weight (LMW) heat-shock proteins (HSPs) were cloned and characterized from white spruce [Picea glauca (Moench) Voss] somatic embryo tissues by differentially screening a cotyledonary embryo cDNA library. Clone PgEMB22 is predicted to encode a putative mitochondria-localized LMW HSP, and PgEMB27 and 29 are predicted to encode different cytoplasmic class II LMW HSPs, although they share 84.7% identity within DNA coding regions and 83.0% identity for predicted proteins. They are developmentally regulated during somatic embryo development and subsequent embryo germination, in addition they show strong response to heat-shock stress. Transcripts of the two kinds of hsp genes could be detected in embryogenic tissues before induction of embryo maturation, but subsequently increased, being most abundant at late embryo stages. Gene expression levels were very low or not detectable in germinated plantlets or needle tissues from older plants. Abscisic acid and polyethylene glycol, stimulators for spruce embryo maturation, could also induce the hsp genes.

Abscisic Acid↗

Sequence of the sheep interleukin-10-encoding cDNA.

The ovine interleukin-10 (oIL-10)-encoding cDNA has been cloned and sequenced using gene amplification by the polymerase chain reaction (PCR). We present the complete coding sequence of the ovine IL-10 gene, as well as the predicted amino acid (aa) sequence. The oIL10 DNA coding sequence is 531 nucleotides long and the mature protein product is predicted to be 18,367 Da, consisting of 158 aa, excluding a 19-aa N-terminal hydrophobic signal peptide. The oIL-10 protein is > 77% identical to pig and human IL-10, > 71% identical to rodent IL-10 and > 68% identical to viral IL-10.

Amino Acid Sequence↗

Phonological and orthographic coding skills in adult readers.

Although less skilled readers are handicapped by their poor phonological skills, this may not be true of their visual and orthographic coding skills. Because of an increasing reliance on visual-orthographic coding with reading experience, the author predicted that there would be smaller differences between skilled and less skilled adult readers on orthographic coding measures than on phonological coding measures. The orthographic and phonological coding measures involved, respectively, judgments of which looks more like a word, filv-fild, and which sounds like a real word, kake-dake? On the orthographic measure, reading groups did not differ in coding speed--although the less skilled readers made more errors, but far fewer than on the phonological coding measure. Differences between reading groups were substantial for both speed and errors on the phonological coding measure. Phonological variables accounted for most of the variance in word recognition, and this was especially true for men. The results suggest that in less skilled adult readers, phonological skills are a primary factor in their reading despite some evidence of visual-orthographic compensation.

Adolescent↗

Computer survey for likely genes in the one megabase contiguous genomic sequence data of Synechocystis sp. strain PCC6803.

Using the computer program GeneMark, the open reading frames (ORFs) previously assigned within the one megabase sequence data of the genome of the cyanobacterium, Synechocystis sp. strain PCC6803 (Kaneko et al., DNA Res. 2: 153-166, 1995), were re-examined. Matrices required by GeneMark for its statistical calculation were generated and modified by running a script termed GeneMark-Genesis that performed recursive application of GeneMark against the Synechocystis data and evaluated the probability scores for optimization. Based on the matrices thus generated, 752 of the 818 previously assigned ORFs (92%) were supported by GeneMark as likely coding sequences, of which 26 were predicted to start at more internal positions than previously assigned. In addition, 50 ORFs were newly identified as likely coding sequences, most of them being shorter than 300 bp. Thus, the procedure was proven to be very powerful to locate likely coding regions within the genomic sequence data of Synechocystis without having prior information concerning their similarity to the genes of other organisms. However, GeneMark did not predict 66 previously assigned ORFs as likely genes: 14 of them showed significant degrees of similarity to known genes and 10 others were found within IS-like elements. It seems that these genes, many of which appear to be exogenous origin, escaped detection by GeneMark as in the case of "class 3 (horizontally transferred) genes" of E. coli, which in turn suggests that genes of different phylogenetic origins might also be detected as such by modifying the matrices.

Base Sequence↗

The Gene-Finder computer tools for analysis of human and model organisms genome sequences.

We present a complex of new programs for promoter, 3'-processing, splice sites, coding exons and gene structure identification in genomic DNA of several model species. The human gene structure prediction program FGENEH, exon prediction-FEXH and splice site prediction-HSPL have been modified for sequence analysis of Drosophila (FGENED, FEXD and DSPL), C.elegance (FGENEN, FEXN and NSPL), Yeast (FEXY and YSPL) and Plant (FGENEA, FEXA and ASPL) genomic sequences. We recomputed all frequency and discriminant function parameters for these organisms and adjusted organism specific minimal intron lengths. An accuracy of coding region prediction for these programs is similar with the observed accuracy of FEXH and FGENEH. We have developed FEXHB and FGENEHB programs combining pattern recognition features and information about similarity of predicted exons with known sequences in protein databases. These programs have approximately 10% higher average accuracy of coding region recognition. Two new programs for human promoter site prediction (TSSG and TSSW) have been developed which use Gosh (1993) and Wingender (1994) data bases of functional motifs, respectively. POLYAH program was designed for prediction of 3'-processing regions in human genes and CDSB program was developed for bacterial gene prediction. We have developed a new approach to predict multiple genes based on double dynamic programming, that is very important for analysis of long genomic DNA fragments generated by genome sequencing projects. Analysis of uncharacterized sequences based on our methods is available through the University of Houston, Weizmann Institute of Science email servers and several Web pages at Baylor College of Medicine.

Animals↗

Pombe: a gene-finding and exon-intron structure prediction system for fission yeast.

A special program developed by the authors, called Pombe, identifies protein coding regions in the Schizosaccharomyces pombe genome. Linear discriminant analysis was applied to predict 5'-terminal, internal, 3'-terminal exons (coding-exon) and introns. The accuracy of the prediction was tested by cross verifications. The sensitivity, specificity and correlation coefficient for the internal exon prediction were 98.5%, 99.9% and 98.3% respectively at the nucleotide level. Open reading frames were studied and used to predict intron-less genes: 99.0% of such genes were identified with correct stopping sites. The gene structure was determined by dynamic programming and the prediction achieved 97.0% correlation coefficient at the nucleotide level. The program is available at http:(/)/clio.cshl.org/genefinder.

Algorithms↗

Sequence divergence among members of a trypanosome variant surface glycoprotein gene family.

We have used analysis of DNA sequence data from four members of a Trypanosoma brucei variant surface glycoprotein gene family to investigate the molecular basis of the generation of antigenic diversity in African trypanosomes. Among these four sequences we find the greatest similarity in the untranslated sequences immediately upstream from the coding region. A complex pattern of nucleic acid and predicted amino acid sequence divergence appears starting at the coding sequence. Two related but highly divergent hydrophobic leaders are associated with different members of this gene family; both forms of these hydrophobic leaders appear to exist in other isolates of T. b. brucei. We find conservative replacements in the first 120 predicted amino acid residues of the mature protein; the following 80 predicted residues show less conservative replacements, and we suggest that this region may be hypervariable and exposed to the aqueous environment.

Amino Acid Sequence↗

Transcranial color-coded duplex sonography in the evaluation of collateral flow through the circle of Willis.

PURPOSE: To determine the sensitivity, specificity, and positive and negative predictive values of transcranial color-coded duplex sonographic (TCCD) evaluation of cross flow through the anterior (ACoA) and posterior (PCoA) communicating arteries in patients with occlusive cerebrovascular disease. METHODS: We studied prospectively 132 patients (37 women, 95 men; mean age, 60 years) with stenoses of more than 69% reduction in vessel diameter (n = 93) and occlusions (n = 52) of the internal carotid artery, and three occlusions of the basilar artery. The sonographer was aware of extracranial sonographic findings but was blinded to the results of cerebral angiography. RESULTS: Nine patients (7%) with thick bones preventing transtemporal insonation and three patients (3%) with occlusions of the middle (n = 3) and anterior (n = 1) cerebral arteries were excluded. Sensitivity of TCCD for detection of collateral flow through the ACoA in patients with occlusive carotid artery disease was 98%, specificity was 100%, positive predictive value was 100%, and negative predictive value was 98%. The corresponding values for the PCoA were 84%, 94%, 94%, and 84%, respectively. All three functional PCoAs were identified in patients with occluded basilar arteries. CONCLUSION: TCCD is a valuable method for noninvasive evaluation of cross flow through the ACoA in patients with adequate sonographic windows. However, TCCD evaluation of cross flow through the PCoA is less reliable, because hemodynamic criteria may cause falsely positive and falsely negative results.

Adult↗

Effects of target fragmentation on evaluation of LET spectra from space radiation in low-earth orbit (LEO) environment: impact on SEU predictions.

Recent improvements in the radiation transport code HZETRN/BRYNTRN and galactic cosmic ray environmental model have provided an opportunity to investigate the effects of target fragmentation on estimates of single event upset (SEU) rates for spacecraft memory devices. Since target fragments are mostly of very low energy, an SEU prediction model has been derived in terms of particle energy rather than linear energy transfer (LET) to account for nonlinear relationship between range and energy. Predictions are made for SEU rates observed on two Shuttle flights, each at low and high inclination orbit. Corrections due to track structure effects are made for both high energy ions with track structure larger than device sensitive volume and for low energy ions with dense track where charge recombination is important. Results indicate contributions from target fragments are relatively important at large shield depths (or any thick structure material) and at low inclination orbit. Consequently, a more consistent set of predictions for upset rates observed in these two flights is reached when compared to an earlier analysis with CREME model. It is also observed that the errors produced by assuming linear relationship in range and energy in the earlier analysis have fortuitously canceled out the errors for not considering target fragmentation and track structure effects.

Aluminum↗

Prenatal screening of pregnant mothers for parenting difficulties: final results from the Queen Mary Child Care Unit.

This paper reports the results of 10 years of research into the prenatal identification of mothers likely to have major parenting problems. Previous published research reported the development of a set of criteria for determining risk status. These criteria were used to classify into four levels of risk a sample of mothers who were consecutive enrollments for prenatal care. The sample was monitored through various social agencies for 2 years. Results of this monitoring indicate the predictive validity of the risk code in an unselected sample. The value of prenatal identification of the 'at risk' is discussed together with the procedures adopted for implementing routine screening in the maternity hospital. The issue of causation, as distinct from prediction, is addressed.

Child↗

Involvement of hyperprolinemia in cognitive and psychiatric features of the 22q11 deletion syndrome.

Microdeletions of the 22q11 region, responsible for the velo-cardio-facial syndrome (VCFS), are associated with an increased risk for psychosis and mental retardation. Recently, it has been shown in a hyperprolinemic mouse model that an interaction between two genes localized in the hemideleted region, proline dehydrogenase (PRODH) and catechol-o-methyl-transferase (COMT), could be involved in this phenotype. Here, we further characterize in eight children the molecular basis of type I hyperprolinemia (HPI), a recessive disorder resulting from reduced activity of proline dehydrogenase (POX). We show that these patients present with mental retardation, epilepsy and, in some cases, psychiatric features. We next report that, among 92 adult or adolescent VCFS subjects, a subset of patients with severe hyperprolinemia has a phenotype distinguishable from that of other VCFS patients and reminiscent of HPI. Forward stepwise multiple regression analysis selected hyperprolinemia, psychosis and COMT genotype as independent variables influencing IQ in the whole VCFS sample. An inverse correlation between plasma proline level and IQ was found. In addition, as predicted from the mouse model, hyperprolinemic VCFS subjects bearing the Met-COMT low activity allele are at risk for psychosis (OR = 2.8, 95% CI = 1.04-7.4). Finally, from the extensive analysis of the PRODH gene coding sequence variations, it is predicted that POX residual activity in the 0-30% range results into HPI, whereas residual activity in the 30-50% range is associated either with normal plasma proline levels or with mild-to-moderate hyperprolinemia.

Adolescent↗