Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Nucleotide sequence of the EcoRI D fragment of adenovirus 2 genome.

The entire nucleotide sequence of the Ad. 2 EcoRI D fragment has been determined using the Maxam and Gilbert method. This sequence of 2678 bp contains informations relative to late mRNAs ending at position 78 and for which an AATAAA sequence corresponding to their 3' ends is found at residue number 833. Position of the PVIII mRNA is determined thus allowing deduction of the probable amino acid sequence of the PVIII protein. The position and the sequence of the first leader of early 3 mRNAs is determined as well as the sequence and position of the second early leader of region 3 mRNAs, which also correspond to the "y" leader of the fiber mRNA. Following the localization of an open reading frame in which an ATG could initiate protein synthesis it can be predicted that 3a, b, c mRNAs code for the 16K early protein and the probable amino acid sequence of this protein can be deduced. The CAGTTT sequence frequently present at the 5' end of a leader or of a mRNA body as well as the GGTGAG sequence which is found at the 3' end of several leaders were used to postulate the position of various early mRNAs of region 3 and to suggest the existence of an additional splicing event during the processing of mRNAs 3a, b and c. They were also used to predict the position of the additional "x" late leaders. The imbrication of information concerning (i) the family of late mRNAs ending at position 78, (ii) the position of the "x" leader and the "y" leader and (iii) the beginning of early region 3 is also depicted.

Adenoviruses, Human↗

GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.

Improving the accuracy of prediction of gene starts is one of a few remaining open problems in computer prediction of prokaryotic genes. Its difficulty is caused by the absence of relatively strong sequence patterns identifying true translation initiation sites. In the current paper we show that the accuracy of gene start prediction can be improved by combining models of protein-coding and non-coding regions and models of regulatory sites near gene start within an iterative Hidden Markov model based algorithm. The new gene prediction method, called GeneMarkS, utilizes a non-supervised training procedure and can be used for a newly sequenced prokaryotic genome with no prior knowledge of any protein or rRNA genes. The GeneMarkS implementation uses an improved version of the gene finding program GeneMark.hmm, heuristic Markov models of coding and non-coding regions and the Gibbs sampling multiple alignment program. GeneMarkS predicted precisely 83.2% of the translation starts of GenBank annotated Bacillus subtilis genes and 94.4% of translation starts in an experimentally validated set of Escherichia coli genes. We have also observed that GeneMarkS detects prokaryotic genes, in terms of identifying open reading frames containing real genes, with an accuracy matching the level of the best currently used gene detection methods. Accurate translation start prediction, in addition to the refinement of protein sequence N-terminal data, provides the benefit of precise positioning of the sequence region situated upstream to a gene start. Therefore, sequence motifs related to transcription and translation regulatory sites can be revealed and analyzed with higher precision. These motifs were shown to possess a significant variability, the functional and evolutionary connections of which are discussed.

Algorithms↗

Use of a template to improve documentation and coding.

BACKGROUND AND OBJECTIVES: Accurate assignment of evaluation and management (E&M) codes is a challenge for physicians. Having guidelines close at hand during patient visits might improve appropriateness and accuracy of E&M coding. We developed a template based on a clinical prediction rule for group A beta-hemolytic streptococcal (GABHS) pharyngitis to improve documentation and coding decisions. METHODS: Fifty office visits for sore throat were documented using templates and were compared with 50 sore throat visits that were documented using progress notes. We counted history and physical examination items and compared the level of service charged to the level of service supported by the note. RESULTS: Significantly more history of present illness and physical examination items were recorded on templates. Decisions related to treatment for patients with a low probability of GABHS were also improved by the templates. Templates had no effect on billing and coding errors. CONCLUSIONS: The template resulted in more-thorough documentation but had no effect on coding and billing errors relative to progress notes.

Adult↗

NSCL-2: a basic domain helix-loop-helix gene expressed in early neurogenesis.

We have identified a new basic domain helix-loop-helix (bHLH) gene, NSCL-2, which was cloned because of its homology to the previously described putative hematopoietic transcription factor, SCL. NSCL-2 has been identified in both human and murine DNA. NSCL-2 complementary DNA clones were obtained from an 11.5-day murine embryo library. The coding region is 405 base pairs and encodes a predicted protein of 15.6 kilodaltons. There is 74% homology at the nucleotide level with the coding region of the murine SCL and 27% protein homology. Unlike the majority of previously described bHLH genes, the NSCL-2 coding region ends only six amino acids beyond the second amphipathic helix of the HLH domain. The NSCL-2 gene shows a markedly restricted pattern of expression predominantly confined to murine embryos at days 11-13 of development, although low level expression can be detected in murine embryos flanking this time point. Examination of 11- and 12-day mouse embryos by tissue in situ hybridization reveals expression of NSCL-2 in the developing nervous system, most likely in developing neurons. The NSCL-2 gene maps to murine chromosome 3. The temporally and tissue restricted pattern of expression of this gene and its identification as a member of a family of transcription factors relevant to growth and development in a wide variety of species suggest a role for NSCL-2 in the development of the eukaryotic nervous system.

Amino Acid Sequence↗

Assessment of asthma using automated and full-text medical records.

Automated medical records systems are used to study clinical outcomes and quality of care, but this requires accurate disease identification and assessment of severity. We sought to determine the reliability of identifying asthmatics through automated medical and pharmacy records, and the adequacy of such data for severity assessment. All adult health maintenance organization (HMO) members who received at least one asthma drug and an asthma diagnosis between April 1988 and September 1991 were identified. Records of a random sample were reviewed to validate the diagnosis and extract clinical information. Asthma drugs were dispensed to 15,491 individuals; 7583 (49%) also received an asthma diagnosis. Asthma drug use was three times greater for persons with diagnosed asthma compared to those with no diagnosis. Record review revealed that a coded asthma diagnosis had a positive predictive value of 86%. Nearly 4000 ambulatory encounters were reviewed, 10% of which were for asthma; the median number of encounters was two. Asthma symptoms were mentioned in 9% of all encounters; wheezing was most common. Peak flow and spirometry were measured in 4% and 1% of encounters, respectively. Records from recipients of asthma drugs who lacked an asthma diagnosis showed that 79% did not have asthma. Automated medical and pharmacy records from an HMO were relatively accurate when used to identify individuals with asthma. Similarly, most asthma drug recipients who lacked a coded diagnosis of asthma did not have asthma. However, conventional full-text records usually do not contain sufficient information to assess asthma severity, limiting the utility of such records for research and quality improvement.

Adult↗

Cloning and further sequence analysis of the spike gene of attenuated porcine epidemic diarrhea virus DR13.

The spike (S) gene of the attenuated porcine epidemic diarrhea virus (PEDV) DR13 was cloned and sequenced to further explore the functions of wild type PEDV and attenuated PEDV. Sequencing revealed a single large ORF of 4,149 nucleotides encoding a protein of 1,382 amino acids with predicted M(r) of 151 kDa. The coding region of the S gene of attenuated PEDV DR13 had 20 nucleotide changes that appeared to be significant determinants of function in that they produced changes in its predicted amino acid sequence. Notably, attenuated PEDV DR13 has previously been found to exhibit reduced pathogenicity in pigs. The regions containing these 20 nucleotide changes may therefore be crucial for PEDV pathogenicity. The attenuated PEDV DR13 S protein contains 28 Asn-Xaa-Ser/Thr sequons, 21 asparagines that are predicted to be N-glycosylated and a stretch of highly hydrophobic residues at positions 1,327-1,347, which is predicted to form an alpha-helix and to function as a membrane anchor. One (from N to K at 378) of the changes in the deduced amino acid sequence destroyed N-linked glycosylation sites, while another change (from N to S at 114) created a new one at a different location. These alterations in N-linked glycosylation sites reflected 3 nucleotide changes, which were related to the above-mentioned nucleotide changes and are suggested to influence the pathogenicity of attenuated PEDV DR13. Attenuated PEDV DR13 has 96.5, 96.4, 96.1, 93.9, 93.5 and 96.6% DNA sequence identities with CV777, Br1/87, JS-2004-2, Spk1, Chinju99 and parent DR13, respectively. Likewise, it shares 95.7, 95.4, 95.6, 92.0, 91.6 and 95.7% identity with those genes at the deduced amino acid sequence level. Phylogenetic analysis suggested that attenuated PEDV DR13 is closely related to CV777, Br1/87, JS-2004-2 and parent DR13, rather than to Spk1 and Chinju99 and is especially close to the Chinese PEDV strain JS-2004-2.

Amino Acid Sequence↗

Measurements of absolute CH concentrations by cavity ring-down spectroscopy and linear laser-induced fluorescence in laminar, counterflow partially premixed and nonpremixed flames at atmospheric pressure.

We report quantitative, spatially resolved measurements of methylidyne concentration ([CH]) in laminar, counterflow partially premixed and nonpremixed flames at atmospheric pressure by using both cavity ring-down spectroscopy (CRDS) and linear laser-induced fluorescence (LIF) in the A-X (0, 0) band. Three partially premixed (phiB = 1.45, 1.6, 2.0) flames plus a single nonpremixed methane-air flame are investigated at a global strain rate of 20 s(-1). These quantitative measurements are compared with predictions from an opposed-flow flame code when utilizing two GRI chemical kinetic mechanisms (versions 2.11 and 3.0). The LIF measurements of [CH] are corrected for variations in the electronic quenching rate coefficient by using predicted major species concentrations and temperatures along with quenching cross sections for CH that are available in the literature. The peak CH concentration obtained by CRDS is used to calibrate the quenching-corrected LIF measurements. Excellent agreement is obtained between CH concentration profiles measured by using the CRDS and LIF techniques. The spatial location of the CH layer is very well predicted by GRI 3.0; moreover, the measured and predicted CH concentrations are in good agreement for all the flames of this study.

Journal Article↗

A non-destructive method to determine the depth of radionuclides in materials in-situ.

A non-destructive method based on in-situ gamma spectroscopy is developed to determine the depth of radiological contamination in media. An innovative algorithm, Gamma Penetration Depth Unfolding Algorithm (GPDUA), uses point kernel techniques to predict the depth of contamination based on the results of the uncollided peak information from the in-situ gamma spectroscopy. The GPDUA is designed and verified through extensive Monte Carlo simulations and validated through laboratory experiments. This innovative tool promises to be "better, faster, safer, and cheaper" than the current practice in decontamination and decommissioning. The method requires the a priori knowledge of the contaminant source distribution. The applicable radiological contaminants of interest are any isotopes that emit two or more gamma rays per disintegration or isotopes that emit a single gamma ray but have gamma-emitting progeny in secular equilibrium with its parent (e.g., 60Co, 235U, and 137Cs to name a few). The predicted depths from the GPDUA algorithm using Monte Carlo N-Particle Transport Code simulations and laboratory experiments using 60Co have consistently produced predicted depths within 20% of the actual or known depth.

Gamma Rays↗

Novel lipoprotein expressed by Neisseria meningitidis but not by Neisseria gonorrhoeae.

The ppk gene, which codes for the enzyme polyphosphate kinase in Neisseria meningitidis strain BNCV, is preceded by an open reading frame coding for a protein with a predicted size of 19.2 kDa with a typical lipoprotein signal sequence of 21 amino acids. The protein has significant homology to the N-terminal portion of an outer membrane protein from Haemophilus somnus (J. Won and R. W. Griffith, Infect. Immun. 61:2813-2821, 1993). Sequencing of the same open reading frame from meningococcus strain M1080 predicted an almost identical protein. Antisera were raised against the lipoprotein, expressed in Escherichia coli as a fusion protein with glutathione S-transferase. The antisera reacted with meningococcal membrane fractions on a Western blot (immunoblot) but did not elicit complement-dependent bactericidal activity. Restriction enzyme digestion demonstrated conservation of this portion of the meningococcal and gonococcal chromosomes. However, antisera raised to the recombinant protein showed that the protein was absent from all strains of gonococcus tested. The sequences of the gene from several strains of Neisseria gonorrhoeae and N. meningitidis were compared and found to be almost identical, except that the coding sequences from all of the gonococcal strains were terminated prematurely as a result of a frameshift mutation. The significance of the remarkable conservation of these gonococcal genes is discussed.

Amino Acid Sequence↗

Challenges in validating CFD-derived inhaled aerosol deposition predictions.

Computational fluid dynamic (CFD) techniques have provided unprecedented opportunity for investigating inhaled particle deposition in realistic human airway geometries. Several recent articles describing local aerosol deposition predictions based upon "validated" CFD models have highlighted the challenges in validating local aerosol deposition predictions. These challenges include: (1) defining what is meant by validation; (2) defining appropriate experimental data for validation; and (3) determining when the agreement is not fortuitous. The term validation has numerous meanings, depending on the field and context in which it is used. For example, in computer programming it means the code executes as intended, to the experimentalist it means predicted results agree with matched experimental measurements, and to the risk assessor it implies that predictions using new parameters can be trusted. Based on the current literature it is not clear that a consensus exists for what constitutes a validated CFD model. It is also not clear what types of experimental data are needed or how closely the CFD input values and experimental conditions should be matched (similar or identical airway geometries, entrance airflow, or aerosol profiles) to validate CFD derived predictions. Due to the complexity of CFD computer codes and the multiplicity of deposition mechanisms, it is possible that total aerosol deposition may be accurately predicted and the resulting local particle deposition patterns are incorrect, or vice versa. Specific examples and suggestions for several challenges to experimentalists and modelers are presented.

Aerosols↗

Complete nucleotide sequence of the freshwater unicellular cyanobacterium Synechococcus elongatus PCC 6301 chromosome: gene content and organization.

The entire genome of the unicellular cyanobacterium Synechococcus elongatus PCC 6301 (formerly Anacystis nidulans Berkeley strain 6301) was sequenced. The genome consisted of a circular chromosome 2,696,255 bp long. A total of 2,525 potential protein-coding genes, two sets of rRNA genes, 45 tRNA genes representing 42 tRNA species, and several genes for small stable RNAs were assigned to the chromosome by similarity searches and computer predictions. The translated products of 56% of the potential protein-coding genes showed sequence similarities to experimentally identified and predicted proteins of known function, and the products of 35% of the genes showed sequence similarities to the translated products of hypothetical genes. The remaining 9% of genes lacked significant similarities to genes for predicted proteins in the public DNA databases. Some 139 genes coding for photosynthesis-related components were identified. Thirty-seven genes for two-component signal transduction systems were also identified. This is the smallest number of such genes identified in cyanobacteria, except for marine cyanobacteria, suggesting that only simple signal transduction systems are found in this strain. The gene arrangement and nucleotide sequence of Synechococcus elongatus PCC 6301 were nearly identical to those of a closely related strain Synechococcus elongatus PCC 7942, except for the presence of a 188.6 kb inversion. The sequences as well as the gene information shown in this paper are available in the Web database, CYORF (http://www.cyano.genome.jp/).

Base Sequence↗

Organization and expression of the double-stranded RNA genome of Helminthosporium victoriae 190S virus, a totivirus infecting a plant pathogenic filamentous fungus.

The complete nucleotide sequence, 5178 bp, of the totivirus Helminthosporium vicotoriae 190S virus (Hv190SV) double-stranded RNA, was determined. Computer-assisted sequence analysis revealed the presence of two large overlapping ORFs; the 5'-proximal large ORF (ORF1) codes for the coat protein (CP) with a predicted molecular mass of 81 kDa, and the 3'-proximal ORF (ORF2), which is in the -1 frame relative to ORF1, codes for an RNA-dependent RNA polymerase (RDRP). Unlike many other totiviruses, the overlap region between ORF1 and ORF2 lacks known structural information required for translational frameshifting. Using an antiserum to a C-terminal fragment of the RDRP, the product of ORF2 was identified as a minor virion-associated polypeptide of estimated molecular mass of 92 kDa. No CP-RDRP fusion protein with calculated molecular mass of 165 kDa was detected. The predicted start codon of the RDRP ORF (2605-AUG-2607) overlaps with the stop codon (2606-UGA-2608) of the CP ORF, suggesting RDRP is expressed by an internal initiation mechanism. Hv190SV is associated with a debilitating disease of its phytopathogenic fungal host. Knowledge of its genome organization and expression will be valuable for understanding its role in pathogenesis and for potential exploitation in the development of biocontrol measures.

Amino Acid Sequence↗

Administrative data accurately identified intensive care unit admissions in Ontario.

BACKGROUND AND OBJECTIVES: To evaluate the accuracy of Ontario administrative health data for identifying intensive care unit (ICU) patients. MATERIALS AND METHODS: Records from the Critical Care Research Network patient registry (CCR-Net) were linked to the Ontario Health Insurance Program (OHIP) database and the Canadian Institute for Health Information (CIHI) database. The CCR-Net was considered the criterion standard for assessing the accuracy of different OHIP or CIHI codes for identifying ICU admission. RESULTS: The highest positive predictive value (PPV) for ICU admission (91%) was obtained using a CIHI special care unit (SCU) code, but its sensitivity was poor (26%). A strategy based on a combination of CIHI SCU codes yielded a lower PPV (84%) but a higher sensitivity (92%). A strategy based purely on OHIP claims yielded further reductions in PPV (73%), gains in specificity (99%), and moderate sensitivity (56%). The highest sensitivity (100%) was obtained using a combination of CIHI and OHIP codes in exchange for poor PPV (32%). CONCLUSIONS: Administrative databases can be used to identify ICU patients, but no single strategy simultaneously provided high sensitivity, specificity, and PPV. Researchers should consider the study purpose when selecting a strategy for health services research on ICU patients.

Databases as Topic↗

The primary structure of the 32-kDa subunit of human replication protein A.

Replication protein A (RP-A) is a complex of three polypeptides of molecular mass 70, 32, and 14 kDa, which is absolutely required for simian virus 40 DNA replication in vitro. We have isolated a cDNA coding for the 32-kDa subunit of RP-A. An oligonucleotide probe was constructed based upon a tryptic peptide sequence derived from whole RP-A, and clones were isolated from a lambda gt11 library containing HeLa cDNA inserts. The amino acid sequence predicted from the cDNA contains the peptide sequence obtained from whole RP-A along with two sequences obtained from tryptic peptides derived from sodium dodecyl sulfate-polyacrylamide gel-purified 32-kDa subunit. The coding sequence predicts a protein of 29,228 daltons, in good agreement with the electrophoretically determined molecular mass of the 32-kDa subunit. No significant homology was found with any of the sequences in the GenBank data base. The protein predicted from the cDNA has an N-terminal region rich in glycine and serine along with two acidic and two basic segments. Monoclonal antibodies have been raised against the 70- and 32-kDa subunits of RP-A. The cloned cDNA has been overexpressed in bacteria using an inducible T7 expression system. The protein made in bacteria is recognized by a monoclonal antibody that is specific for the 32-kDa subunit of RP-A. This monoclonal antibody against the 32-kDa subunit inhibits DNA replication in vitro.

Amino Acid Sequence↗

The role of combined allelic imbalance and mutations of p53 in tumor progression and survival following surgery for colorectal carcinoma.

The role of p53 mutations in disease progression and survival of colorectal cancer is unclear, since numerous studies have reported different conclusions. However, few reports, if any, have evaluated disease progression and survival in relationship to 'functional' and 'non-functional' p53 status defined by genetic and molecular indications. Malignant colorectal tumors, from 72 unselected patients who underwent primary and potentially curative elective tumor resections, were either classified as p53 functional (p53+/+, p53+/-) or non-functional (p53-/-) based on DNA sequence analysis of all p53 exons, including determination of allelic imbalance of p53 (LOH), according to four DNA markers; 2 within the coding gene and two markers in the immediate flanking regions of p53. Tumor frequency of microsatellite instability was also analyzed according to Dukes' A-D stages. Dukes' staging predicted survival as expected, while the conceptual p53 status, functional p53 vs non-functional p53, did not clear-cut predict disease specific survival. p53 mutations alone or allelic imbalance inside the reading frame of the gene were unpredictive of survival, while allelic imbalance downstream of p53 predicted reduced survival (p < 0.05). The present study demonstrates that base mutations in combination with allelic imbalance within the reading frame of p53 do not predict survival or progression of colorectal cancer, while allelic imbalance upstream coding parts of the gene predicted disease-specific survival in univariate analysis. Thus, structural alterations within the gene seem less important than alterations in regions with potential control elements.

Adult↗

The accuracy of predicting cardiac arrest by emergency medical services dispatchers: the calling party effect.

OBJECTIVES: To analyze the accuracy of paramedic emergency medical services (EMS) dispatchers in predicting cardiac arrest and to assess the effect of the caller party on dispatcher accuracy in an advanced life support, public utility model EMS system, with greater than 90,000 calls and greater than 60,000 transports per year. METHODS: This was a retrospective analysis from January 1, 2000, through June 30, 2000, of 911 calls with dispatcher-assigned presumptive patient condition (PPC) or field diagnosis of cardiac arrest. Sensitivity and positive predictive value (PPV) of the PPC code for cardiac arrest by calling parties were calculated. Homogeneity of sensitivity and PPV of the PPC code for cardiac arrest by calling parties was studied with chi-square analysis. Relevant proportions, relative risk ratios, and associated 95% confidence intervals (95% CIs) were calculated. Student's t-test was used to compare quality assurance scores between calling parties. RESULTS: There were 506 patients included in the study. Overall sensitivity for dispatcher-assigned PPC of cardiac arrest was 68.3% (95% CI = 63.3% to 73.0%) with a PPV of 65.0% (95% CI = 60.0% to 69.7%). There was a significant difference in the PPV for the EMS dispatcher diagnosis of cardiac arrest depending on the type of caller (chi(2) = 17.34, p < 0.001). CONCLUSIONS: A higher level of medical training may improve dispatch accuracy for predicting cardiac arrest. The type of calling party influenced the PPV of dispatcher-assigned condition.

Emergency Medical Services↗

Positively selected amino acid sites in the entire coding region of hepatitis C virus subtype 1b.

To predict the amino acid sites important for the clearance of hepatitis C virus (HCV) subtype 1b in vivo, positively selected amino acid sites were detected by analyzing the sequence data collected from the international DNA databank. The rate of nonsynonymous substitutions per nonsynonymous site was compared with that of synonymous substitutions per synonymous site for each codon site in the entire coding region. As a result, 13 out of 3010 amino acid sites were found to be positively selected. Among the 13 positively selected amino acid sites, eight were located in the structural proteins and five were in the nonstructural proteins. Moreover, eight were located in B-cell epitopes and two were in T-cell epitopes. These observations suggest that both the antibody and the cytotoxic T lymphocyte are involved in the clearance of HCV subtype 1b in vivo. These positively selected amino acid sites represent candidate vaccination targets for HCV subtype 1b.

Amino Acids↗

Computer code for the optimization of performance parameters of mixed explosive formulations.

LOTUSES is a novel computer code, which has been developed for the prediction of various thermodynamic properties such as heat of formation, heat of explosion, volume of explosion gaseous products and other related performance parameters. In this paper, we report LOTUSES (Version 1.4) code which has been utilized for the optimization of various high explosives in different combinations to obtain maximum possible velocity of detonation. LOTUSES (Version 1.4) code will vary the composition of mixed explosives automatically in the range of 1-100% and computes the oxygen balance as well as the velocity of detonation for various compositions in preset steps. Further, the code suggests the compositions for which least oxygen balance and the higher velocity of detonation could be achieved. Presently, the code can be applied for two component explosive compositions. The code has been validated with well-known explosives like, TNT, HNS, HNF, TATB, RDX, HMX, AN, DNA, CL-20 and TNAZ in different combinations. The new algorithm incorporated in LOTUSES (Version 1.4) enhances the efficiency and makes it a more powerful tool for the scientists/researches working in the field of high energy materials/hazardous materials.

Algorithms↗