Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

The phylogeny of tRNAs seems to confirm the predictions of the coevolution theory of the origin of the genetic code.

An extensive analysis of the evolutionary relationships existing between transfer RNAs, performed using parsimony algorithms, is presented. After building up an estimate of the tRNA ancestral sequences, these sequences are then compared using certain methods. The results seem to suggest that the coevolution hypothesis (Wong, J.T., 1975, Proc. Natl. Acad. Sci. USA 72, 1909-1912) that sees the genetic code as a map of the biosynthetic relationships between amino acids is further supported by these results, as compared to the hypotheses that see the physicochemical properties of amino acids as the main adaptative theme that led to the structuring of the genetic code.

Amino Acid Sequence↗

Complex promoter and coding region beta 2-adrenergic receptor haplotypes alter receptor expression and predict in vivo responsiveness.

The human beta(2)-adrenergic receptor gene has multiple single-nucleotide polymorphisms (SNPs), but the relevance of chromosomally phased SNPs (haplotypes) is not known. The phylogeny and the in vitro and in vivo consequences of variations in the 5' upstream and ORF were delineated in a multiethnic reference population and an asthmatic cohort. Thirteen SNPs were found organized into 12 haplotypes out of the theoretically possible 8,192 combinations. Deep divergence in the distribution of some haplotypes was noted in Caucasian, African-American, Asian, and Hispanic-Latino ethnic groups with >20-fold differences among the frequencies of the four major haplotypes. The relevance of the five most common beta(2)-adrenergic receptor haplotype pairs was determined in vivo by assessing the bronchodilator response to beta agonist in asthmatics. Mean responses by haplotype pair varied by >2-fold, and response was significantly related to the haplotype pair (P = 0.007) but not to individual SNPs. Expression vectors representing two of the haplotypes differing at eight of the SNP loci and associated with divergent in vivo responsiveness to agonist were used to transfect HEK293 cells. beta(2)-adrenergic receptor mRNA levels and receptor density in cells transfected with the haplotype associated with the greater physiologic response were approximately 50% greater than those transfected with the lower response haplotype. The results indicate that the unique interactions of multiple SNPs within a haplotype ultimately can affect biologic and therapeutic phenotype and that individual SNPs may have poor predictive power as pharmacogenetic loci.

Base Sequence↗

Isolation and characterization of two distinct growth hormone cDNAs from the tetraploid smallmouth buffalofish (Ictiobus bubalus).

The growth hormone (GH) gene has been characterized for a number of fishes and used to establish phylogenetic relationships and population structures. Analysis of tetraploid fishes, such as salmon and some Asian cyprinids, has shown the presence of two GH genes. Fishes in the sucker family (Catostomidae, Cypriniformes) are also tetraploid, and the present study reports the isolation and characterization of two GH cDNAs from a representative species, the smallmouth buffalofish (Ictiobus bubalus). The GH cDNAs of smallmouth buffalofish are 1272 and 1273nt in length, and each codes for a polypeptide of 210 amino acids, predicted to be cleaved to a final product of 188 aa. The GH cDNAs of smallmouth buffalofish are 6% divergent in nt sequence in the coding region, and there are 16 differences in predicted aa sequence. Because the cDNAs have distinct sequences in coding regions and in UTRs, which differed by more than 10%, they were identified as GHI and GHII. The predicted GHI protein contains 4 Cys residues, homologous to other vertebrate GH sequences. On the other hand, GHII has 5 Cys residues, homologous to other ostariophysan sequences. GHI and GHII are most similar to other cypriniform fishes for both nt and protein sequences. Phylogenetically, the sequences of smallmouth buffalofish GH consistently grouped with Asian cyprinids, but not loaches, consistent with morphological evidence suggesting that suckers are most closely related to minnows.

Amino Acid Sequence↗

Modal decomposition method for modeling the interaction of Lamb waves with cracks.

The interaction of the low-order antisymmetric (a0) and symmetric (s0) Lamb waves with vertical cracks in aluminum plates is studied. Two types of slots are considered: (a) internal crack symmetrical with respect to the middle plane of the plate and (b) opening crack. The modal decomposition method is used to predict the reflection and transmission coefficients and also the through-thickness displacement fields on both sides of slots of various heights. The model assumes strip plates and cracks, thus considering two-dimensional plane strain conditions. However, mode conversion (a0 into s0 and vice versa) that occurs for single opening cracks is considered. The energy balance is always calculated from the reflection and transmission coefficients, in order to check the validity of the results. These coefficients together with the through-thickness displacement fields are also compared to those predicted using a finite element code widely used in the past for modeling Lamb mode diffraction problems. Experiments are also made for measuring the reflection and transmission coefficients for incident a0 or s0 lamb modes on opening cracks, and compared to the numerical predictions.

Journal Article↗

Statistical evaluation of the coding capacity of complementary DNA strands.

Two independent methods are used to evaluate the protein-coding information content in different classes of DNA sequences. The first method allows to evaluate the statistical relevance of finding unidentified reading frames, longer than 100 codons, on both DNA strands of: a) 117 DNA sequences that code for 142 nuclear proteins; b) 39 stable RNA coding sequences and c) 36 other DNA sequences which include regulatory and as yet unknown function sequences. The finding of 50 reading frames longer than 100 codons (complementary inverted proteins or c.i.p. genes) located on the DNA strand complementary to the protein-coding one is drastically in excess of the number predicted by chance alone. An independent method (testcode) applied to c.i.p. gene sequences, which assigns the probability of coding to a given sequence, predicts that more than 50% of these genes are translated in a functional product. These analyses indicate the existence of a new class of protein-coding genes, located on the DNA sequences complementary to the protein-coding DNA strand.

Animals↗

Diagnostic coding and medical rehabilitation length of stay: their relationship.

OBJECTIVE: To determine if diagnostic information provided in the form of International Classification of Diseases, 9th Revision, Clinical Modification (ICD-9-CM) codes improves rehabilitation length of stay (LOS) prediction when used in combination with the Functional Independence Measure-Function Related Groups (FIM-FRGs) classification system. DESIGN: Various models characterizing diagnostic information using ICD-9-CM codes were created that included individual ICD-9-CM codes and groupings of those codes by organ or etiology involved. Each method was evaluated using linear regression with the natural logarithm of LOS as the dependent variable. Separate validation data sets were held back to quantify the incremental effect of diagnosis when combined with the FIM-FRG classification system. SETTING: Records from 252 rehabilitation facilities and hospital units across the nation. PATIENTS: Analyses were undertaken using 82,646 records from patients discharged in 1992. RESULTS: The addition of ICD-9-CM diagnostic information to the FIM-FRG classification system increased the variance explained by a maximum of 1.9%, from 31.5% to 33.4%. CONCLUSIONS: Refinement of the FIM-FRGs to include ICD-9-CM diagnoses does not appear warranted on the basis of the small increase in the percentage of explained variance in LOS. We believe the lack of improved prediction with the addition of ICD-9-CM codes relates primarily to incomplete coding practices and to the effect of patients' diagnoses being absorbed in variables as already expressed by the FIM-FRG system. Although ICD-9-CM codes, overall, did not greatly improve LOS prediction, they appeared to have some impact in certain impairment categories.

Brain Injuries↗

Prediction of exact boundaries of exons.

It is known that while the programs used to predict genes are good at determining coding nucleotides, there are considerable inaccuracies in the determination of the gene structural elements. Among them, the most notable is that of the exact boundaries of exons. In order to assess this, we had earlier reviewed various programs that predict potential splice sites and exons. The results led to the following two observations: (i) a high proportion of false positive splice sites from computational predictions occur in the vicinity of real splice sites; and (ii) current algorithms are misled to predict wrong splice sites more often when the coding potential ends within +/-25 nucleotides from real sites than when it ends at farther positions. In this report, we review decision tree models for human splice sites and the resultant software tool, namely SpliceProximalCheck, that discriminates such'proximal' false positives from real splice sites. Further presented is an integrated system (MZEF-SPC) with Splice ProximalCheck (SPC) as a front-end tool operating on the results of Michael Zhang's exon finder program. Examination of the output of the integrated program on an illustrative gene set revealed that as much as 61 of 93 MZEF-predicted false positive exons could be eliminated by SPC for a loss of only 3 out of 33 MZEF-predicted true positive exons.

Algorithms↗

Separation of sequences from host-pathogen interface using triplet nucleotide frequencies.

The identification of genes involved in host-pathogen interactions is important for the elucidation of mechanisms of disease resistance and host susceptibility. A traditional way to classify the origin of genes sampled from a pool of mixed cDNA is through sequence similarity to known genes from either the pathogen or host organism or other closely related species. This approach does not work when the identified sequence has no close homologues in the sequence databases. In our previous studies, we classified genes using their codon frequencies. This method, however, explicitly required the prediction of CDS regions and thus could not be applied to sequences composed from the non-coding regions of genes. In this study, we show that the use of sliding-window triplet frequencies extends the application of the algorithm to both coding and non-coding sequences and also increases the prediction accuracy of a Support Vector Machine classifier from 95.6+/-0.3 to 96.5+/-0.2. Thus the use of the triplet frequencies increased the prediction accuracy of the new method by more than 20% compared to our previous approach. A functional analysis of sequences detected gene families having significantly higher or lower probability to be correctly classified compared to the average accuracy of the method is described. The server to perform classification of EST sequences using triplet frequencies is available at (URL: http://mips.gsf.de/proj/est3).

Algorithms↗

Finding genes in DNA with a Hidden Markov Model.

This study describes a new Hidden Markov Model (HMM) system for segmenting uncharacterized genomic DNA sequences into exons, introns, and intergenic regions. Separate HMM modules were designed and trained for specific regions of DNA: exons, introns, intergenic regions, and splice sites. The models were then tied together to form a biologically feasible topology. The integrated HMM was trained further on a set of eukaryotic DNA sequences and tested by using it to segment a separate set of sequences. The resulting HMM system which is called VEIL (Viterbi Exon-Intron Locator), obtains an overall accuracy on test data of 92% of total bases correctly labelled, with a correlation coefficient of 0.73. Using the more stringent test of exact exon prediction, VEIL correctly located both ends of 53% of the coding exons, and 49% of the exons it predicts are exactly correct. These results compare favorably to the best previous results for gene structure prediction and demonstrate the benefits of using HMMs for this problem.

Algorithms↗

A gene-based model of fitness and its implications for genetic variation: Linkage disequilibrium.

A widely used model of the effects of mutations on fitness (the "sites" model) assumes that heterozygous recessive or partially recessive deleterious mutations at different sites in a gene complement each other, similarly to mutations in different genes. However, the general lack of complementation between major effect allelic mutations suggests an alternative possibility, which we term the "gene" model. This assumes that a pair of heterozygous deleterious mutations in trans behave effectively as homozygotes, so that the fitnesses of trans heterozygotes are lower than those of cis heterozygotes. We examine the properties of the two different models, using both analytical and simulation methods. We show that the gene model predicts positive linkage disequilibrium (LD) between deleterious variants within the coding sequence, under conditions when the sites model predicts zero or slightly negative LD. We also show that focussing on rare variants when examining patterns of LD, especially with Lewontin's´ measure, is likely to produce misleading results with respect to inferences concerning the causes of the sign of LD. Synergistic epistasis between pairs of mutations was also modeled; it is less likely to produce negative LD under the gene model than the sites model. The theoretical results are discussed in relation to patterns of LD in natural populations of several species.

complementation↗

Patient and disease profile of emergency medical readmissions to an Irish teaching hospital.

OBJECTIVE: To determine whether there was a relationship between coded diseases at the time of hospital discharge, a pattern of ordering investigations, and hospital readmission in a major teaching hospital. DESIGN: Systematic review of data relating to emergency medical patients admitted to St James' Hospital Dublin between 1 January and 31 December 2002. DATA SOURCES AND METHODS: Data on discharges from hospital recorded in the Hospital In-Patient Enquiry (HIPE) system. The value of HIPE data in describing the relationship between the pattern of resource utilisation, diagnostic related groups, and hospital readmission has not previously been examined. RESULTS: Of 5038 episodes recorded among 4050 patients admitted, the number of readmissions was up to 15. Age and male gender were factors associated with readmission, and readmitted patients remained in hospital for longer. No particular test request predicted readmission, but computed tomography of the brain was associated with a reduced readmission rate. Discharge diagnostic related group coding at first discharge predicted readmission-codes related to heart failure, respiratory system, alcohol, malignancy, and anaemia. CONCLUSIONS: It was found that clinical coding using the HIPE database strongly predicted hospital readmission. It may be argued that early hospital readmission reflects unsatisfactory patient care, alternatively that many readmissions are not preventable, representing either new events in elderly patients with chronic illnesses and frequent co-morbidity or related to social factors. The utility of specific interventions, in patients at high risk for hospital readmission, could be explored.

Adult↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Utility of Scottish morbidity and mortality data for epidemiological studies of motor neuron disease.

OBJECTIVES: To determine the accuracy of (1) hospital discharge data and (2) death certificates, coded as motor neuron disease (MND). DESIGN: Comparison of data from The Scottish Motor Neuron Disease Register (SMNDR) with routinely collected Scottish Hospital In-Patient Statistics (SHIPS) and death certificate coding. SETTING: Scotland UK. PATIENTS: 1) 379 adults (> 15 years) discharged for the first time from a Scottish hospital in 1989-90 and (2) 281 deaths in the same period assigned to the International Classification of Diseases (ICD)-9, category 335 (MND). MAIN OUTCOME MEASURES: The sensitivity and positive predictive value of a diagnosis of MND as retrieved by (1) the Information and Statistics Division of the Common Services Agency for the Scottish Health Service for morbidity data and (2) the Registrar General's office for mortality data, using the SMNDR as the 'gold standard'. RESULTS: (1) Thirty per cent of adult patients identified as having MND by SHIPS did not have this disease and 23% of patients with MND did not appear on SHIPS. The sensitivity of a diagnosis of MND, as retrieved by SHIPS, was 84% and the positive predictive value was 70% overall. Miscoding of patients with pseudobulbar palsy caused by cerebrovascular disease was the major source of false positive error. The incidence of adult onset sporadic MND was over estimated by SHIPS by a factor of 1.6. (2) Mortality data were more accurate, with a false negative rate of 6% and a positive predictive value of 90%. CONCLUSIONS: Coded hospital discharge data are an inaccurate record of a diagnosis of MND and cannot, in their present form, be used as a reliable measure of disease incidence in Scotland. Greater care is required in the preparation of discharge summaries and coding if these data are to be useful for health care planning and epidemiological research. SHIPS is, however, an important source of information to achieve a complete sample of patients with MND. There is also a problematic false positive rate for mortality data but this source more closely approximates true incidence.

Adolescent↗

Predicting recovery from aphasia with connectionist networks: preliminary comparisons with multiple regression.

We trained a series of simulated neural networks with the raw scores on the Western Aphasia Battery from 91 aphasic patients. Patients were tested at 3 and at 12 months post onset. The most successful network we trained is able to predict AQ for an individual in 12 months from the raw scores at 3 months post-onset to a tolerance of + or -4.5. We then compared the relative success of a small range of trained networks to predict recovery with linear multiple regression. With the small groups of subjects involved in this preliminary study, the networks appeared to be more successful at predicting recovery.

Aged↗

Structure of a genome region of the Lactobacillus gasseri temperate phage phiadh covering a repressor gene and cognate promoters.

By sequencing the DNA regions which flank the intG gene encoding integrase of the temperate Lactobacillus (Lb.) gasseri bacteriophage phiadh, a continuous sequence of 6590 bp was established. It encompasses five newly identified ORFs, of which four are located upstream, and one (orfC) downstream of intG. Proteins corresponding to the expected products of the intG upstream coding regions, orfA (33 kDa), orf2 (14 kDa), rad (12.1 kDa), and tec (7.9 kDa), were identified by in vitro expression of subcloned DNA fragments. Rad shares homology with transcription regulators, including SinR of Bacillus species and the repressor of phage phi105. The gporf2 is similar to predicted products of topologically equivalent coding regions of the Lactococcus lactis phage TP901-1 and the B. subtilis phage phi105. Promoters for the divergently oriented rad and tec genes were mapped within the 435-bp region between them and specify overlapping transcripts with extended 5'-untranslated sequences. As shown with lacZ fusions, Rad repressed transcription from the tec and rad promoters 20- and 5-fold, respectively. In Lb. gasseri, weak expression of cloned rad ws sufficient to mediate immunity towards phiadh.

Amino Acid Sequence↗

UFold-X: an enhanced Dual & Dynamic U-Mamba model for long-range RNA secondary structure prediction.

RNA secondary structure is essential for understanding the functions of non-coding RNAs, ribosomal RNAs, and viral genomes. However, accurate prediction of long RNA structures remains challenging due to complex long-range interactions and the limited availability of long-RNA training data. We present UFold-X, a dual-branch deep learning framework that combines a convolutional encoder for local structure modeling with a Mamba-based Visual State Space Module for capturing long-range dependencies. A dynamic gating mechanism adaptively integrates the two branches according to sequence length. UFold-X was evaluated on multiple benchmark datasets containing RNAs up to 5000 nucleotides. To rigorously assess generalization, we introduced a cross-clan benchmark for long RNAs. Under this stringent setting, UFold-X achieved performance comparable to state-of-the-art classical approaches while achieving the best performance among deep learning-based methods. Additional cross-family and within-family evaluations further demonstrated robust transferability and competitive predictive performance. UFold-X also maintained excellent computational efficiency, requiring only 0.08 s per sequence on average. To assess biological consistency, we developed a SHAPE-based reactivity prediction variant (UFold-X-R) and an integrated metric, the Hybrid Reactivity-Pairing Score (HRPS). UFold-X-R showed strong agreement with experimental icSHAPE data and achieved the highest HRPS among all evaluated methods. A user-friendly web server is available at https://ufold-x.ai4bread.com.

Nucleic Acid Conformation↗

Molecular evolution and phylogenetic application of DMC1.

The protein encoded by the single-copy nuclear gene DMC1 belongs to the recA-like group of proteins involved in meiosis. Partial nucleotide sequence, spanning exon 10 to exon 15, was used to test the applicability of the gene to phylogenetic studies in higher plants and used to assess its molecular evolution. The sequences produced from the Triticeae (Poaceae) show that most of the variation is confined to the introns. If a wider taxon sampling is used, alignment problems may be predicted. Comparisons including four complete coding sequences from GenBank reveal that the exons are more than twice as variable as rbcL, but easy to align, and hence may be valuable at higher taxonomic levels. Substitution rates are variable within the Triticeae, though local subclades show rate constancy. The relationships between exon variation and predicted protein structure are briefly discussed. In general, none of the observed nucleotide substitutions can be predicted to cause major structural or functional changes.

Amino Acid Sequence↗

The validity of hospital discharge register data on coronary heart disease in Finland.

We studied the validity of the Finnish hospital discharge register data on coronary heart disease (CHD) for the purposes of epidemiologic studies and health services research. The Finnish nationwide hospital discharge register (HDR) was linked with the FINMONICA acute myocardial infarction (AMI) register for the years 1983-1990. The frequency of errors in the HDR was assessed separately. Between 8% and 13% of hospitalized AMI events registered in the AMI Register were not found in the HDR with an ICD code for CHD. Problems with the register linkage and the use of some ICD code other than one of the codes for CHD explained these missing events. The frequency of errors in the personal identification number was about 5% in the early 1980s. After 1986 errors were found only occasionally. The diagnosis recorded in the HDR was the same as that in the discharge sheet in about 95% of hospitalizations. The positive predictive value of the ICD code 410 (AMI), compared with the FINMONICA definite+possible AMI category, was very high and stable, about 90% in all areas and all hospitals, but the sensitivity varied from 50% at local hospitals to 80% at central hospitals. In summary, data on CHD obtained from the Finnish hospital discharge register give, on average, a correct picture on changes in the occurrence of AMI in Finland and can, with necessary caution, be used in epidemiological studies and health services research. However, the classification of individual cases is not standardized in the HDR, but varies over time, between geographical areas and the levels of care. Therefore, these data should not be used without confirmation in studies where correct classification of individual outcomes is of crucial importance, such as follow-up studies and case-control studies.

Adult↗