Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Cloning, sequencing, and expression of the Lactobacillus casei thymidylate synthase gene.

The thymidylate synthase (TS) gene from Lactobacillus casei has been isolated, cloned, and sequenced. The coding sequence is 948 bp and predicts a primary structure, identical to that reported by protein sequencing methods (Maley et al., 1979b). The gene has been placed in several expression systems which complement TS-deficient Escherichia coli, and express the catalytically active enzyme at levels of 10-20% of the soluble protein of E. coli. The expressed TS has kinetic and structural properties consistent with its being identical to the authentic enzyme from L. casei.

Amino Acid Sequence↗

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1 = 0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n = 305) demonstrated generalizability (AUROC = 0.8105; F1 = 0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC) = 0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article↗

Accelerating inference in genomic and proteomic foundation models via speculative decoding.

MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.

Genomics↗

Generating consensus sequences from partial order multiple sequence alignment graphs.

MOTIVATION: Consensus sequence generation is important in many kinds of sequence analysis ranging from sequence assembly to profile-based iterative search methods. However, how can a consensus be constructed when its inherent assumption-that the aligned sequences form a single linear consensus-is not true? RESULTS: Partial Order Alignment (POA) enables construction and analysis of multiple sequence alignments as directed acyclic graphs containing complex branching structure. Here we present a dynamic programming algorithm (heaviest_bundle) for generating multiple consensus sequences from such complex alignments. The number and relationships of these consensus sequences reveals the degree of structural complexity of the source alignment. This is a powerful and general approach for analyzing and visualizing complex alignment structures, and can be applied to any alignment. We illustrate its value for analyzing expressed sequence alignments to detect alternative splicing, reconstruct full length mRNA isoform sequences from EST fragments, and separate paralog mixtures that can cause incorrect SNP predictions. AVAILABILITY: The heaviest_bundle source code is available at http://www.bioinformatics.ucla.edu/poa

Algorithms↗

Identification and analysis of a class 2 alpha-mannosidase from Aspergillus nidulans.

A Class 2 alpha-mannosidase gene was cloned and sequenced from the filamentous fungus Aspergillus nidulans. A portion of the gene was amplified using degenerate oligonucleotide primers which were designed based on similarity between the Saccharomyces cerevisiae vacuolar and rat ER/cytosolic Class 2 protein sequences. The PCR amplification product was used to isolate the full length gene, and DNA sequencing revealed a 3383 bp coding region containing three introns. The predicted 1049 amino acid reading frame contained six potential N-glycosylation sites and encoded a protein of 118 kDa. The protein sequence did not appear to encode a typical fungal signal sequence or membrane spanning domain. Although the cellular location of the A.nidulans mannosidase was not determined, experimental evidence suggested that it was located within a subcellular organelle. The Matchbox sequence similarity matrix indicated that the A.nidulans protein sequence was more highly similar to the rat ER/cytosolic (Rij = 0.33) and S.cerevisiae vacuolar alpha-mannosidases (Rij = 0.43) than the rat and yeast sequences were to each other (Rij = 0.29). These three enzymes were found to be distantly related to other Class 2 sequences, and compose a third subgroup of Class 2 alpha-mannosidases, as shown by ClustalW sequence alignment.

Amino Acid Sequence↗

Positional dissociation between the genetic mutation responsible for pseudohypoparathyroidism type Ib and the associated methylation defect at exon A/B: evidence for a long-range regulatory element within the imprinted GNAS1 locus.

Pseudohypoparathyroidism type Ib (PHP-Ib) is a paternally imprinted disorder which maps to a region on chromosome 20q13.3 that comprises GNAS1 at its telomeric boundary. Exon A/B of this gene was recently shown to display a loss of methylation in several PHP-Ib patients. In nine unrelated PHP-Ib kindreds, in whom haplotype analysis and mode of inheritance provided no evidence against linkage to this chromosomal region, we confirmed lack of exon A/B methylation for affected individuals, while unaffected carriers showed no epigenetic abnormality at this locus. However, affected individuals in one kindred (Y2) displayed additional methylation defects involving exons NESP55, AS and XL, and unaffected carriers in this family showed an abnormal methylation at exon NESP55, but not at other exons. Taken together, current evidence thus suggests that distinct mutations within or close to GNAS1 can lead to PHP-Ib and the associated epigenetic changes. To further delineate the telomeric boundary of the PHP-Ib locus, the previously reported kindred F, in which patient F-V/51 is recombinant within GNAS1, was investigated with several new markers and direct nucleotide sequence analysis. These studies revealed that F-V/51 remains recombinant at a single nucleotide polymorphism (SNP) located 1.2 kb upstream of XL. No heterozygous mutation was identified between exon XL and an SNP approximately 8 kb upstream of NESP55, where this affected individual becomes linked, suggesting that the genetic defect responsible for parathyroid hormone resistance in kindred F, and probably other PHP-Ib patients, is located >or=56 kb centromeric of the abnormally methylated exon A/B. A region upstream of the known coding exons of GNAS1 is therefore predicted to exert, presumably through imprinting of exon A/B, long-range effects on G(s)alpha expression.

Chromosome Mapping↗

A national record linkage to study acute myocardial infarction incidence and case fatality in Sweden.

BACKGROUND: During the last decades substantial temporal changes, as well as population differences, in coronary heart disease mortality have occurred in Sweden. There is little information to what extent these changes and differences also apply to myocardial infarction incidence. The aim of this paper was to describe the methods used to identify cases in a recently developed National Acute Myocardial Infarction Register in Sweden, and to present estimates of incidence and case fatality in Sweden. MATERIAL AND METHODS: Incident cases of acute myocardial infarction (AMI) were identified by record linkage of routinely collected data on hospital discharges and deaths. Case fatality within 28 days was ascertained by linkage of incident cases to the National Cause of Death Register. RESULTS: About 40 000 new cases of AMI per year were recorded in Sweden during 1987-1995. Well-known differences in incidence with regard to age and gender were observed, as well as a decline in incidence between 1987 and 1995. A similar case fatality was seen in men and women aged 30-89 among hospitalized cases. When fatal cases outside hospital were also considered the case fatality was somewhat higher in men. Examination of medical records for a national sample of ischaemic heart disease patients suggested a high sensitivity (94%) and a high positive predictive value (86%) for ICD-9 code 410 in hospital discharge data with regard to definite AMI. CONCLUSIONS: The National Acute Myocardial Infarction Register offers a new possibility to study the incidence of AMI, as well as case fatality, in Sweden.

Acute Disease↗

Molecular characterization of a parasite antigen in sera from onchocerciasis patients that is immunologically cross-reactive with human keratin.

Onchocerca volvulus is a nematode that causes severe dermatitis and blindness in humans. Prior studies have shown that immune complexes in sera from onchocerciasis patients contain parasite antigens with M(r) of 23,000 and 65,000-70,000. Monoclonal antibody OV-1 binds to these antigens and to corresponding antigens in adult worm extracts and in vitro culture supernatants. OV-1 was used to immunoscreen an O. volvulus adult worm cDNA library. Clone OV1CF contains a 1632-bp open-reading frame that codes for a protein with a predicted M(r) of 63,000. The deduced protein sequence of OV1CF has 78% identity with Ascaris lumbricoides intermediate filament A and 40%-50% identity with several mammalian intermediate filaments. Antibodies raised to OV1CF bind to human keratin, and some sera from onchocerciasis patients contain antibodies to OV1CF and to keratin. Thus, immune complexes from onchocerciasis patients contain a parasite antigen that is immunologically cross-reactive with human intermediate filament proteins.

Adolescent↗

Sequence and characteristics of IS900, an insertion element identified in a human Crohn's disease isolate of Mycobacterium paratuberculosis.

The complete sequence of an insertion element IS900 in Mycobacterium paratuberculosis is reported. This is the first characterised example of a mycobacterial insertion element. IS900 consists of 1451bp of which 66% is G + C. It lacks terminal inverted and direct repeats, characteristic of Escherichia coli insertion elements but shows a degree of target sequence specificity. A single open reading frame (ORF 1197) coding for 399 amino acids is predicted. This amino acid sequence, and to a lesser extent the nucleotide sequence, show significant homologies to IS110, an insertion element of Streptomyces coelicolor A3(2). It is proposed that IS900, IS110, and similar insertion elements recently identified in disease isolates of Mycobacterium avium are members of a phylogenetically related family. IS900 will provide highly specific markers for the precise identification of Mycobacterium paratuberculosis, useful in defining its relationship to animal and human diseases.

Amino Acid Sequence↗

Nucleotide sequence of the Xdh region in Drosophila pseudoobscura and an analysis of the evolution of synonymous codons.

The nucleotide sequence of the Xdh region of Drosophila pseudoobscura is presented. The Xdh gene structure and organization are compared with the homologous region in D. melanogaster. This locus is shown to have similar organization in the two species, although an additional intron and three insertion/deletion events are described for the D. pseudoobscura coding region. The encoded proteins are predicted to have very similar charges and hydrophobic/hydrophilic domains even though 11% of the amino acids are different. A gene 5' to Xdh, putative l(3)s12, is suggested from sequence similarity between the species. Synonymous differences at the Xdh locus between the two species are analyzed using a new method described in the preceding paper by Lewontin. This analysis shows that synonymous positions within the Xdh locus are evolving at very different rates, being dependent on level of codon redundancy. A comparison of synonymous divergence between D. melanogaster and D. pseudoobscura in five additional genes reveals variation in the level of synonymous substitution.

Amino Acid Sequence↗

In-flight measured and predicted ambient dose equivalent and latitude differences on effective dose estimates.

The results from 2 years (2001-2002) of experimental measurements of in-board radiation doses received at IBERIA commercial flights are presented. The routes studied cover the most significant destinations and provide a good estimate of the route doses as required by the new Spanish regulations on air crew radiation protection. Details on the experimental procedures and calibration methods are given. The experimental measurements from the different instruments (Tissue Equivalent Proportional Counter and the combination of a high pressure ion chamber and a high-energy neutron compensated rem-counter) and their comparison with the predictions from some route-dose codes (CARI-6, EPCARD 3.2) are discussed. In contrast with the already published data, which are mainly focused on North latitudes over parallel 50, many of the data presented in this work have been obtained for routes from Spain to Central and South America.

Aircraft↗

Impact of high-energy nuclear data on radioprotection in spallation sources.

The high-energy programme of the HINDAS European project has provided a large amount of experimental data and led to a better understanding of the spallation reaction mechanism and the development of more reliable spallation models. These data, or the new models, which have been implemented into high-energy transport codes, can be now used to predict with a larger confidence or, at least with a known uncertainty, some important quantities for the design of spallation sources. In this paper, examples concerning the residue production in a Pb-Bi target and the high-energy neutrons escaping the target are presented. In the first case, the activity and the amount of radioactive volatile elements that can be released, in case of a containment failure, are calculated and the level of confidence of the calculation is assessed. The second example shows that the models correctly predict the high-energy tail of the neutron spectrum, which is important for radioprotection in the facility.

Algorithms↗

NASA Space Radiation Transport Code Development Consortium.

Recently, NASA established a consortium involving the University of Tennessee (lead institution), the University of Houston, Roanoke College and various government and national laboratories, to accelerate the development of a standard set of radiation transport computer codes for NASA human exploration applications. This effort involves further improvements of the Monte Carlo codes HETC and FLUKA and the deterministic code HZETRN, including developing nuclear reaction databases necessary to extend the Monte Carlo codes to carry out heavy ion transport, and extending HZETRN to three dimensions. The improved codes will be validated by comparing predictions with measured laboratory transport data, provided by an experimental measurements consortium, and measurements in the upper atmosphere on the balloon-borne Deep Space Test Bed (DSTB). In this paper, we present an overview of the consortium members and the current status and future plans of consortium efforts to meet the research goals and objectives of this extensive undertaking.

Algorithms↗

Simulations of neutron transport at low energy: a comparison between GEANT and MCNP.

The use of the simulation tool GEANT for neutron transport at energies below 20 MeV is discussed, in particular with regard to shielding and dose calculations. The reliability of the GEANT/MICAP package for neutron transport in a wide energy range has been verified by comparing the results of simulations performed with this package in a wide energy range with the prediction of MCNP-4B, a code commonly used for neutron transport at low energy. A reasonable agreement between the results of the two codes is found for the neutron flux through a slab of material (iron and ordinary concrete), as well as for the dose released in soft tissue by neutrons. These results justify the use of the GEANT/MICAP code for neutron transport in a wide range of applications, including health physics problems.

Computer Simulation↗

Nucleotide sequence of the coding and flanking regions of the human parainfluenza virus type 3 fusion glycoprotein gene.

The complete nucleotide sequence of the human parainfluenza virus type 3 (HPIV3) fusion (F) protein gene has been determined. The HPIV3 F gene is 1851 nucleotides long including six U residues in the genomic RNA, which probably direct synthesis of the first few nucleotides in the F mRNA polyadenylate tail. The HPIV3 F gene contains a single long open reading frame coding for 539 amino acids. The predicted molecular weight of the unglycosylated precursor F0 protein was 60031. Four potential carbohydrate acceptor sites were identified. Comparison of the HPIV3 F protein sequence with the F gene sequences of two other paramyxoviruses, Sendai virus and simian virus 5, indicated a very close evolutionary relationship between HPIV3 and Sendai virus. Sequence analysis of HPIV3 F gene flanking regions identified signals which appear to be responsible for polymerase recognition and polyadenylation.

Amino Acid Sequence↗

Efficient migration of complex off-line computer vision software to real-time system implementation on generic computer hardware.

This paper addresses the problem of migrating large and complex computer vision code bases that have been developed off-line, into efficient real-time implementations avoiding the need for rewriting the software, and the associated costs. Creative linking strategies based on Linux loadable kernel modules are presented to create a simultaneous realization of real-time and off-line frame rate computer vision systems from a single code base. In this approach, systemic predictability is achieved by inserting time-critical components of a user-level executable directly into the kernel as a virtual device driver. This effectively emulates a single process space model that is nonpreemptable, nonpageable, and that has direct access to a powerful set of system-level services. This overall approach is shown to provide the basis for building a predictable frame-rate vision system using commercial off-the-shelf hardware and a standard uniprocessor Linux operating system. Experiments on a frame-rate vision system designed for computer-assisted laser retinal surgery show that this method reduces the variance of observed per-frame central processing unit cycle counts by two orders of magnitude. The conclusion is that when predictable application algorithms are used, it is possible to efficiently migrate to a predictable frame-rate computer vision system.

Algorithms↗

Structural basis for the enantiospecificities of R- and S-specific phenoxypropionate/alpha-ketoglutarate dioxygenases.

(R)- and (S)-dichlorprop/alpha-ketoglutarate dioxygenases (RdpA and SdpA) catalyze the oxidative cleavage of 2-(2,4-dichlorophenoxy)propanoic acid (dichlorprop) and 2-(4-chloro-2-methyl-phenoxy)propanoic acid (mecoprop) to form pyruvate plus the corresponding phenol concurrent with the conversion of alpha-ketoglutarate (alphaKG) to succinate plus CO2. RdpA and SdpA are strictly enantiospecific, converting only the (R) or the (S) enantiomer, respectively. Homology models were generated for both enzymes on the basis of the structure of the related enzyme TauD (PDB code 1OS7). Docking was used to predict the orientation of the appropriate mecoprop enantiomer in each protein, and the predictions were tested by characterizing the activities of site-directed variants of the enzymes. Mutant proteins that changed at residues predicted to interact with (R)- or (S)-mecoprop exhibited significantly reduced activity, often accompanied by increased Km values, consistent with roles for these residues in substrate binding. Four of the designed SdpA variants were (slightly) active with (R)-mecoprop. The results of the kinetic investigations are consistent with the identification of key interactions in the structural models and demonstrate that enantiospecificity is coordinated by the interactions of a number of residues in RdpA and SdpA. Most significantly, residues Phe171 in RdpA and Glu69 in SdpA apparently act by hindering the binding of the wrong enantiomer more than the correct one, as judged by the observed decreases in Km when these side chains are replaced by Ala.

2-Methyl-4-chlorophenoxyacetic Acid↗

Mutations within the protein Z-dependent protease inhibitor gene are associated with venous thromboembolic disease: a new form of thrombophilia.

Protein Z-dependent protease inhibitor (ZPI) is a serpin that inhibits the activated coagulation factors X and XI. The precise physiological significance of ZPI in the control of haemostasis is unknown although a deficiency of ZPI may be predicted to alter this balance. The coding region of the ZPI gene was screened for mutations using denaturing high-performance liquid chromatography. 16 mutations/polymorphisms within the coding region of ZPI were identified including two mutations, which generated stop codons at residues R67 and W303. We observed nonsense mutations within the ZPI gene in 4.4% of thrombosis patients (n = 250) compared with 0.8% of controls (n = 250). The difference in distribution of stop codon mutations between thrombosis patients and controls was significant (P = 0.02) with an odds ratio of 5.7 (95% confidence interval, 1.25-26.0). Our results suggest an association between ZPI deficiency and venous thrombosis and we propose that ZPI deficiency is potentially a new form of thrombophilia.

Adult↗