Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Rice tungro bacilliform virus encodes reverse transcriptase, DNA polymerase, and ribonuclease H activities.

Rice tungro bacilliform virus (RTBV) is a newly described badnavirus and proposed member of the plant pararetrovirus group. RTBV open reading frame 3 is predicted to encode a capsid protein, protease (PR), and reverse transcriptase (RT) and has the capacity to encode other proteins of as yet unknown function. To study the possible enzymatic activities encoded by open reading frame 3, a DNA fragment containing the putative PR and RT domains was used to construct the recombinant baculovirus PR/RT-BBac. Trichoplusia ni insect cells infected with PR/RT-BBac were used in pulse-labeling experiments and demonstrated synthesis of an 87-kDa polyprotein that corresponds in molecular mass to that predicted from the PR/RT DNA coding sequence. The 87-kDa polyprotein was processed with concomitant accumulation of 62-kDa (p62) and 55-kDa (p55) proteins. Amino-terminal sequencing of p62 and p55 determined that they mapped to the PR/RT domain and shared common amino termini. p62 and p55 were purified and exhibited both RT and DNA polymerase activities using synthetic primer/template substrates. Only p55 had detectable ribonuclease H activity, an activity intrinsic to all reverse transcriptases studied to date. Characterization of the RTBV RT provides a biochemical basis for classifying RTBV as a pararetrovirus and will lead to further studies of these proteins and their role in virus replication.

Amino Acid Sequence↗

Codon pair utilization biases influence translational elongation step times.

Two independent assays capable of measuring the relative in vivo translational step times across a selected codon pair in a growing polypeptide in the bacterium Escherichia coli have been employed to demonstrate that codon pairs observed in protein coding sequences more frequently than predicted (over-represented codon pairs) are translated slower than pairs observed less frequently than expected (under-represented codon pairs). These results are consistent with the findings that translational step times are influenced by codon context and that these context effects are related to the compatabilities of adjacent tRNA isoacceptor molecules on the surface of a translating ribosome. These results also support our previous suggestion that the frequency of one codon next to another has co-evolved with the structure and abundance of tRNA isoacceptors in order to control the rates of translational step times without imposing additional constraints on amino acid sequences or protein structures.

Amino Acid Sequence↗

Identification of a novel cysteine string protein variant and expression of cysteine string proteins in non-neuronal cells.

Cysteine string proteins (Csps) are synaptic vesicle proteins thought to be involved in calcium-dependent neurotransmitter release at nerve endings. Here, we report the cloning of two Csp variants, termed Csp1 and Csp2, from bovine adrenal medullary chromaffin cells. The bovine Csp1 appears to be the homologue of rat brain Csp, sharing 95% identity at the amino acid level. The nucleotide sequence of csp2 is identical with that of csp1 except for a 72-base insert which introduces a stop codon into the coding sequence, which would be predicted to result in a truncated protein 3.3 kDa smaller than Csp1. Furthermore, polymerase chain reaction analysis detected homologues of Csp1 and Csp2 in rat kidney, liver, pancreas, spleen, lung, and adrenal gland. Expression of Csps in non-neuronal tissues was confirmed by Northern blotting and by immunoblotting with anti-Csp1 antiserum which also demonstrated expression of both full-length and truncated Csps in spleen. The widespread tissue distribution is inconsistent with a role of Csps as specific regulators of presynaptic calcium channels as previously proposed. We suggest that Csps may have a more general role in membrane traffic in non-neuronal as well as neuronal cells.

Adrenal Medulla↗

Two distinct isoforms of cDNA encoding rainbow trout androgen receptors.

Androgens play an important role in male sexual differentiation and development. The activity of androgens is mediated by an androgen receptor (AR), which binds to specific DNA recognition sites and regulates transcription. We describe here the isolation of two distinct rainbow trout cDNA clones, designated rtAR-alpha and rtAR-beta, which contain the entire androgen receptor coding region. Comparison of the predicted amino acid sequence of rtAR-alpha to that of rtAR-beta revealed 85% identity. Interestingly, despite this high homology, rtAR-alpha activated transcription of an androgen-responsive reporter gene in co-transfection assays, but rtAR-beta did not. These results suggest that rainbow trout contains two distinct isoforms of androgen receptors whose functions differ. The region of rtAR-beta responsible for its inactivity was mapped to its ligand binding domain by analyzing chimeras of the rtAR-alpha, rtAR-beta, and rtGR-I (glucocorticoid) receptors. Alteration of any one of three out of four segments within this domain restored activity. Extracts made from COS-1 cells transfected with an rtAR-alpha expression plasmid produced a high level of [3H]mibolerone binding, whereas no binding was observed by extracts of cells transfected with an rtAR-beta expression plasmid. These data demonstrate that the lack of transactivation activity of rtAR-beta is due to its inability to bind hormone.

Amino Acid Sequence↗

Isolation of a new member of the S100 protein family: amino acid sequence, tissue, and subcellular distribution.

A low molecular mass protein which we term S100L was isolated from bovine lung. S100L possesses many of the properties of brain S100 such as self association, Ca++-binding (2 sites per subunit) with moderate affinity, and exposure of a hydrophobic site upon Ca++-saturation. Antibodies to brain S100 proteins, however, do not cross react with S100L. Tryptic peptides derived from S100L were sequenced revealing similarity to other members of the S100 family. Oligonucleotide probes based on these sequences were used to screen a cDNA library derived from a bovine kidney cell line (MDBK). A 562-nucleotide cDNA was sequenced and found to contain the complete coding region of S100L. The predicted amino acid sequence displays striking similarity, yet is clearly distinct from other members of the S100 protein family. Polyclonal and monoclonal antibodies were raised against S100L and used to determine the tissue and subcellular distribution of this molecule. The S100L protein is expressed at high levels in bovine kidney and lung tissue, low levels in brain and intestine, with intermediate levels in muscle. The MDBK cell line was found to contain both S100L and the calpactin light chain, another member of this protein family. S100L was not found associated with a higher molecular mass subunit in MDBK cells while the calpactin light chain was tightly bound to the calpactin heavy chain. Double label immunofluorescence microscopy confirmed the observation that the calpactin light chain and S100L have a different distribution in these cells.

Amino Acid Sequence↗

Diagnostic x-ray dosimetry using Monte Carlo simulation.

An Electron Gamma Shower version 4 (EGS4) based user code was developed to simulate the absorbed dose in humans during routine diagnostic radiological procedures. Measurements of absorbed dose using thermoluminescent dosimeters (TLDs) were compared directly with EGS4 simulations of absorbed dose in homogeneous, heterogeneous and anthropomorphic phantoms. Realistic voxel-based models characterizing the geometry of the phantoms were used as input to the EGS4 code. The voxel geometry of the anthropomorphic Rando phantom was derived from a CT scan of Rando. The 100 kVp diagnostic energy x-ray spectra of the apparatus used to irradiate the phantoms were measured, and provided as input to the EGS4 code. The TLDs were placed at evenly spaced points symmetrically about the central beam axis, which was perpendicular to the cathode-anode x-ray axis at a number of depths. The TLD measurements in the homogeneous and heterogenous phantoms were on average within 7% of the values calculated by EGS4. Estimates of effective dose with errors less than 10% required fewer numbers of photon histories (1 x 10(7)) than required for the calculation of dose profiles (1 x 10(9)). The EGS4 code was able to satisfactorily predict and thereby provide an instrument for reducing patient and staff effective dose imparted during radiological investigations.

Humans↗

Optimization of accelerator target and detector for portal imaging using Monte Carlo simulation and experiment.

Megavoltage portal images suffer from poor quality compared to those produced with kilovoltage x-rays. Several authors have shown that the image quality can be improved by modifying the linear accelerator to generate more low-energy photons. This work addresses the problem of using Monte Carlo simulation and experiment to optimize the beam and detector combination to maximize image quality for a given patient thickness. A simple model of the whole imaging chain was developed for investigation of the effect of the target parameters on the quality of the image. The optimum targets (6 mm thick aluminium and 1.6 mm copper) were installed in an Elekta SL25 accelerator. The first beam will be referred to as A16 and the second as Cu1.6. A tissue-equivalent contrast phantom was imaged with the 6 MV standard photon beam and the experimental beams with standard radiotherapy and mammography film/screen systems. The arrangement with a thin Al target/mammography system improved the contrast from 1.4 cm bone in 5 cm water to 19% compared with 2% for the standard arrangement of a thick, high-Z target/radiotherapy verification system. The linac/phantom/detector system was simulated with the BEAM/EGS4 Monte Carlo code. Contrast calculated from the predicted images was in good agreement with the experiment (to within 2.5%). The use of MC techniques to predict images accurately, taking into account the whole imaging system, is a powerful new method for portal imaging system design optimization.

Bone and Bones↗

Photoneutron yields from tungsten in the energy range of the giant dipole resonance.

Photoneutron production on the nuclei of high-Z components of medical accelerator heads can lead to a significant secondary dose during a course of bremsstrahlung radiotherapy. However, a quantitative evaluation of secondary neutron dose requires improved data on the photoreaction yields. These have been measured as a function of photon energy, neutron energy and neutron angle for natW, using tagged photons at the MAX-Lab photonuclear facility in Sweden. This work presents neutron yields for natW(gamma, n) and compares these with the predictions of the Monte Carlo code MCNP-GN, developed specifically to simulate photoneutron production at medical accelerators.

Monte Carlo Method↗

Molecular cloning, chromosomal localization, and expression analysis of CYRN1 and CYRN2, two human genes coding for cyritestin, a sperm protein involved in gamete interaction.

Germ cell cyritestin is a membrane-anchored protein belonging to the ADAM family of proteins. Sequencing of eight human cyritestin cDNA clones revealed that they are identical at their 5' and 3' ends but differ from each other in the length of an internal deletion, suggesting that the human cyritestin mRNA is alternatively spliced. Internal deletions that are present in some cDNA isoforms do not cause a frameshift in the C-terminal coding region. Analysis of the predicted amino acid sequences demonstrated that the human cyritestin is a polymorphic protein that could include membrane-anchored and soluble forms. Southern blot analysis and characterization of human cyritestin genomic fragments revealed that the human genome contains two copies of the cyritestin gene instead of one as in the mouse. The human CYRN1 and CYRN2 genes were assigned to the region p12-21 of chromosome 8 and q12 of chromosome 16, respectively. Northern blot and RT-PCR analyses revealed that both human genes are expressed in the testis. Amino acid sequence comparisons between cyritestin and other members of the metalloprotease-disintegrin family of proteins suggested that human and mouse cyritestin and monkey tMDCI are homologous molecules.

ADAM Proteins↗

Overproduction of penicillin-binding protein 7 suppresses thermosensitive growth defect at low osmolarity due to an spr mutation of Escherichia coli.

Escherichia coli delta prc mutants lacking periplasmic protease Prc, which was originally found involved in the C-terminal processing of penicillin-binding protein (PBP) 3, show thermosensitive growth at low osmolarity. We isolated thermoresistant revertants containing extragenic suppressor (spr) mutations. In the prc+ background the mutations also caused thermosensitivity at low osmolarity. They were all mapped at about 48 min on the chromosome and most probably allelic to one another. From this chromosomal region we cloned a gene that could correct the thermosensitive defect of an spr mutant, which turned out to be a multicopy suppressor of spr. Analysis of the nucleotide sequence predicted that the gene would code for a low-molecular-weight PBP, and penicillin-binding experiments revealed the product to be PBP 7. Disruption of the gene on the chromosome caused no apparent growth defect. PBP 7 seemed to be degraded by protease Prc. Overproduction of mutant PBP 7 that had the active site serine residue replaced with alanine did not correct the spr thermosensitivity, suggesting importance of the DD-endopeptidase activity in the multicopy suppression.

Alleles↗

The metalloproteinase matrilysin is preferentially expressed by epithelial cells in a tissue-restricted pattern in the mouse.

To explore the role of the matrix metalloproteinase matrilysin (MAT) in normal tissue remodeling, we cloned the murine homologue of MAT from postpartum uterus using RACE polymerase chain reaction and examined its pattern of expression in embryonic, neonatal, and adult mice. The murine coding sequence and the corresponding predicted protein sequence were found to be 75% and 70% identical to the human sequences, respectively, and organization of the six exons comprising the gene is similar to the human gene. Northern analysis and in situ hybridization revealed that MAT is expressed in the normal cycling, pregnant, and postpartum uterus, with levels of expression highest in the involuting uterus at early time points (6 h to 1.5 days postpartum). The mRNA was confined to epithelial cells lining the lumen and some glandular structures. High constitutive levels of MAT transcripts were also detected in the small intestine, where expression was localized to the epithelial Paneth cells at the base of the crypts. Similarly, MAT expression was found in epithelial cells of the efferent ducts, in the initial segment and cauda of the epididymis, and in an extra-hepatic branch of the bile duct. MAT transcripts were detectable only by reverse transcription-polymerase chain reaction in the colon, kidney, lung, skeletal muscle, skin, stomach, juvenile uterus, and normal, lactating, and involuting mammary gland, as was expression primarily late in embryogenesis. Analysis of MAT expression during postnatal development indicated that although MAT is expressed in the juvenile small intestine and reproductive organs, the accumulation of significant levels of MAT mRNA appears to correlate with organ maturation. These results show that MAT expression is restricted to specific organs in the mouse, where the mRNA is produced exclusively by epithelial cells, and suggest that in addition to matrix degradation and remodeling, MAT may play an important role in the differentiated function of these organs.

Amino Acid Sequence↗

Refining sequence-to-expression modelling with chromatin accessibility.

MOTIVATION: Sequence-to-expression models typically do not consider chromatin accessibility, a major factor limiting gene regulation. We hypothesized that supplying accessibility as an input feature would allow a sequence-to-expression model to focus on important open regions of the genome. RESULTS: We found that the performance of such an augmented model was significantly better than that of sequence-only or accessibility-only models with similar architectures. Specifically, its ability to predict the expression of highly variable genes and gene expression in other cell types improved, and higher attribution scores in the input DNA sequences of the augmented model conformed to accessibility, enabling the learning of cell type-specific sequence patterns. Additionally, we show that fine-tuning a pre-trained sequence-only model with both sequence and accessibility can boost performance further and highlight the importance of sequencing depth in sequence-to-expression prediction. AVAILABILITY AND IMPLEMENTATION: Source code is available on GitHub at https://github.com/lapohosorsolya/accessible_seq2exp.

Chromatin↗

GFPE: gene-finding program evaluation.

Gene-finding program evaluation (GFPE) is a set of Java classes for evaluating gene-finding programs. A command-line interface is also provided. Inputs to the program include the sequence data (in FASTA format), annotations of "actual" sequence features, and annotations of "predicted" sequence features. Annotation files are in the General Feature Format promoted by the Sanger center. GFPE calculates a number of metrics of accuracy of predictions at three levels:the coding level, the exon level, and the protein level.

Computational Biology↗

Integrating time-course microarray gene expression profiles with cytotoxicity for identification of biomarkers in primary rat hepatocytes exposed to cadmium.

MOTIVATION: DNA microarrays can provide information about the expression levels of thousands of genes simultaneously at the transcriptomic level, while conventional cell viability and cytotoxicity measurement methods provide information about the biological functions at the cellular level. Integrating these data at different levels provides a promising approach for evaluating or predicting how cells respond to chemical exposure. It is important to investigate the multi-scale biological system in a systematic way to better understand the gene regulation networks and signal transduction pathways involved in the cellular responses to environmental factors. RESULTS: Primary rat hepatocytes were exposed to cadmium acetate at 0, 1.25 and 2 microM. mRNA expression profiles at 0, 3, 6, 12 and 24 h were measured using the Affymetrix RatTox U34 GeneChip arrays. Simultaneously, cytotoxicity was assessed by lactase dehydrogenase leakage assay. Gene expression profiles at different time points were used to evaluate cytotoxicity at subsequent time points using partial least squares, and it was found that gene expression profiles at 0 h had the best prediction accuracy for the cytotoxicity observed at 12 h. Some biomarkers whose expression profiles showed strong relationship with cytotoxicity were identified and the underlying pathways were reconstructed to illustrate how hepatocytes respond to cadmium exposure. Permutation studies were also applied to assess the reliability of the predictive models. AVAILABILITY: Matlab source code is available upon request and DNA microarray data are available at GEO (http://www.ncbi.nlm.nih.gov/geo).

Algorithms↗

Structural analysis of a Lotus japonicus genome. III. Sequence features and mapping of sixty-two TAC clones which cover the 6.7 Mb regions of the genome.

A total of sixty-two clones were selected from a TAC (transformation-competent artificial chromosome) genomic library of the Lotus japonicus accession MG-20 based on the sequence information of expressed sequence tags (ESTs), cDNA and gene information, and their nucleotide sequences were determined. The length of the sequenced regions in this study is 6,682,189 bp, and the total length of the regions sequenced so far is 18,711,484 bp together with the nucleotide sequences of 121 TAC clones previously reported. By comparison with the sequences in protein and EST databases and analysis with computer programs for gene modeling, a total of 573 potential protein-coding genes with known or predicted functions, 91 gene segments and 272 pseudogenes were identified in the newly sequenced regions. Each of the sequenced clones was localized onto the linkage map of two accessions of L. japonicus, Gifu B-129 and Miyakojima MG-20, using simple sequence repeat length polymorphism (SSLP) or derived cleaved amplified polymorphic sequence (dCAPS) markers generated based on the nucleotide sequences of the clones. The sequence data, gene information and mapping information are available through the World Wide Web at http://www.kazusa.or.jp/lotus/.

Base Sequence↗

Structural analysis of Arabidopsis thaliana chromosome 5. VIII. Sequence features of the regions of 1,081,958 bp covered by seventeen physically assigned P1 and TAC clones.

A total of 17 Pl and TAC clones each representing an assigned region of chromosome 5 were isolated from P1 and TAC genomic libraries of Arabidopsis thaliana Columbia, and their nucleotide sequences were determined. The length of the clones sequenced in this study summed up to 1,081,958 bp. As we have previously reported the sequence of 9,072,622 bp by analysis of 125 P1 and TAC clones, the total length of the sequences of chromosome 5 determined so far is now 10,154,580 bp. The sequences were subjected to similarity search against protein and EST databases and analysis with computer programs for gene modeling. As a consequence, a total of 253 potential protein-coding genes with known or predicted functions were identified. The positions of exons which do not show apparent similarity to known genes were also assigned using computer programs for exon prediction. The average density of the genes identified in this study was 1 gene per 4277 bp. Introns were observed in 74% of the potential protein genes, and the average number per gene and the average length of the introns were 4.3 and 168 bp, respectively. The sequence data and gene information are available on the World Wide Web database KAOS (Kazusa Arabidopsis data Opening Site) at http://www.kazusa.or.jp/arabi/.

Arabidopsis↗

A Chromosome-Level Genome Assembly of the Potato Leafhopper Empoasca fabae (Hemiptera: Cicadellidae).

The potato leafhopper, Empoasca fabae (Harris, 1841), is a highly polyphagous, migratory insect pest of eastern North America that feeds on more than 200 herbaceous and woody plant species, causing substantial losses to forage and field crops. Despite its agricultural and ecological importance, no genome has been available for this species. Here, we present the first chromosome-level genome assembly of E. fabae, generated from Oxford Nanopore long reads, Illumina short reads, and Omni-C proximity-ligation data. The final assembly spans 908 Mb across 132 scaffolds, with 99.8% of the assembly captured in ten chromosome-length scaffolds (nine autosomes and an X chromosome) with a scaffold N50 of 96.2 Mb. The assembly is highly complete, recovering 92.9% of conserved hemipteran single-copy orthologs from protein annotations, and is composed of 47.6% repetitive sequence, dominated by long terminal repeat retrotransposons and unclassified elements. Read-depth comparison between male and female individuals supports assignment of a single sex-linked chromosome, consistent with an XO sex determination system. BRAKER3 gene annotation predicted 31,406 protein-coding genes after retaining the longest isoform per locus. Comparative genome analysis of the two closest related Typhlocybinae species with genomes available, Matsumurasca onukii and Hebata decipiens, revealed extensive chromosome-scale collinearity while defining a shared core gene repertoire. This reference genome provides a foundation for comparative and population genomic studies and for investigating genetic traits in this economically important crop pest species.

Animals↗

Identification and cloning of the CHL4 gene controlling chromosome segregation in yeast.

A collection of chl mutants characterized by decreased fidelity of chromosome transmission and by minichromosome nondisjunction in mitosis was examined for the ability to maintain nonessential dicentric plasmids. In one of the seven mutants analyzed, chl4, dicentric plasmids did not depress cell division. Moreover, nonessential dicentric plasmids were maintained stably without any rearrangements during many generations in the chl4 mutant. The rate of mitotic heteroallelic recombination in the chl4 mutant was not increased compared to that in an isogenic wild-type strain. Analysis of the segregation of a marked chromosome indicated that sister chromatid nondisjunction and sister chromatid loss contributed equally to chromosome malsegregation in the chl4 mutant. A genomic clone of CHL4 was isolated by complementation of the chl4-1 mutation and was physically mapped to the right arm of chromosome IV near the SUP2 gene. Nucleotide sequence analysis of CHL4 clone revealed a 1.4-kb open reading frame coding for a 53-kD predicted protein which does not have homology to published proteins. A strain containing a null allele of CHL4 is viable under standard growth conditions but has a temperature-sensitive phenotype (conditional lethality at 36 degrees). We suggest that the CHL4 gene is required for kinetochore function in the yeast Saccharomyces cerevisiae.

Amino Acid Sequence↗