Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A gene required for very short patch repair in Escherichia coli is adjacent to the DNA cytosine methylase gene.

Deamination of 5-methylcytosine in DNA results in T/G mismatches. If unrepaired, these mismatches can lead to C-to-T transition mutations. The very short patch (VSP) repair process in Escherichia coli counteracts the mutagenic process by repairing the mismatches in favor of the G-containing strand. Previously we have shown that a plasmid containing an 11-kilobase fragment from the E. coli chromosome can complement a chromosomal mutation defective in both cytosine methylation and VSP repair. We have now mapped the regions essential for the two phenotypes. In the process, we have constructed plasmids that complement the chromosomal mutation for methylation, but not for repair, and vice versa. The genes responsible for these phenotypes have been identified by DNA sequence analysis. The gene essential for cytosine methylation, dcm, is predicted to code for a 473-amino-acid protein and is not required for VSP repair. It is similar to other DNA cytosine methylases and shares extensive sequence similarity with its isoschizomer, EcoRII methylase. The segment of DNA essential for VSP repair contains a gene that should code for a 156-amino-acid protein. This gene, named vsr, is not essential for DNA methylation. Remarkably, the 5' end of this gene appears to overlap the 3' end of dcm. The two genes appear to be transcribed from a common promoter but are in different translational registers. This gene arrangement may assure that Vsr is produced along with Dcm and may minimize the mutagenic effects of cytosine methylation.

Amino Acid Sequence

Clinical, microbiological, and biochemical factors in recurrent bacterial vaginosis.

Because so little is known about the pathogenesis of recurrent bacterial vaginosis (BV), a longitudinal microbiological study was conducted on 13 women with recurrent BV treated sequentially with conventional metronidazole therapy. A rapid clinical response characterized by disappearance of mal odor and an improvement in vaginal discharge occurred in 92% of 31 clinical episodes of BV, with patients no longer satisfying the composite clinical criteria for the diagnosis of BV. However, prospective evaluation of these asymptomatic women revealed profound residual biochemical and microbial abnormalities which were best evident on Gram stain and wet mount examination of vaginal secretions. Other common residual abnormalities included mild persistent elevation of vaginal pH and polyamine and fatty acid levels and the presence of clue cells in small numbers. Residual abnormalities could be quantified to create an overall symptom code which predicted recurrence, and it was found that the severity of residual abnormalities was inversely related to the time required until the next recurrence occurred. The severity and prevalence of residual abnormalities following clinically successful therapy support the concept that BV recurrence, especially when it is early, represents a relapse rather than a reinfection. This concept may have important therapeutic implications.

Adult

Herpes simplex virus 1 protein kinase is encoded by open reading frame US3 which is not essential for virus growth in cell culture.

Earlier reports have described a novel protein kinase in cells infected with herpes simplex or pseudorabies viruses. These novel enzymes were characterized by their acceptance of protamine as a substrate and by their differential chromatographic behavior in anion-exchange chromatography. We report that this activity was not present in extracts of uninfected cells or of cells infected with a mutant constructed so as to contain a deletion in the US3 open reading frame mapping in the small component of herpes simplex virus 1 DNA. The activity was present in extracts of cells infected with wild-type virus and with a recombinant in which the US3 open reading frame had been rescued. Our results are consistent with the observation reported earlier that the coding sequences predict an amino acid motif common to protein kinases and lead to the conclusion that the US3 open reading frame encodes a virus-specific protein kinase that is not required for virus growth in cells in culture.

Amino Acid Sequence

Identification of Epstein-Barr virus terminal protein 1 (TP1) in extracts of four lymphoid cell lines, expression in insect cells, and detection of antibodies in human sera.

The terminal proteins TP1 and TP2 are putative products of Epstein-Barr virus (EBV) genes expressed during the latent cycle of the virus. They are predicted to code for 53- and 40-kilodalton integral membrane proteins. We used the baculovirus Autographa californica nuclear polyhedrosis virus as an expression vector to produce TP1 in large amounts in insect cells. The DNA sequences used to express TP1 originated from a TP1 cDNA derived from an M-ABA/CBL1 cDNA library. Rabbit antisera raised against procaryotic TP1 fusion proteins recognized a monomer and a dimer of the recombinant TP1 protein in the infected insect cells. Immunofluorescence studies of living insect cells showed that the recombinant protein is located in the plasma membrane. The insect cells infected with the recombinant baculovirus producing TP1 provided a test system to screen human antisera for TP1 antibodies. A total of 168 human EBV-positive and EBV-negative antisera were studied. TP1 antibodies were detected only in sera from nasopharyngeal carcinoma patients (16 out of 42). Rabbit antiserum raised against the recombinant TP1 protein expressed in the baculovirus system specifically recognized a protein of about 54 kilodaltons in the lymphoblastoid cell lines M-ABA and M-ABA/CBL1 and in the Burkitt's lymphoma cell lines BL18 and BL72. This protein could be located in the total membrane fraction of M-ABA cells and is upregulated by treating the cells with 12-O-tetradecanoylphorbol-13-acetate.

Animals

RPC53 encodes a subunit of Saccharomyces cerevisiae RNA polymerase C (III) whose inactivation leads to a predominantly G1 arrest.

RPC53 is shown to be an essential gene encoding the C53 subunit specifically associated with yeast RNA polymerase C (III). Temperature-sensitive rpc53 mutants were generated and showed a rapid inhibition of tRNA synthesis after transfer to the restrictive temperature. Unexpectedly, the rpc53 mutants preferentially arrested their cell division in the G1 phase as large, round, unbudded cells. The RPC53 DNA sequence is predicted to code for a hydrophilic M(r)-46,916 protein enriched in charged amino acid residues. The carboxy-terminal 136 amino acids of C53 are significantly similar (25% identical amino acid residues) to the same region of the human BN51 protein. The BN51 cDNA was originally isolated by its ability to complement a temperature-sensitive hamster cell mutant that undergoes a G1 cell division arrest, as is true for the rpc53 mutants.

Amino Acid Sequence

The two alpha-tubulin genes of Chlamydomonas reinhardi code for slightly different proteins.

Full-length cDNA clones corresponding to the transcripts of the two alpha-tubulin genes in Chlamydomonas reinhardi were isolated. DNA sequence analysis of the cDNA clones and cloned gene fragments showed that each gene contains 1,356 base pairs of coding sequence, predicting alpha-tubulin products of 451 amino acids. Of the 27 nucleotide differences between the two genes, only two result in predicted amino acid differences between the two gene products. In the more divergent alpha 2 gene, a leucine replaces an arginine at amino acid 308, and a valine replaces a glycine at amino acid 366. The results predicted that two alpha-tubulin proteins with different net charges are produced as primary gene products. The predicted amino acid sequences are 86 and 70% homologous with alpha-tubulins from rat brain and Schizosaccharomyces pombe, respectively. Each gene had two intervening sequences, located at identical positions. Portions of an intervening sequence highly conserved between the two beta-tubulin genes are also found in the second intervening sequence of each of the alpha genes. These results, together with our earlier report of the beta-tubulin sequences in C. reinhardi, present a picture of the total complement of genetic information for tubulin in this organism.

Amino Acid Sequence

Verbal recoding of visual stimuli impairs mental image transformations.

Two experiments were carried out to test the hypothesis that verbal recoding of visual stimuli in short-term memory influences long-term memory encoding and impairs subsequent mental image operations. Easy and difficult-to-name stimuli were used. When rotated 90 degrees counterclockwise, each stimulus revealed a new pattern consisting of two capital letters joined together. In both experiments, subjects first learned a short series of stimuli and were then asked to rotate mental images of the stimuli in order to detect the hidden letters. In Experiment 1, articulatory suppression was used to prevent subjects from subvocal rehearsal when learning the stimuli, whereas in Experiment 2, verbal labels were presented with each stimulus during learning to encourage a reliance on the verbal code. As predicted, performance in the imagery task was significantly improved by suppression when the stimuli were easy to name (Experiment 1) but was severely disrupted by labeling when the stimuli were difficult to name (Experiment 2). We concluded that verbal recoding of stimuli in short-term memory during learning disrupts the ability to generate veridical mental images from long-term memory.

Attention

Isolation, expression, and mutation of a rabbit skeletal muscle cDNA clone for troponin I. The role of the NH2 terminus of fast skeletal muscle troponin I in its biological activity.

A cDNA for rabbit fast skeletal muscle troponin I (TnI) was isolated and sequenced. The clone contains a coding sequence predicting a 182-amino-acid protein with a molecular mass of 21,162 daltons. The translated sequence is different from that reported by Wilkinson and Grand (Wilkinson, J. M., and Grand, R. J. A. (1978) Nature 271, 31-35) in that Arg-153, Asp-154, and Leu-155 must be inserted into their original sequence. Amino acid sequencing of adult rabbit TnI confirmed this result. In order to investigate the role of the NH2 terminus of TnI in its biological activity, we have expressed a recombinant deletion mutant (TnId57), which lacks residues 1-57, in a bacterial expression system. Both wild type TnI (WTnI) and TnId57 inhibited acto-S1-ATPase activity and this inhibition could be fully reversed by troponin C (TnC) in the presence of Ca2+. Additionally both WTnI and TnId57 bound to an actin affinity column. Thus, both inhibitory actin binding and Ca(2+)-dependent neutralization by TnC were retained in TnId57. TnC affinity chromatography was used to compare the binding of TnI and TnId57 to TnC. Using this method, two types of interaction between TnC and TnI were observed: 1) one which is metal independent (or structural) and 2) one dependent on Ca2+ or Mg2+ binding to the Ca(2+)-Mg2+ sites of TnC. The same experiments with TnId57 demonstrated that the type 1 interaction was weakened, and type 2 binding was lost. This method also revealed an interaction between TnC and TnI which is dependent upon Ca2+ binding to the Ca(2+)-specific sites of TnC and which is retained in TnId57. Taken together, these results suggest that the NH2 terminus of TnI may constitute a Ca(2+)-Mg(2+)-dependent interaction site between TnC and TnI and play, in part, a structural role in maintaining the stability of the troponin complex while the COOH terminus of TnI contains a Ca(2+)-specific site-dependent interaction site for TnC as well as the previously demonstrated Ca(2+)-sensitive inhibitory and actin binding activities.

Actins

Bacillus subtilis alkaline phosphatases III and IV. Cloning, sequencing, and comparisons of deduced amino acid sequence with Escherichia coli alkaline phosphatase three-dimensional structure.

Bacillus subtilis has an alkaline phosphatase multigene family. Two members of this gene family, phoAIII and phoAIV, were cloned, taking advantage of in vitro constructed strains containing a plasmid insertion within one or the other of the structural genes. The DNA sequences of the two genes showed approximately 64% identity at the DNA level and 63% identity in the deduced primary amino acid sequences. The phoAIII and phoAIV genes code for predicted proteins of 47,149 and 45,935 Da, respectively. Comparison of the deduced primary amino acid sequence of the mature proteins with other sequenced alkaline phosphatases from Escherichia coli, yeast, and humans shows 25-30% identity. Based on the refined crystal structure of E. coli alkaline phosphatase, it appears that the active site and the core of the structure are retained in both Bacillus alkaline phosphatases. However, both proteins are truncated at the amino terminus compared with other mature alkaline phosphatases, three sizable surface loops of E. coli are deleted, and a minidomain is replaced with a larger domain in the model. Neither Bacillus alkaline phosphatase sequenced contains any cysteine residues, an amino acid implicated in intrachain disulfide bond formation in other alkaline phosphatases.

Alkaline Phosphatase

Single amino acid substitution defines a naturally occurring genetic variant of human thymidylate synthase.

Previously, we identified an altered structural form of thymidylate synthase (TS) in a human colonic tumor cell line. This form, which is encoded by a variant structural gene, renders cells relatively resistant to 5-fluoro-2'-deoxyuridine as a result of the reduced affinity of the enzyme for the active metabolite 5-fluoro-2'-deoxyuridylic acid. We have isolated a cDNA clone specific to the altered TS and have determined its sequence. Two point mutations distinguish the normal from the altered TS mRNAs. One, a (A----G) change, is located within the 3'-untranslated region; the other, a T----C change within the amino acid-coding region, predicts replacement of tyrosine by histidine at residue 33 of the polypeptide. This sequence change was confirmed by direct analysis of cDNA amplified by the polymerase chain reaction and was further verified using allele-specific oligonucleotides as probes in Northern blots. These results, along with studies by other laboratories showing Tyr33 to be evolutionarily conserved, suggest that this residue plays an important role in TS function.

Amino Acid Sequence

Bovine alpha 1----3-galactosyltransferase: isolation and characterization of a cDNA clone. Identification of homologous sequences in human genomic DNA.

We have isolated, by immunological screening of a lambda gt11 expression library, a cDNA clone that represents the complete coding sequence for bovine alpha 1----3-galactosyltransferase. The coding sequence predicts a membrane-bound protein with three distinct structural features: a large, potentially glycosylated COOH-terminal domain (346 amino acids), a single transmembrane domain (16 amino acids), and a short NH2-terminal domain (6 amino acids). Thus, the domain structure for this transferase is similar to that deduced for beta 1----4-galactosyltransferase (Shaper, N. L., Hollis, G. F., Douglas, J. G., Kirsch, I. R., and Shaper, J. H. (1988) J. Biol. Chem. 263, 10420-10428) and alpha 2----6-sialyltransferase (Weinstein, J., Lee, E. V., McEntee, K., Lai, P.-H., and Paulson, J. C. (1987) J. Biol. Chem. 262, 17735-17743). S1 analysis demonstrates that two sets of mRNAs, which are heterogeneous at their 5' ends, are transcribed. Because both sets initiate upstream of the translational start site, only one protein is encoded by this gene. alpha 1----3-Galactosyltransferase is widely expressed in different mammalian species, with the notable exception of man and Old World monkeys (Galili, U., Shohet, S. B., Kobrin, E., Stults, C.L.M., and Macher, B. A. (1988) J. Biol. Chem. 263, 17755-17762). By Northern blot analysis we were indeed unable to detect transcripts for this enzyme in various human and Old World monkey cell lines; transcripts were readily detected in other mammalian species. However, by Southern blot analysis, homologous sequences for alpha 1----3-galactosyltransferase were identified in human genomic DNA. This suggests that the gene, although present in the human genome, is normally not expressed. These observations have potential medical implications. Because many humans have high levels of circulating antibodies directed against the enzymatic product of alpha 1----3-galactosyltransferase (Gal alpha 1----3Gal beta 1----4GlcN Ac) (Galili, U., Clark, M. R., Shohet, S. B., Buehler, J., and Macher, B. A. (1987) Proc. Natl. Acad. Sci. U. S. A. 84, 1369-1373), it has been suggested that activation of this normally silent gene may play a role in autoimmune disease in man (Etienne-Decerf, J., Malaise, M., Mahieu, P., and Winand, R. (1987) Acta Endocrinol. 115, 67-74).

Amino Acid Sequence

cDNA and gene sequence of Manduca sexta arylphorin, an aromatic amino acid-rich larval serum protein. Homology to arthropod hemocyanins.

The serum (storage) proteins produced by insect larvae at the end of the feeding cycle are hexameric blood proteins with one or more type of subunits. The cDNA and gene structure of the aromatic amino acid-rich larval serum protein arylphorin from the tobacco hornworm, Manduca sexta, has been determined. In M. sexta arylphorin there are two subunits alpha and beta, which have 686 and 687 amino acids, respectively, and whose amino acid sequences are 68% identical. The two genes, separated by 7.1 kilobases of chromosomal DNA, are transcribed in the same direction. Based on the alignment of the amino acid sequence, the rate of nucleotide substitution between the two coding regions predicts that the two genes diverged about 100 million years ago. Both genes contain 5 exons and the upstream region contains a sequence, TGATAAA, which is similar to a sequence found in all other storage protein genes for which information is available. When the National Biomedical Research Foundation protein sequence data base was searched, it was found that the arylphorin subunits showed significant similarity to the arthropod hemocyanins, which are hexameric oxygen-carrying proteins. Based on the alignment of the sequence of M. sexta arylphorin and the hemocyanin from the spiny lobster (Panulirus interruptus), for which a 3.2 A structure has been determined, it was observed that the highest concentration of conserved residues were found in those regions of the sequence which are involved in subunit interactions in the hexameric protein. It is suggested that the insect storage proteins and the arthropod hemocyanins have evolved from a common ancestor.

Amino Acid Sequence

Extensive sequence homology of the goldfish ras gene to mammalian ras genes.

We cloned ras-related sequences from goldfish genomic libraries constructed as recombinants using the lambda phage. Restriction enzyme mapping of the clones obtained revealed three kinds of ras-related sequences among approximately 350,000 genomic clones. One of these clones was partially sequenced. Comparison with the nucleotide sequences of mammalian ras genes showed that the determined sequences covered the predicted amino acid coding regions and parts of the intervening regions. The predicted amino acid sequences of the cloned ras-related goldfish gene suggested that the coding region is localized separately in DNA, and that its exon-intron boundaries are exactly the same as those of corresponding mammalian genes. The nucleotide and amino acid sequences of the goldfish ras-related gene may have extensive homologies to mammalian p 21 protein. Among the three mammalian ras proteins, the predicted amino acid sequence of the sequenced ras-related goldfish clone is most closely homologous (96%) to the Kirsten ras protein. Differences in the predicted amino acid sequence were greatest in the sequence predicted from the fourth exon; fewer differences were found in the sequence from the third exon, and only slight or no differences were found in the sequence predicted for the first and second exons. The 12th and 61st amino acids from the N-terminal of the protein, which are thought to be critical positions for GTP binding and catalysis, are both conserved in the goldfish protein.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Cloning and sequencing the HinfI restriction and modification genes.

The HinfI restriction and modification genes were cloned on a 3.9-kb PstI fragment inserted into the PstI site of plasmid pBR322. Both genes are confined to an internal 2.3-kb BclI-AvaI subfragment. This subfragment was sequenced. Two large open reading frames (ORF's) are present. ORF1 codes for the methylase [predicted 359 amino acids (aa)] and ORF2 codes for the endonuclease (predicted 262 or 272 aa).

Amino Acid Sequence

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny

Chromosome-level genome assembly of Nothapodytes nimmoniana.

Nothapodytes nimmoniana is a plant species belonging to the genus Nothapodytes in the family Icacinaceae. This species holds significant medicinal value due to its camptothecin content. In this study, we present the first chromosome-level genome assembly of N. nimmoniana constructed using NGS, Hi-C, and HiFi sequencing technologies. The assembled genome spans 3.53&#x2009;Gb across 14 chromosomes, with an N50 length of 248.74&#x2009;Mb. Genome annotation revealed that repetitive sequences constitute 80.82% of the genome size, and 83,269 protein-coding genes were predicted. Additionally, 4,360,538&#x2009;bp of non-coding RNA were annotated. This genomic resource provides a foundation for further investigation into camptothecin biosynthesis pathways and plant phylogeny in N. nimmoniana.

Genome, Plant

Concordance of experimentally mapped or predicted Z-DNA sites with positions of selected alternating purine-pyrimidine tracts.

The recent electronmicroscopic and biochemical mapping of Z-DNA sites in phi X174, SV40, pBR322 and PM2 DNAs has been used to determine two sets of criteria for identification of potential Z-DNA sequences in natural DNA genomes. The prediction of potential Z-DNA tracts and corresponding statistical analysis of their occurrence have been made on a sample of 14 DNA genomes. Alternating purine and pyrimidine tracts longer than 5 base pairs in length and their clusters (quasi alternating fragments) in the 14 genomes studied are under-represented compared to the expectation from corresponding random sequences. The fragments [d(G X C)]n and [d(C X G)]n (n greater than or equal to 3) in general do not occur in circular DNA genomes and are under-represented in the linear DNAs of phages lambda and T7, whereas in linear genomes of adenoviruses they are strongly over-represented. With minor exceptions, potential Z-DNA sites are also under-represented compared to random sequences. In the 14 genomes studied, predicted Z-DNA tracts occur in non-coding as well as in protein coding regions. The predicted Z-DNA sites in phi X174, SV40, pBR322 and PM2 correspond well with those mapped experimentally. A complete listing together with a compact graphical representation of alternating purine-pyrimidine fragments and their Z-forming potential are presented.

Animals

Isolation and characterization of a partial cDNA for a human sialyltransferase.

A probe generated from the coding sequence of the rat hepatic beta-galactoside alpha 2,6-sialyltransferase was used to screen a human cDNA library constructed of human submaxillary gland mRNA lambda gt-11. We report the isolation and characterization of a human cDNA, HSM-ST1, that is putatively the human homolog of the beta-galactoside alpha 2,6-sialyltransferase. The largest human clone contains a 1.3 kb cDNA insert and is predicted to encompass 75% of the coding sequence as well as a small portion of the 3' untranslated region. Comparative analysis of this insert with the rat hepatic alpha 2,6-sialyltransferase sequence indicates 79% nucleotide similarity between the two sequences in the predicted coding region. On the amino acid level, the degree of conservation is 86%. Substantial sequence similarity is observed in the 3'-untranslated region between the rat and human sequences as well. S1 nuclease analysis was performed to demonstrate the expression of HSM-ST1 transcripts in the human hepatoma cell line, HepG2, and in the human colonic adenocarcinoma cell lines, LS174T.

Amino Acid Sequence