Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Isolation and characterization of a partial cDNA for a human sialyltransferase.

A probe generated from the coding sequence of the rat hepatic beta-galactoside alpha 2,6-sialyltransferase was used to screen a human cDNA library constructed of human submaxillary gland mRNA lambda gt-11. We report the isolation and characterization of a human cDNA, HSM-ST1, that is putatively the human homolog of the beta-galactoside alpha 2,6-sialyltransferase. The largest human clone contains a 1.3 kb cDNA insert and is predicted to encompass 75% of the coding sequence as well as a small portion of the 3' untranslated region. Comparative analysis of this insert with the rat hepatic alpha 2,6-sialyltransferase sequence indicates 79% nucleotide similarity between the two sequences in the predicted coding region. On the amino acid level, the degree of conservation is 86%. Substantial sequence similarity is observed in the 3'-untranslated region between the rat and human sequences as well. S1 nuclease analysis was performed to demonstrate the expression of HSM-ST1 transcripts in the human hepatoma cell line, HepG2, and in the human colonic adenocarcinoma cell lines, LS174T.

Amino Acid Sequence

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55 Gb, with a scaffold N50 of 93.38 Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

Variable deletion of exon 9 coding sequences in cystic fibrosis transmembrane conductance regulator gene mRNA transcripts in normal bronchial epithelium.

The predicted protein domains coded by exons 9-12 and 19-23 of the 27 exon cystic fibrosis transmembrane conductance regulator (CFTR) gene contain two putative nucleotide-binding fold regions. Analysis of CFTR mRNA transcripts in freshly isolated bronchial epithelium from 12 normal adult individuals demonstrated that all had some CFTR mRNA transcripts with exon 9 completely deleted (exon 9- mRNA transcripts). In most (9 of 12), the exon 9- transcripts represented less than or equal to 25% of the total CFTR transcripts. However, in three individuals, the exon 9- transcripts were more abundant, comprising 39, 62 and 66% of all CFTR transcripts. Re-evaluation of the same individuals 2-4 months later showed the same proportions of exon 9- transcripts. Of the 24 CFTR alleles in the 12 individuals, the sequences of the exon-intron junctions relevant to exon 9 deletion (exon 8-intron 8, intron 8-exon 9, exon 9-intron 9, and intron 9-exon 10) were identical except for the intron 8-exon 9 region sequences. Several individuals had varying lengths of a TG repeat in the region between splice branch and splice acceptor consensus sites. Interestingly, one allele in each of the two individuals with 62 and 66% exon 9- transcripts had a TT deletion in the splice acceptor site for exon 9. These observations suggest either the unlikely possibility that sequences in exon 9 are not critical for the functioning of the CFTR or that only a minority of the CFTR mRNA transcripts need to contain exon 9 sequences to produce sufficient amounts of a normal CFTR to maintain a normal clinical phenotype.

Base Sequence

Nucleotide sequence of the pilin gene of Bacteroides nodosus 340 (serogroup D) and implications for the relatedness of serogroups.

The gene encoding pilin of Bacteroides nodosus 340 has been isolated and the nucleotide sequence determined. The gene is present as a single copy within the B. nodosus genome and a protein of Mr 16683 can be predicted from the proposed coding region. A comparison of the predicted amino acid sequence with pilin from other strains of B. nodosus indicated that the protein of strain 340 (serogroup D) has a high degree of similarity with pilin of strain 265 (serogroup H). The degree of similarity between pilins from these strains and from other B. nodosus serogroups is no greater than that between B. nodosus pilins and the homologous proteins of several different bacterial species. These findings suggest that serogroups D and H may form a subset of B. nodosus serogroups.

Amino Acid Sequence

Cloning and nucleotide sequence analysis of the dog insulin gene. Coded amino acid sequence of canine preproinsulin predicts an additional C-peptide fragment.

A 4.0-kilobase HindIII/EcoRI-cleaved dog genomic DNA fragment was shown to contain the dog insulin gene by restriction mapping using a human insulin cDNA probe. This fragment was subsequently cloned in a lambda vector, and the nucleotide sequence of the dog insulin gene was determined. As in several other species, the insulin gene of the dog is interrupted by two intervening sequences, one of 151 base pairs located in the 5' untranslated region and the other of 264 base pairs occurring within the codon of the 7th amino acid of the C-peptide. Translation of the nucleotide sequence in one frame revealed the primary structure of canine preproinsulin. An interesting feature of the coded amino acid sequence is that it predicts a C-peptide of 31 amino acids, 8 residues longer than that reported by Peterson et al. (Peterson, J. D., Nehrlich, S., Oyer, P. E., and Steiner, D. F. (1973) J. Biol. Chem. 247, 4866-4871). The additional octapeptide sequence, Glu-Val-Glu-Asp-Leu-Gln-Val-Arg, is located NH2-terminal to the 23-residue C-peptide sequence described in the earlier report. Its coding sequence is interrupted by the second intervening sequence. The arginine at position 8 suggests that a trypsin-like cleavage may separate the NH2-terminal octapeptide from the remainder of the C-peptide during the post-translational processing of dog proinsulin in the pancreas. The revised C-peptide sequence suggests that the proinsulin C-peptide is more highly conserved in length and overall sequence than was previously supposed.

Amino Acid Sequence

The hemoglobin of Urechis caupo. The cDNA-derived amino acid sequence.

The nucleotide sequence of a cDNA transcript containing part of the 5' noncoding region, the entire coding region, and the entire 3' noncoding region has been determined. The protein sequence predicted from the coding region matches almost exactly the aminoterminal sequence and the sequence of several peptides from Urechis caupo F-I globin. Only 11-20% of the amino acid positions are identical with those of other known globins.

Amino Acid Sequence

Leveraging functional annotations to map rare variants associated with Alzheimer disease with gruyere.

Increased availability of whole-genome sequencing (WGS) has facilitated the study of rare variants (RVs) in complex diseases. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most do not fully leverage the availability of variant-level functional annotations. We propose genome-wide rare variant enrichment evaluation (gruyere), an empirical Bayesian framework that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimer's Disease Sequencing Project to identify Alzheimer disease (AD)-associated genes and annotations. Growing evidence suggests that the disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not examined rare non-coding effects that incorporate such cell-type-specific information. To address this gap, we (1) define per-gene non-coding RV test sets using predicted enhancer and promoter regions in microglia and other brain cell types (oligodendrocytes, astrocytes, and neurons) and (2) include cell-type-specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 13 significant genetic associations not detected by other RV methods, four of which remain significant in omnibus tests. We find that deep-learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a robust framework incorporating functional annotations, coding RVs, and cell-type-associated non-coding RVs to perform genome-wide association tests, uncovering AD-relevant genes and annotations.

Alzheimer Disease

A 10,400-molecular-weight membrane protein is coded by region E3 of adenovirus.

Previous studies with adenovirus mutants have indicated that a 10,400-molecular-weight (10.4K) protein predicted to be coded by an open reading frame in region E3 of adenovirus functions to down regulate the epidermal growth factor receptor (C. R. Carlin, A. E. Tollefson, H. A. Brady, B. L. Hoffman, and W. S. M. Wold, Cell 57:135-144, 1989). We now demonstrate that the 10.4K protein is in fact synthesized in cells infected by group C adenoviruses. This was done by immunoprecipitation of 10.4K from cells infected by a variety of E3 mutants, using antisera against three different synthetic peptides corresponding to the predicted 10.4K sequence. The 10.4K protein was translated primarily from E3 mRNA f, as indicated by cell-free translation of mRNA purified by hybridization from cells infected with an RNA processing mutant that synthesizes predominantly mRNA f. The 10.4K protein was overproduced or underproduced in vivo, respectively, by mutants that overproduce or underproduce E3 mRNA f, also indicating that the 10.4K protein is translated primarily from mRNA f. The 10.4K protein migrated as two bands with apparent molecular weights of 16,000 and 11,000 (10 to 18% gradient gels); both bands contained 10.4K epitopes, as shown by Western blot (immunoblot). Only the 16K band was obtained by cell-free translation, suggesting that the 16K protein is the precursor to the 11K protein. The 10.4K protein is a membrane protein, as shown by cell fractionation experiments and as predicted from its sequence. The predicted 10.4K sequence as well as a putative N-terminal signal sequence and 30-residue transmembrane domain are conserved in adenovirus types 2 and 5 (group C) and in types 3, 7, and 35 (group B).

Adenoviruses, Human

Sequence characterization of the membrane protein gene of paramyxovirus simian virus 5.

The complete nucleotide sequence of the membrane (M) protein gene of the paramyxovirus simian virus 5 (SV5) was determined from cDNA clones of viral mRNAs. The M gene boundaries were determined by (i) primer extension sequencing on M mRNA; (ii) nuclease S1 analysis; and (iii) primer extension sequencing on viral genomic RNA. The M gene mRNA consisted of 1371 templated nucleotides. It contains a single large open reading frame that can encode a protein of 377 amino acids with a predicted Mr = 42,253. The authenticity of the predicted M protein coding sequence was confirmed by synthesis of the M protein from mRNA synthesized from cDNA. The predicted M amino acid sequence indicated it is an overall hydrophobic protein carrying a net positive charge. Alignment of the SV5 protein amino acid sequence with the M protein sequences of other paramyxoviruses indicated that these viruses fall into the following two groups: (1) SV5, mumps virus, and Newcastle disease virus; or (2) Sendai, parainfluenza virus type 3, measles virus, and canine distemper virus, with mumps virus M sequence being the most closely related to SV5.

Amino Acid Sequence

Molecular cloning of cDNA and analysis of protein secondary structure of Candida albicans enolase, an abundant, immunodominant glycolytic enzyme.

We isolated and sequenced a clone for Candida albicans enolase from a C. albicans cDNA library by using molecular genetic techniques. The 1.4-kbp cDNA encoded one long open reading frame of 440 amino acids which was 87 and 75% similar to predicted enolases of Saccharomyces cerevisiae and enolases from other organisms, respectively. The cDNA included the entire coding region and predicted a protein of molecular weight 47,178. The codon usage was highly biased and similar to that found for the highly expressed EF-1 alpha proteins of C. albicans. Northern (RNA) blot analysis showed that the enolase cDNA hybridized to an abundant C. albicans mRNA of 1.5 kb present in both yeast and hyphal growth forms. The polypeptide product of the cloned cDNA, which was purified as a recombinant protein fused to glutathione S-transferase, had enolase enzymatic activity and inhibited radioimmunoprecipitation of a single C. albicans protein of molecular weight 47,000. Analysis of the predicted C. albicans enolase showed strong conservation in regions of alpha helices, beta sheets, and beta turns, as determined by comparison with the crystal structure of apo-enolase A of S. cerevisiae. The lack of cysteine residues and a two-amino-acid insertion in the main domain differentiated C. albicans enolase from S. cerevisiae enolase. Immunofluorescence of whole C. albicans cells by using a mouse antiserum generated against the purified fusion protein showed that enolase is not located on the surface of C. albicans. Recombinant C. albicans enolase will be useful in understanding the pathogenesis and host immune response in disseminated candidiasis, since enolase is an immunodominant antigen which circulates during disseminated infections.

Amino Acid Sequence

The human gene for vascular endothelial growth factor. Multiple protein forms are encoded through alternative exon splicing.

Vascular endothelial growth factor (VEGF) is an apparently endothelial cell-specific mitogen that is structurally related to platelet-derived growth factor. By Northern blot and protein analyses, we show that VEGF is produced by cultured vascular smooth muscle cells. Analysis of VEGF transcripts in these cells by polymerase chain reaction and cDNA cloning revealed three different forms of the VEGF coding region, as had been reported in HL60 cells. The three forms of the human VEGF protein chain predicted from these coding regions are 189, 165, and 121 amino acids in length. Comparison of cDNA nucleotide sequences with sequences derived from human VEGF genomic clones indicates that the VEGF gene is split among eight exons and that the various VEGF coding region forms arise from this gene by alternative splicing: the 165-amino-acid form of the protein is missing the residues encoded by exon 6, whereas the 121-amino-acid form is missing the residues encoded by exons 6 and 7. Analysis of the VEGF gene promoter region revealed a single major transcription start, which lies near a cluster of potential Sp1 factor binding sites. The promoter region also contains several potential binding sites for the transcription factors AP-1 and AP-2; consistent with the presence of these sites, Northern blot analysis demonstrated that the level of VEGF transcripts is elevated in cultured vascular smooth muscle cells after treatment with the phorbol ester 12-O-tetradecanoyl-phorbol-13-acetate.

Amino Acid Sequence

The nucleotide sequence of a soybean mosaic virus coat protein-coding region and its expression in Escherichia coli, Agrobacterium tumefaciens and tobacco callus.

A DNA complementary to the 3'-terminal 1168 nucleotides of the genome of the N strain of soybean mosaic virus (SMV) has been cloned and sequenced. cDNA sequence and coat protein analyses indicate that the SMV coat protein-coding region is at the 3' end of the genome, and that the coat protein is processed from a larger protein. The coat protein-coding sequence is predicted to be 795 nucleotides in length, encoding a protein of 265 amino acids with a calculated Mr of 29,857. The 3' untranslated region is 259 nucleotides in length and is followed by a polyadenylate tract. The SMV coat protein-coding region, along with a small amount of upstream sequence, has been expressed in Escherichia coli as a beta-galactosidase fusion protein. The size of the protein was less than predicted for the fusion protein, suggesting processing in E. coli. The coat protein-coding region has also been expressed in Agrobacterium tumefaciens and transgenic tobacco callus as an unfused protein under the control of the cauliflower mosaic virus 35S promoter. The coat protein produced in transgenic tobacco callus had an electrophoretic mobility identical to that of SMV coat protein and constituted approximately 0.05% (w/w) of the total extracted protein.

Amino Acid Sequence

cDNA cloning of the B cell membrane protein CD22: a mediator of B-B cell interactions.

We have cloned a full-length cDNA for the B cell membrane protein CD22, which is referred to as B lymphocyte cell adhesion molecule (BL-CAM). Using subtractive hybridization techniques, several B lymphocyte-specific cDNAs were isolated. Northern blot analysis with one of the clones, clone 66, revealed expression in normal activated B cells and a variety of B cell lines, but not in normal activated T cells, T cell lines, Hela cells, or several tissues, including brain and placenta. One major transcript of approximately 3.3 kb was found in B cells although several smaller transcripts were also present in low amounts (approximately 2.6, 2.3, and 1.6 kb). Sequence analysis of a full-length cDNA clone revealed an open reading frame of 2,541 bases coding for a predicted protein of 847 amino acids with a molecular mass of 95 kD. The BL-CAM cDNA is nearly identical to a recently isolated cDNA clone for CD22, with the exception of an additional 531 bases in the coding region of BL-CAM. BL-CAM has a predicted transmembrane spanning region and a 140-amino acid intracytoplasmic domain. Search of the National Biological Research Foundation protein database revealed that this protein is a member of the immunoglobulin super family and that it had significant homology with three homotypic cell adhesion proteins: carcinoembryonic antigen (29% identity over 460 amino acids), myelin-associated glycoprotein (27% identity over 425 amino acids), and neural cell adhesion molecule (21.5% over 274 amino acids). Northern blot analysis revealed low-level BL-CAM mRNA expression in unactivated tonsillar B cells, which was rapidly increased after B cell activation with Staphylococcus aureus Cowan strain 1 and phorbol myristate acetate, but not by various cytokines, including interleukin 4 (IL-4), IL-6, and gamma interferon. In situ hybridization with an antisense BL-CAM RNA probe revealed expression in B cell-rich areas in tonsil and lymph node, although the most striking hybridization was in the germinal centers. COS cells transfected with a BL-CAM expression vector were immunofluorescently stained positively with two different CD22 antibodies, each of which recognizes a different epitope. Additionally, both normal tonsil B cells and a B cell line were found to adhere to COS transfected with BL-CAM in the sense but not the antisense direction.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

A DNA fragment hybridizing to a nif probe in Rhodobacter capsulatus is homologous to a 16S rRNA gene.

We have sequenced the Rhodobacter capsulatus nifH and nifD genes. The nifH gene, which codes for the dinitrogenase reductase protein, is 894 bp long and codes for a polypeptide of predicted Mr 32,412. The nifD gene, which codes for the alpha subunit of dinitrogenase, is 1,500 bp long and codes for a protein of predicted Mr 56,113. A 776-bp BglII-XhoI fragment containing only nif sequences was used as a hybridization probe against R. capsulatus genomic DNA. Two HindIII fragments, 11.8 kb and 4.7 kb in length, hybridize to this probe. Both fragments have been cloned from a cosmid library. The 11.8-kb fragment contains the nifH, D and K genes, as previously demonstrated (Scolnik and Haselkorn, 1984). In this paper we present evidence that suggests that the 4.7-kb HindIII fragment contains a gene coding for 16S rRNA, and that although homology between nif and this fragment can be observed in filter hybridization experiments, a second copy of the nif structural genes seems not to be present in this region.

Base Sequence

The murine complement receptor gene family. IV. Alternative splicing of Cr2 gene transcripts predicts two distinct gene products that share homologous domains with both human CR2 and CR1.

The murine Cr2 gene encodes at least two related proteins. The first of these is predicted to include 1408 amino acids from a transcript including 4224 coding nucleotides. This protein is predicted to contain 21 60-amino acid repeats plus those residues encoding transmembrane and cytoplasmic regions for a total peptide m.w. of 155,307. The first six of these repeats are similar to human CR1 in sequence and organization. The second protein is encoded by an alternatively spliced Cr2 transcript that is lacking those sequences which encode the first six 60-amino acid repeats of the larger protein. This second, smaller protein, encoded by a transcript of 3096 coding nucleotides, is predicted to include 15 60-amino acid repeats, plus the transmembrane and cytoplasmic regions. This smaller protein contains 1032 amino acids for a peptide m.w. of 113,328. This second protein is extremely homologous in size and sequence to human CR2. Both proteins share the same signal sequence for membrane insertion. DNA sequence analysis, RNA protection studies and genomic phage mapping indicate the transcripts which encode these proteins are derived from the Cr2 gene via alternative splicing.

Amino Acid Sequence

Simulation of CRISPR/Cas9-mediated gene editing for the Vitellogenin gene in Apis mellifera.

CRISPR/Cas9 genome editing provides a powerful framework for interrogating gene function in Apis mellifera. Yet, empirical application remains challenging due to biological constraints, including haplodiploid genetics, narrow embryonic injection window, and the social rearing requirements that complicate functional validation. These constraints necessitate in silico pre-screening to maximize editing success before resource-intensive wet-lab implementation. Within the omnigenic framework, which distinguishes core regulatory genes from peripheral loci buffered by network effects, vitellogenin (Vg) represents an optimal target which is ancestrally dedicated to yolk provisioning; it has been co-opted to orchestrate diverse non-reproductive functions including longevity, stress resistance, immunity, and social behavior. We developed a computational pipeline to design a list of 57 and 56 candidate guide RNAs (gRNA) for targeted Vg knockout, evaluating candidate sites in both functional exons 2 and 3 based on structural accessibility and frameshift efficiency. Comparative analysis revealed complementary strengths in two top-best candidates from initial target pool of predicted gRNAs. The gRNA targeting exon 2 exhibits weaker secondary structure (ΔG = -0.25 kcal/mol versus -2.10 kcal/mol for exon 3), aligning with empirical evidence that sites with ΔG > -1.0 kcal/mol achieve 2-5 × higher Cas9 binding efficiency. This site yielded moderate frameshift frequency (77.8%; 61.9 percentile). Conversely, the predicted editing outcome for the gRNA targeting exon 3, despite stronger structural constraints, demonstrated superior functional disruption metrics demonstrating very high frameshift frequency (88.3%; 95.2 percentile), high in silico editing precision, minimal microhomology-mediated repair bias, and reproducible outcomes wherein nearly all predicted indels disrupt the coding sequence. Protein structure and domain analyses further predict that frameshift edits will generate a truncated protein missing all downstream functional domains. We recommend parallel empirical validation of both exon 2 and exon 3 targets to resolve the trade-off between structural accessibility (favoring higher editing rates) and frameshift efficacy (favoring complete loss-of-function). This dual-target strategy accommodates uncertainty in in vivo performance while maximizing the probability of generating informative phenotypes. Our in silico framework enables rational CRISPR design in non-model organisms by computationally balancing biophysical accessibility with functional impact, accelerating functional genomics in species where empirical optimization faces substantial biological constraints.

Animals

Real-time distortionless high-factor compression scheme.

Nowadays, digital subtraction angiography systems must be able to sustain real-time acquisition (30 frames per second) of 512 x 512 x 8 bit images and store several sequences of such images on low cost and general-purpose mass memories. Concretely, that means a 7.8 Mbytes per second rate and about 780 Mbytes disk space to hold a 100-s cardiac examination. To fulfill these requirements at competitive cost, a distortionless compressor/decompressor system can be designed: during acquisition, the real-time compressor transforms the input images into a lower quantity of coded information through a predictive coder and a variable-length Huffman code. The process is fully reversible because during review, the real-time decompressor exactly recovers the acquired images from the stored compressed data. Test results on many raw images demonstrate that real-time compression is feasible and takes place with absolutely no loss of information. The designed system indifferently works on 512 or 1024 formats, and 256 or 1024 gray levels.

Angiography, Digital Subtraction

Loci for act recall: contextual influence on the processing of action events.

A problem-solving account of act memory predicts stronger impacts of context than theories that explain act memory by reference to automatic processing or by reference to the operation of modality-specific code systems. This prediction was tested in three experiments, all using the loci memotechnique to provide contexts for memorization of subject-performed tasks (SPTs). The results of the three experiments did not provide unambiguous evidence for or against any of the rival theories. Most consistent, however, was the observation that memory under motor-encoding conditions profits less on contexts than memory under nonmotor-encoding conditions, a finding which by itself lends more support to a multicode than to a problem-solving interpretation.

Adolescent