Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “cDNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Human mb-1 gene: complete cDNA sequence and its expression in B cells bearing membrane Ig of various isotypes.

The transmembrane protein, IgM-alpha, a product of mb-1 gene, has been shown to be specifically associated with membrane-bound IgM on the plasma membrane of B lymphocytes. Recent studies have suggested that IgM-alpha may play a role in transducing signals from the Ag receptors during the activation of B cells. A large amount of information has been obtained in the mouse system regarding IgM-alpha and other components of the newly conceived B cell Ag receptor complex. Here we report the cloning and the nucleotide sequencing of cDNA clones of human mb-1, covering the entire length of the mRNA. At the amino acid sequence level, human and murine mb-1 share a high homology in their transmembrane and intracytoplasmic segments, suggesting an important biologic function for these regions of mb-1. A major difference, mainly in the 3' untranslated part, exists between our cDNA sequence and the published partial human mb-1 cDNA sequence. It has also been observed that human mb-1 is expressed not only by B cell lines expressing membrane-bound Ig of mu and delta isotypes but also those expressing membrane-bound Ig of alpha and gamma isotypes.

Amino Acid Sequence↗

The cDNA sequence of horse transferrin.

The cDNA sequence of horse transferrin was determined by sequencing clones isolated from a horse liver cDNA library and clones obtained by PCR. The 2305 bp horse transferrin cDNA sequence included part of the 5' untranslated region and extended to the poly(A) tail. It had 80% sequence identity with the human transferrin cDNA, and encoded a protein of 706 residues, including a signal sequence of 19 amino acids. The horse transferrin sequence had the duplicated structure and conserved iron binding and cysteine residues which are characteristic of the transferrin family.

Amino Acid Sequence↗

Amino acid sequence of a major human amelogenin protein employing Edman degradation and cDNA sequencing.

The abundant hydrophobic, proline-glutamine, and histidine-rich (over 90%) amelogenins constitute the major class of proteins in forming extracellular enamel matrix. These are thought to play a major role in the structural organization and mineralization of developing enamel. The present report describes the successful sequencing of the major human amelogenin protein, by use of both Edman degradation and cDNA sequencing. When Edman degradation was used, over 75% of the primary structure of the protein was determined. This sequence was supplemented with cDNA sequencing studies, which revealed the predicted sequence of this protein. Together, they provide the complete sequence of an important human enamel protein. The information complements recent studies on bovine and human amelogenin genes. A comparison between the present results and the protein sequences predicted from the corresponding human amelogenin genomic coding regions and that of cDNA sequences of other species is described.

Amelogenin↗

Isolation, sequencing and expression of cDNA sequences encoding uroporphyrinogen decarboxylase from tobacco and barley.

We have cloned and sequenced a full-length cDNA for uroporphyrinogen decarboxylase (UROD, EC 4.1.1.37) from tobacco (Nicotiana tabacum L.) and a partial cDNA clone from barley (Hordeum vulgare L.). The cDNA of tobacco encodes a protein of 43 kDa, which has 33% overall similarity to UROD sequences determined from other organisms. We propose that tobacco UROD has an N-terminal extension of 39 amino acid residues. This extension is most likely a chloroplast transit sequence. The in vitro translation product of UROD was imported into pea chloroplasts and processed to ca. 39 kDa. A truncated cDNA, from which the putative transit peptide had been deleted, was used to over-express the mature UROD in Escherichia coli. Purified protein showed UROD activity, thus providing an adequate source for subsequent enzymatic characterization and inhibition studies. Expression of UROD was investigated by northern and western blot analysis during greening of etiolated barley seedlings, and in segments of barley primary leaves grown under day/night cycles. The amount of RNA and protein increased during illumination. Maximum UROD-RNA levels were detected in the basal segments relative to the top of the leaf.

Amino Acid Sequence↗

Baboon lecithin cholesterol acyltransferase (LCAT): cDNA sequences of two alleles, evolution, and gene expression.

Lecithin cholesterol acyltransferase (LCAT) is a key enzyme of cholesterol metabolism that catalyzes esterification of cholesterol for packaging in high-density lipoprotein (HDL) particles. In this study, we cloned and sequenced LCAT cDNA from baboon, a nonhuman primate model of atherosclerosis. LCAT sequences have been highly conserved over approximately 25 million years since the divergence of the baboon and human lineages. The baboon and human sequences are 97% identical at the nucleotide (nt) level and 98% identical at the amino acid (aa) level. Only 18% of the nt substitutions change the aa sequence (nonsynonymous substitutions). The substitutions between baboon and human LCAT do not alter key functional sites including the interfacial substrate active site, asparagine-linked glycosylation sites, or sites at which rare mutations cause human familial LCAT deficiencies. We also sequenced LCAT cDNA for a less common allele that is associated with higher LCAT activities and altered lipoprotein phenotypes. There were no sequence differences between the two alleles, which suggests that genotypic effects are most likely due to allelic differences in gene expression. The tissue specificity of LCAT expression was investigated using an RNase protection assay calibrated with known amounts of synthetic human LCAT RNA. In a survey of baboon tissues, the highest levels of LCAT mRNA were found in the cerebellum and liver and trace amounts in the ileum, spleen and cerebral cortex.

Alleles↗

cDNA sequence of bovine thioredoxin.

In this paper, we report the cDNA sequence of bovine thioredoxin. We determined the full-length cDNA sequence of bovine thioredoxin by RT-PCR, 5'-RACE and 3'-RACE methods. Currently, the thioredoxin cDNA sequences of only five mammalian species (human, macaca, mouse, ovine and rat) are registered in the GenBank database. We performed sequence comparisons on the total cDNA sequence and the coding region, and produced a multialignment between the amino acid sequences of bovine and other mammalian thioredoxins. The amino acid sequences of thioredoxins are highly conserved among mammalian species, for example, only one difference exists between the amino acid sequences of bovine and ovine thioredoxin.

Animals↗

cDNA sequence of bovine thioredoxin.

In this paper, we report the cDNA sequence of bovine thioredoxin. We determined the full-length cDNA sequence of bovine thioredoxin by RT-PCR, 5'-RACE and 3'-RACE methods. Currently, the thioredoxin cDNA sequences of only five mammalian species (human, macaca, mouse, ovine and rat) are registered in the GenBank database. We performed sequence comparisons on the total cDNA sequence and the coding region, and produced a multialignment between the amino acid sequences of bovine and other mammalian thioredoxins. The amino acid sequences of thioredoxins are highly conserved among mammalian species, for example, only one difference exists between the amino acid sequences of bovine and ovine thioredoxin.

Amino Acid Sequence↗

Amino acid sequence of Coprinus macrorhizus peroxidase and cDNA sequence encoding Coprinus cinereus peroxidase. A new family of fungal peroxidases.

Sequence analysis and cDNA cloning of Coprinus peroxidase (CIP) were undertaken to expand the understanding of the relationships of structure, function and molecular genetics of the secretory heme peroxidases from fungi and plants. Amino acid sequencing of Coprinus macrorhizus peroxidase, and cDNA sequencing of Coprinus cinereus peroxidase showed that the mature proteins are identical in amino acid sequence, 343 residues in size and preceded by a 20-residue signal peptide. Their likely identity to peroxidase from Arthromyces ramosus is discussed. CIP has an 8-residue, glycine-rich N-terminal extension blocked with a pyroglutamate residue which is absent in other fungal peroxidases. The presence of pyroglutamate, formed by cyclization of glutamine, and the finding of a minor fraction of a variant form lacking the N-terminal residue, indicate that signal peptidase cleavage is followed by further enzymic processing. CIP is 40-45% identical in amino-acid sequence to 11 lignin peroxidases from four fungal species, and 42-43% identical to the two known Mn-peroxidases. Like these white-rot fungal peroxidases, CIP has an additional segment of approximately 40 residues at the C-terminus which is absent in plant peroxidases. Although CIP is much more similar to horseradish peroxidase (HRP C) in substrate specificity, specific activity and pH optimum than to white-rot fungal peroxidases, the sequences of CIP and HRP C showed only 18% identity. Hence, CIP qualifies as the first member of a new family of fungal peroxidases. The nine invariant residues present in all plant, fungal and bacterial heme peroxidases are also found in CIP. The present data support the hypothesis that only one chromosomal CIP gene exists. In contrast, a large number of secretory plant and fungal peroxidases are expressed from several peroxidase gene clusters. Analyses of three batches of CIP protein and of 49 CIP clones revealed the existence of only two highly similar alleles indicating less peroxidase polymorphism in C. cinereus strains than observed in plants and white-rot fungi.

Amino Acid Sequence↗

Amino acid translation program for full-length cDNA sequences with frameshift errors.

Here we present an amino acid translation program designed to suggest the position of experimental frameshift errors and predict amino acid sequences for full-length cDNA sequences having phred scores. Our program generates artificial insertions into artificial deletions from low-accuracy positions of the original sequence, thereby generating many candidate sequences. The validity of the most probable sequence (the likelihood that it represents the actual protein) is evaluated by using a score (V(a)) that is calculated in light of the Kozak consensus, preferred codon usage, and position of the initiation codon. To evaluate the software, we have used a database in which, out of 612 cDNA sequences, 524 (86%) carried 773 frameshift errors in the coding sequence. Our software detected and corrected 48% of the total frameshift errors in 62% of the total cDNA sequences with frameshift errors. The false positive rate of frameshift correction was 9%, and 91% of the suggested frameshifts were true.

Base Composition↗

Phytotoxic protein PcF, purification, characterization, and cDNA sequencing of a novel hydroxyproline-containing factor secreted by the strawberry pathogen Phytophthora cactorum.

A novel protein factor, named PcF, has been isolated from the culture filtrate of Phytophthora cactorum strain P381 using a highly sensitive leaf necrosis bioassay with tomato seedlings. Isolated PcF protein alone induced leaf necrosis on its host strawberry plant. The primary structure and cDNA sequence of this novel phytotoxic protein was determined, and BLAST searches of Swiss-Prot, EMBL, and GenBank(TM)/EBI data banks showed that PcF shared no significant homology with other known sequences. The 52-residue PcF protein, which contains a 4-hydroxyproline residue along with three S-S bridges, exhibits a high content of acidic sidechains, accounting for its isoelectric point of 4.4. The molecular mass of isolated PcF is 5,622 +/- 0.5 Da as determined by mass spectrometry and matches that calculated from the deduced amino acid sequence with cDNA sequencing. The cDNA sequence indicates that PcF is first produced as a larger precursor, comprising an additional N-terminal, 21-residue secretory signal peptide. Maturation of this protein involves the hydroxylation of proline 49, a feature that is unique among other known secreted fungal phytopathogenic proteins.

Amino Acid Sequence↗

Cloning of cDNA sequences of human adenosine deaminase.

Cloned cDNA sequences of human adenosine deaminase (ADA; adenosine aminohydrolase, EC 3.5.4.4) have been isolated from a cDNA library constructed in bacteriophage lambda gt10. The cDNA for the library was prepared from poly(A)+ RNA isolated from a human T-lymphoblast cell line, CCRF-CEM. The library was initially screened by differential plaque hybridization to labeled cDNA prepared from human T- and B-lymphoblast cell lines with a 21-fold difference in levels of translatable ADA mRNA. Two recombinants containing cloned cDNA sequences for ADA were identified by hybridization-selected translation. Both recombinants contained approximately 1,600 base pairs of inserted human DNA. Restriction maps of the two inserts were not identical. One contained approximately 40 base pairs of additional DNA toward the center of the cDNA. The cloned cDNA specifically hybridized to five fragments generated by HindIII digestion of human genomic DNA. It also hybridized to human lymphoblast RNA species 1.6 and 5.8 kilobases in length. The cDNA was used as a probe to estimate ADA mRNA levels in human lymphoblast cell lines. ADA mRNA levels correlate closely with levels of ADA catalytic activity and ADA protein in cell lines containing structurally normal ADA. A leukemic T-lymphoblast line produced 6 to 9 times as much ADA protein and ADA mRNA as transformed B-lymphoblast lines. Two mutant B-lymphoblast lines from patients with hereditary ADA deficiency contained unstable ADA protein but had 3 to 4 times the normal level of ADA mRNA.

Adenosine Deaminase↗

Assessing protein coding region integrity in cDNA sequencing projects.

MOTIVATION: In cDNA sequencing projects, it is vital to know whether the protein coding region of a sequence is complete, or whether errors have occurred during library construction. Here we present a linear discriminant approach that predicts this completeness by estimating the probability of each ATG being the initiation codon. RESULTS: Because of the current shortage of full-length cDNA data on which to base this work, tests were performed on a non-redundant set of 660 initiation codon-containing DNA sequences that had been conceptually spliced into mRNA/cDNA. We also used an edited set of the same sequences that only contained the region following the initiation codon as a negative control. Using the criterion that only a single prediction is allowed for each sequence, a cut-off was selected at which discrimination of both positive and negative sets was equal. At this cut-off, 67% of each set could be correctly distinguished, with the correct ATG codon also being identified in the positive set. Reliability could be increased further by raising the cut-off or including homologues, the relative merits of which are discussed. AVAILABILITY: The prediction program, called ATGpr, and other data are available at http://www.hri.co.jp/atgpr CONTACT: swintech@hri.co.jp

Base Sequence↗

Recovery of upstream cDNA sequences by a PCR-based biotin-capture method.

The world-wide, large-scale sequencing efforts have generated an abundance of partial cDNA sequences, i.e., expressed sequence tags (ESTs), accessible in the public databases. To enable functional characterization of these partial cDNA sequences, general and robust methods for recovery of upstream full-coding cDNA sequences are needed. Here, a novel biotin- and PCR-assisted capture method was used directly on poly(A)+ RNA for the purpose of generating a full-coding sequence of a gene with only partially known sequence and for which a full-length clone of the gene was not found in existing cDNA libraries. The presented method involves linear extension by reverse transciptase from a biotinylated primer annealing in a region with known sequence. After capture of the generated single-stranded cDNA onto paramagnetic beads, unspecifically annealing primers, i.e., arbitrary primers, were used to generate cDNA fragments that could be amplified by PCR and thereafter directly sequenced without subcloning. By using the presented strategy, which is to be seen as a complement to rapid amplification of cDNA ends (RACE)-related methods, we were able to recover full-coding sequence versions of two potential splice variants of the target gene. The general applicability of the novel method for recovery and sequencing of cDNA sequences is discussed.

Animals↗

Amino acid and cDNA sequences of a methionine-rich 2S protein from sunflower seed (Helianthus annuus L.).

The amino acid sequence of a methionine-rich 2S seed protein from sunflower (Helianthus annuus L.) and the sequence of a cDNA clone which codes for the entire primary translation product have been determined. The mature protein consists of a single polypeptide chain of 103 amino acids (molecular mass 12133 Da) which contains 16 residues of methionine and 8 residues of cysteine. The cDNA sequence established that the protein is synthesized as a precursor of 141 residues with a typical hydrophobic signal sequence of 25 residues followed by a further 13-residue hydrophobic pro-sequence which is presumably removed by post-translational cleavage. The sequence of the mature protein and that deduced from the cDNA were identical with no evidence of processing at the C-terminus. Comparison of the sunflower methionine-rich protein sequence with sequences of other seed 2S proteins from dicotyledons and monocotyledons showed limited but distinct sequence similarities; in particular the arrangement of the cysteine residues was conserved. The sunflower protein shows 34% identity with the methionine-rich Brazil nut 2S protein and the prepro regions of the precursors of these two proteins show about 50% identity. This similarity indicates that these methionine-rich 2S proteins have diverged as a subclass of the 2S superfamily of proteins which contain only 2-3% methionine. While the related 2S proteins from other dicotyledons are processed to a small and large subunit, the sunflower protein is not cleaved in this way.

2S Albumins, Plant↗

Identification by polymerase chain reaction and hybridization with sequence-specific oligoprobes and complete exon-2 cDNA sequencing of a new DRB1 (DRB1*1131) allele.

HLA-DRB1 is the most polymorphic gene described so far, and their encoded molecules disclose a major role in allogeneic responses. We describe in this report a new DRB1 allele in a Spanish Caucasian bone marrow donor, initially defined by PCR-SSO as a DRB1*11-like allele. Complete exon 2 cDNA-sequencing reveled that this allele was identical to DRB1*1119 except for a single substitution at position 178, which generates an amino acid change (Tyr-His) at position 60. This residue is shared by several DRB1*14 subtypes and DRB1*0808. The new allele was officially named DRB1*1131.

Alleles↗

Primary structure of a lipoxygenase from barley grain as deduced from its cDNA sequence.

A full length cDNA sequence for a barley grain lipoxygenase was obtained. It includes a 5' untranslated region of 69 nucleotides, an open reading frame of 2586 nucleotides encoding a protein of 862 amino acid residues and a 3' untranslated region of 142 nucleotides. The molecular mass of the encoded polypeptide was calculated to be 96.392. Its amino acid sequence shows a high homology with that of other plant lipoxygenases identified to date.

Amino Acid Sequence↗

Human erythrocyte membrane sialoglycoprotein beta. The cDNA sequence suggests the absence of a cleaved N-terminal signal sequence.

We have isolated cDNA clones corresponding to the human erythrocyte membrane sialoglycoprotein beta. The clones encompass the coding region for the protein, 120 residues of the 5' non-coding region and the 3' non-coding region. The cDNA sequence suggests that sialoglycoprotein beta is not translated with the cleaved N-terminal signal sequence usual in a membrane protein of this type. Sialoglycoprotein beta or a closely related homologue is present in human kidney as well as erythroid cells.

Base Sequence↗

An evolutionary perspective on glutathione transferases inferred from class-theta glutathione transferase cDNA sequences.

We report the cDNA sequence for rat glutathione transferase (GST) subunit 5, which is one of at least three class Theta subunits in this species. This sequence, when compared with that of subunit 12 recently published by Ogura, Nishiyama, Okada, Kajita, Narihata, Watabe, Hiratsuka & Watabe [(1991) Biochem. Biophys. Res. Commun. 181, 1294-1300] proves that Theta is a separate multigene class of GST with little amino acid sequence identity with Mu-, Alpha- or Pi-class enzymes. The amino acid sequence identity of class-Theta subunits is highly conserved in rat, the fruitfly Drosophila, maize (Zea mays) and Methylobacterium, which suggests that this family is representative of the ancient progenitor GST gene and originates from the endosymbioses of a purple bacterium leading to the mitochondrion. The high conservation of class Theta brings into prominence that Alpha-, Mu- and Pi-class enzymes, which are not present in plants, derive from a Theta-class gene duplication before the divergence of fungi and animals and, given the binding properties of the Alpha-, Mu- and Pi-classes, suggests a role for these in the evolution of fungi and animals.

Amino Acid Sequence↗