Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “cDNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Genomic and cDNA sequence tags of the hyperthermophilic archaeon Pyrobaculum aerophilum.

The hyperthermophilic archaeum, Pyrobaculum aerophilum, grows optimally at 100 degrees C with a doubling time of 180 min. It is a member of the phylogenetically ancient Thermoproteales order, but differs significantly from all other members by its facultatively aerobic metabolism. Due to its simple cultivation requirements and its nearly 100% plating efficiency, it was chosen as a model organism for studying the genome organization of hyperthermophilic ancient archaea. By a G+C content of the DNA of 52 mol%, sequence analysis was easily possible. At least some of the mRNA of P. aerophilum carried poly-A tails facilitating the construction of a cDNA library. 245 sequence tags of a poly-A primed cDNA library and 55 sequence tags from a 1-2 kb Sau3AI-fragment containing genomic library were analyzed and the corresponding amino acid sequences compared with protein sequences from databases. Fourteen percent of the cDNA and >9% of genomic DNA sequence tags revealed significant similarities to proteins in the databases. Matches were obtained to proteins from archaeal, bacterial and eukaryal sources. Some sequences showed greatest similarity to eukaryal rather than to bacterial versions of proteins, other matches were found to proteins which had previously only been found in eukaryotes.

Archaea↗

cDNA sequences of two inducible T-cell genes.

We have previously described a set of human T-lymphocyte-specific cDNA clones isolated by a modified differential screening procedure. Apparent full-length cDNAs containing the sequences of 14 of the 16 initial isolates were sequenced and were found to represent five different species of mRNA; three of the five species were identical to previously reported cDNA sequences of preproenkephalin, T-cell-replacing factor, and a serine esterase, respectively. The other two species, 4-1BB and L2G25B, were inducible sequences found in mRNA from both a cytolytic T-lymphocyte and a helper T-lymphocyte clone and were not previously described in T-cell mRNA; these mRNA sequences encode peptides of 256 and 92 amino acids, respectively. Both peptides contain putative leader sequences. The protein encoded by 4-1BB also has a potential membrane anchor segment and other features also seen in known receptor proteins.

Adjuvants, Immunologic↗

Analysis of cDNA sequences from mouse testis.

Few mammalian proteins involved in chromosome structure and function during meiosis have been characterized. As an approach to identify such proteins, cDNA clones expressed in mouse testis were analyzed by sequencing and Northern blotting. Various cDNA library screening methods were used to obtain the clones. First, hybridization with cDNA from testis or brain allowed selection of either negative or differentially expressed plaques. Second, positive plaques were identified by screening with polyclonal antisera to prepubertal testis nuclear proteins. Most clones were selected by negative hybridization to correspond to a low abundance class of mRNAs. A PCR-based solid-phase DNA sequencing protocol was used to rapidly obtain 306 single-pass cDNA sequences totaling more than 104 kb. Comparison with nucleic acid and protein databases showed that 56% of the clones have no significant match to any previously identified sequence. Northern blots indicate that many of these novel clones are testis-enriched in their expression. Further evidence that the screening strategies were appropriate is that a high proportion of the clones which do have a match encode testis-enriched or meiosis-specific genes, including the mouse homolog of a rat gene that encodes a synaptonemal complex protein.

Amino Acid Sequence↗

Human cytosolic asparaginyl-tRNA synthetase: cDNA sequence, functional expression in Escherichia coli and characterization as human autoantigen.

The cDNA for human cytosolic asparaginyl-tRNA synthetase (hsAsnRSc) has been cloned and sequenced. The 1874 bp cDNA contains an open reading frame encoding 548 amino acids with a predicted M r of 62 938. The protein sequence has 58 and 53% identity with the homologous enzymes from Brugia malayi and Saccharomyces cerevisiae respectively. The human enzyme was expressed in Escherichia coli as a fusion protein with an N-terminal 4 kDa calmodulin-binding peptide. A bacterial extract containing the fusion protein catalyzed the aminoacylation reaction of S.cerevisiae tRNA with [14C]asparagine at a 20-fold efficiency level above the control value confirming that this cDNA encodes a human AsnRS. The affinity chromatography purified fusion protein efficiently aminoacylated unfractionated calf liver and yeast tRNA but not E.coli tRNA, suggesting that the recombinant protein is the cytosolic AsnRS. Several human anti-synthetase sera were tested for their ability to neutralize hsAsnRSc activity. A human autoimmune serum (anti-KS) neutralized hsAsnRSc activity and this reaction was confirmed by western blot analysis. The human asparaginyl-tRNA synthetase appears to be like the alanyl- and histidyl-tRNA synthetases another example of a human Class II aminoacyl-tRNA synthetase involved in autoimmune reactions.

Acylation↗

Opsin cDNA sequences of a UV and green rhodopsin of the satyrine butterfly Bicyclus anynana.

The cDNAs of an ultraviolet (UV) and long-wavelength (LW) (green) absorbing rhodopsin of the bush brown Bicyclus anynana were partially identified. The UV sequence, encoding 377 amino acids, is 76-79% identical to the UV sequences of the papilionids Papilio glaucus and Papilio xuthus and the moth Manduca sexta. A dendrogram derived from aligning the amino acid sequences reveals an equidistant position of Bicyclus between Papilio and Manduca. The sequence of the green opsin cDNA fragment, which encodes 242 amino acids, represents six of the seven transmembrane regions. At the amino acid level, this fragment is more than 80% identical to the corresponding LW opsin sequences of Dryas, Heliconius, Papilio (rhodopsin 2) and Manduca. Whereas three LW absorbing rhodopsins were identified in the papilionid butterflies, only one green opsin was found in B. anynana.

Amino Acid Sequence↗

Sequence variations in the envelope protein of the hepatitis C virus: comparison with partial cDNA sequence of a new variant virus obtained by the polymerase chain reaction.

It has been reported that the envelope region located at the 3' portion of the structural protein coding region is one of the most variable regions at both nucleotide and amino acid sequence levels in the hepatitis C virus (HCV) genome. We cloned HCV cDNA fragments of an envelope protein coding region (HCVNK), which were derived from serum of a Japanese patient with hepatocellular carcinoma and were amplified by polymerase chain reaction. After determining the nucleotide sequence, deduced amino acid sequence of the envelope protein region was compared with those of six HCV strains already published (HCJ1, HCVUS, HCJ4, HCVJH, HCVJ and HCVBK). Homology analysis among the strains revealed that the seven strains were classified into two subtypes; a US subtype (HCJ1 and HCVUS) and a Japanese subtype (HCJ4, HCVJH, HCVJ, HCVBK and HCVNK), since percentage homologies between two subtypes (70.3-77.3%) were significantly lower than those within each subtype (83.9-93.5%). Detailed analysis of the amino acid sequences also indicates that the region at aa246-aa258, tentatively named intersubtype variable region-1, may distinguish the US subtype from the Japanese subtype.

Amino Acid Sequence↗

cDNA sequence and chromosomal localization of human enterokinase, the proteolytic activator of trypsinogen.

Enterokinase is a serine protease of the duodenal brush border membrane that cleaves trypsinogen and produces active trypsin, thereby leading to the activation of many pancreatic digestive enzymes. Overlapping cDNA clones that encode the complete human enterokinase amino acid sequence were isolated from a human intestine cDNA library. Starting from the first ATG codon, the composite 3696 nt cDNA sequence contains an open reading frame of 3057 nt that encodes a 784 amino acid heavy chain followed by a 235 amino acid light chain; the two chains are linked by at least one disulfide bond. The heavy chain contains a potential N-terminal myristoylation site, a potential signal anchor sequence near the amino terminus, and six structural motifs that are found in otherwise unrelated proteins. These domains resemble motifs of the LDL receptor (two copies), complement component Clr (two copies), the metalloprotease meprin (one copy), and the macrophage scavenger receptor (one copy). The enterokinase light chain is homologous to the trypsin-like serine proteinases. These structural features are conserved among human, bovine, and porcine enterokinase. By Northern blotting, a 4.4 kb enterokinase mRNA was detected only in small intestine. The enterokinase gene was localized to human chromosome 21q21 by fluorescence in situ hybridization.

Amino Acid Sequence↗

The major form of the murine asialoglycoprotein receptor: cDNA sequence and expression in liver, testis and epididymis.

Northern blot analysis of poly(A)+RNAs isolated from mouse liver or mouse testis (Te)/epididymis (Ep) reveals that both tissues express 1.5- and 7.5-kb transcripts which have extensive homology to the major form of the rat asialoglyco-protein receptor (ASGP-R). In situ hybridization studies have localized the expression of this ASGP-R-like transcript to late-stage sperm from Te and Ep of several different strains of mice. Swiss Webster mice express this ASGP-R-like transcript in late-stage spermatids at the time of release into the seminiferous tubule and in Ep sperm, while Balb/C, NIH Swiss and C57Bl/6 mice express this ASGP-R-like transcript predominantly in Ep sperm. cDNAs containing the entire coding region for this ASGP-R-like transcript have been cloned from mouse liver and mouse Te/Ep. These cDNAs are 100% identical in the coding region and 3'-untranslated region (UTR), but differ in the 5'-UTR. The gene encoding these cDNAs is called MHL-1, designating the major form of the mouse ASGP-R. The deduced amino acid (aa) sequence of MHL-1 shares 88% homology to the rat hepatic (He) lectin form 1 (RHL-1) and 78% homology to the human asialoglycoprotein receptor form 1 (H1). The three sites for N-linked glycosylation in the RHL-1 sequence are all conserved in the deduced MHL-1 sequence. Taken collectively, these data describe the cloning and sequencing of the MHL-1 cDNA and illustrate its deduced aa homology to RHL-1 and H1.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Insect chitin synthase cDNA sequence, gene organization and expression.

Chitin is a major component of the cuticle of arthropods. However, the synthesis of chitin is poorly understood. Feeding larvae of the insect Lucilia cuprina on the fungal chitin synthase competitive inhibitor, nikkomycin Z resulted in strong concentration-dependent mortality of the larvae (LD50 = 280 nM). This result demonstrates that chitin is an essential component of this insect. The complete cDNA and deduced amino-acid sequences of the first arthropod chitin synthase-like protein, LcCS-1, from the larvae of the insect L. cuprina have been determined. The cDNA sequence is 5757 bp in length and codes for a large complex protein containing 1592 amino acids (Mr = 180 717). Analysis of the whole protein sequence reveals low, but significant, similarity to yeast chitin synthases with stronger areas of conservation centred on local regions implicated in the active sites of the yeast enzymes. Strikingly, LcCS-1 contains 15-18 potential transmembrane segments, indicating that the protein is an integral membrane protein. Two alternative topographical models of LcCS-1 are described, which involve its association with either the plasma membrane or the membrane of intracellular vesicles. LcCS-1 mRNA is produced in all life stages of the insect with expression in the larval stage limited to the integument and trachea. In a third instar larva the mRNA was localized to a single layer of epidermal cells immediately underlying the procuticle region of the integument. cDNA or genomic sequences that are highly related to fragments of LcCS-1 were demonstrated in three insect orders, one arachnid and Caenorhabditis elegans, thereby attesting to the importance of this enzyme in these chitin-producing organisms. Bioinformatics has been used to deduce the gene sequence and organization of the highly homologous Drosophila melanogaster orthologue of LcCS-1, DmCS-1.

Amino Acid Sequence↗

cDNA sequence and genomic structure of the murine p55 (Mpp1) gene.

MPP1 is an X-linked human gene encoding a heavily palmitoylated membrane protein (p55) with homology to the Drosophila tumor suppressor gene lethal(1) discs-large. As a first step toward studying the effects of mutations in this gene in a mammalian system, the nucleotide sequence of the mouse Mpp1 cDNA has been determined along with the intron-exon boundaries. Mpp1 is ubiquitously expressed and encodes a p55 protein of 466 amino acids with 93 and 65% identity to the human and puffer fish (Fugu rubripes) p55 sequences, respectively. The genomic structure of the Mpp1 gene is likewise conserved with 12 exons. The location of the Mpp1 gene, on the X chromosome, is also conserved between the human and the mouse. Conservation of the Mpp1 gene between mouse and human gives support to the notion that construction and study of a mouse knockout model may help establish the function of the human MPP1 gene, a potential tumor suppressor gene.

Amino Acid Sequence↗

Rubella virus cDNA. Sequence and expression of E1 envelope protein.

A cDNA clone encoding the entire E1 envelope protein (410 amino acid residues) and a portion of the C-terminal end of the E2 envelope protein of the rubella virus has been isolated and characterized. DNA sequence analysis has revealed a region 20 nucleotides in length at the 3' end of the cloned cDNA which may be a replicase recognition site or a recognition site for encapsidation. The proteolytic cleavage site between the E1 and E2 proteins was localized based on the known amino-terminal sequence of the isolated E1 protein (Kalkkinen, N., Oker-Blom, C., and Pettersson, R. F. (1984) J. Gen. Virol. 65, 1549-1557) and the deduced amino acid sequence. The mature E1 protein is preceded by a set of 20 highly hydrophobic amino acid residues possessing characteristics of a signal peptide. This "signal peptide" is flanked on both sides by typical protease cleavage sites for trypsin-like enzyme and signal peptidase. The presence of a leader sequence in the E1 protein precursor may facilitate its translocation through the host cell membrane. The E1 protein of rubella virus shows no significant homology with alphavirus E1 envelope proteins. However, a stretch of 39 amino acids in the E1 protein of rubella virus (residues 262-300) was found to share a significant homology with the first 39 residues of bovine sperm histone. The position of 4 half-cystines and 8 arginines overlaps. The E1 protein of rubella virus has been successfully expressed in COS cells after transfecting them with rubella virus cDNA in simian virus 40-derived expression vector. This protein is antigenically similar to the one expressed by cells infected with rubella virus.

Amino Acid Sequence↗

cDNA sequence and molecular modeling of a nerve growth factor from Bothrops jararacussu venomous gland.

The complete nucleotide sequence of a nerve growth factor precursor from Bothrops jararacussu snake (Bj-NGF) was determined by DNA sequencing of a clone from cDNA library prepared from the poly(A) + RNA of the venom gland of B. jararacussu. cDNA encoding Bj-NGF precursor contained 723 bp in length, which encoded a prepro-NGF molecule with 241 amino acid residues. The mature Bj-NGF molecule was composed of 118 amino acid residues with theoretical pI and molecular weight of 8.31 and 13,537, respectively. Its amino acid sequence showed 97%, 96%, 93%, 86%, 78%, 74%, 76%, 76% and 55% sequential similarities with NGFs from Crotalus durissus terrificus, Agkistrodon halys pallas, Daboia (Vipera) russelli russelli, Bungarus multicinctus, Naja sp., mouse, human, bovine and cat, respectively. Phylogenetic analyses based on the amino acid sequences of 15 NGFs separate the Elapidae family (Naja and Bungarus) from those Crotalidae snakes (Bothrops, Crotalus and Agkistrodon). The three-dimensional structure of mature Bj-NGF was modeled based on the crystal structure of the human NGF. The model reveals that the core of NGF, formed by a pair of beta-sheets, is highly conserved and the major mutations are both at the three beta-hairpin loops and at the reverse turn.

Amino Acid Sequence↗

Two types of new ferritin cDNA sequences from Xenopus laevis germinal vesicle oocytes.

Some of Xenopus ferritin cDNA family genes have already been sequenced. In this study, we report that two ferritin cDNA genes have been cloned from the Xenopus laevis germinal vesicle (GV) oocytes. The deduced proteins have different lengths with varied sequences when compared with the published Xenopus ferritins. One of them is the ferritin light chain homologous (LCH), which is reported for the first time in Xenopus and the other is the ferritin heavy chain homologous (HCH) that is first reported in Xenopus GV oocyte.

Animals↗

Partial characterization of human complement factor H by protein and cDNA sequencing: homology with other complement and non-complement proteins.

Factor H, a control protein of the human complement system, is closely related in functional activity to two other complement control proteins, C4b-binding protein (C4bp) and complement receptor type 1 (CR1). C4bp is known to have an unusual primary structure consisting of eight homologous units each about 60 amino acids long. Such units also occur in the N-terminal regions of the complement proteins C2 and factor B, and in the non-complement serum glycoprotein beta 2I. Amino acid sequencing, and sequencing of a factor H cDNA clone, show that factor H also contains internal repeating units, and is homologous to the proteins listed above.

Amino Acid Sequence↗

DNA polymorphism detector: an automated tool that searches for allelic matches in public databases for discrepancies found in clone or cDNA sequences.

SUMMARY: DNA polymorphism detector (DPD) is a new web application developed to help automate the process of cDNA clone validation. DPD identifies and highlights discrepancies between any cDNA clone sequence and its expected reference sequence. To determine if these differences correspond to natural genetic polymorphisms (versus artifacts introduced during clone production or evaluation), DPD uses the discrepancies, along with flanking sequences, to search GenBank for identical matching strings. If matching DNA sequences are found, DPD verifies that they are from the same gene. The application then reports the discrepancy as a polymorphism along with the corresponding GenBank reference information. AVAILABILITY: DPD is currently hosted by the Harvard Institute of Proteomics at http://www.hip.harvard.edu

Cloning, Molecular↗

Complete cDNA sequence and tissue localization of N-RAP, a novel nebulin-related protein of striated muscle.

We have cloned and sequenced the full-length cDNA of N-RAP, a novel nebulin-related protein, from mouse skeletal muscle. The N-RAP message is specifically expressed in skeletal and cardiac muscle, but is not detected by Northern blot in non-muscle tissues. The full-length N-RAP cDNA contains an open reading frame of 3,525 base pairs which is predicted to encode a protein of 133 kDa. A 587 amino acid region near the C-terminus is 45% identical to the actin binding region of human nebulin, containing more than 2 complete 245 residue nebulin super repeats. The N-terminus contains the consensus sequence of a cysteine-rich LIM domain, which may function in mediating protein-protein interactions. These data suggest that the encoded protein may link actin filaments to some other proteins or structure. We expressed full-length N-RAP in Escherichia coli, as well as the nebulin-like super repeat region of N-RAP (N-RAP-SR) and the region between the LIM domain and N-RAP-SR (N-RAP-IB). An anti-N-RAP antibody raised against a 30 amino acid peptide corresponding to sequence from N-RAP-IB detected recombinant N-RAP and N-RAP-IB, but failed to detect N-RAP-SR. This antibody specifically identified a 185 kDa band as N-RAP on immunoblots of mouse skeletal and cardiac muscle proteins. In an assay of actin binding to electrophoresed and blotted proteins, we detected significant actin binding to expressed nebulin super repeats and N-RAP-SR, but only a trace amount of binding to N-RAP-IB. In immunofluorescence experiments, N-RAP was found to be localized at the myotendinous junction in mouse skeletal muscle and at the intercalated disc in cardiac muscle. Based on its domain organization, actin binding properties, and tissue localization, we propose that N-RAP plays a role in anchoring the terminal actin filaments in the myofibril to the membrane and may be important in transmitting tension from the myofibrils to the extracellular matrix.

Actins↗

cDNA sequence and deduced primary structure of an alpha-amylase inhibitor from a bruchid-resistant wild common bean.

alpha-Amylase inhibitor-2 (alpha AI-2), a seed storage protein present in a bruchid-resistant wild common bean (Phaseolus vulgaris), inhibits the growth of bruchid pests. The authors isolated and determined the sequence of an 852 nucleotide cDNA, designated as alpha ai2, and found it to contain a 720 base open reading frame (ORF). This ORF encodes a 240 amino-acid alpha AI-2 polypeptide 75.8% identical with alpha-amylase inhibitor-1 (alpha AI-1) and 50.6-55.6% with arcelin-1, phytohemagglutinin (PHA)-L and PHA-E of common bean. The high degree of sequence homology suggests that there is an evolutionary relationship among these genes.

Amino Acid Sequence↗

Structure of mouse protein S as determined by PCR amplification and DNA sequencing of cDNA.

The cDNA sequence of mouse protein S was derived by conventional PCR amplification from liver mRNA, initially using primers derived from the human cDNA sequence, followed by direct DNA sequencing. Seven overlapping PCR fragments covering all of the mature protein, part of the propeptide, and the 3' noncoding region were generated and sequenced. In some cases primers based upon the human cDNA sequence were ineffective. Subsequent successful amplification with mouse-derived primers to the same regions and comparison of the mouse and human sequences in these regions suggest that the failure of the human primers was due to insufficient degree of heterospecies identity. The mouse protein S cDNA sequence of the coding region shares 82% identity to human. The 3' noncoding region of mouse protein S cDNA has several small deletions and insertions compared to human protein S cDNA. Mature mouse protein S consists of 634 amino acids in a single polypeptide chain and displays domain organization similar to that for other species. The amino acid sequence of mouse protein S is about 80% identical to that of other species. Eleven glutamic acid residues were found in the amino terminal region and are predicted to be sites of gamma-carboxylation. Amino acid residues #80-244 are defined as four cysteine-rich repeat sequences homologous to epidermal growth factor. The remainder of the molecule is homologous to plasma sex steroid binding protein. The mouse protein S contains two potential N-glycosylation sites at positions #458 and 468 and is lacking the putative glycosylation site at #490 found in human protein S.

Amino Acid Sequence↗