Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “third generation sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

The FLP recombinase of the 2 micron circle DNA of yeast: interaction with its target sequences.

We have studied the interaction of purified FLP protein with restriction fragments from the substrate 2mu circle DNA of yeast. We find that FLP protects about 50 bp of DNA from nonspecific nuclease digestion. The protected site consists of two 13 bp inverted repeat sequences separated by an 8 bp spacer region. A third 13 bp element is also protected by binding of the FLP protein. We demonstrate that FLP introduces single- and double-strand breaks into the substrate DNA. This site-specific cleavage occurs at the margins of the spacer region, generating 8 bp 5' protruding ends with 5'-OH and 3'-protein-bound termini. Binding to mutant sites and half-sites demonstrates that the third symmetry element is not important for binding and cleavage by the FLP protein. The integrity of the core region is important for the cleavage activity of FLP.

Base Composition↗

Posttranslational processing and differential glycosylation of Tractin, an Ig-superfamily member involved in regulation of axonal outgrowth.

Tractin is a novel member of the Ig-superfamily which has a highly unusual structure. It contains six Ig domains, four FNIII-like domains, an acidic domain, 12 repeats of a novel proline- and glycine-rich motif with sequence similarity to collagen, a transmembrane domain, and an intracellular tail with an ankyrin and a PDZ domain binding motif. By generating domain-specific antibodies, we show that Tractin is proteolytically processed at two cleavage sites, one located in the third FNIII domain, and a second located just proximal to the transmembrane domain resulting in the formation of four fragments. The most NH(2)-terminal fragment which is glycosylated with the Lan3-2, Lan4-2, and Laz2-369 glycoepitopes is secreted, and we present evidence which supports a model in which the remaining fragments combine to form a secreted homodimer as well as a transmembrane heterodimer. The extracellular domain of the dimers is mostly made up of the collagen-like PG/YG-repeat domain but also contains 11/2 FNIII domain and the acidic domain. The collagen-like PG/YG-repeat domain could be selectively digested by collagenase and we show by yeast two-hybrid analysis that the intracellular domain of Tractin can interact with ankyrin. Thus, the transmembrane heterodimer of Tractin constitutes a novel protein domain configuration where sequence that has properties similar to that of extracellular matrix molecules is directly linked to the cytoskeleton through interactions with ankyrin.

Animals↗

Sequence survey of the genome of the opportunistic microsporidian pathogen, Vittaforma corneae.

The microsporidian Vittaforma corneae has been reported as a pathogen of the human stratum corneum, where it can cause keratitis, and is associated with systemic infections. In addition to this direct role as an infectious, etiologic agent of human disease, V. corneae has been used as a model organism for another microsporidian, Enterocytozoon bieneusi, a frequent and problematic pathogen of HIV-infected patients that, unlike V. corneae, is difficult to maintain and to study in vitro. Unfortunately, few molecular sequences are available for V. corneae. In this study, seventy-four genome survey sequences (GSS) were obtained from genomic DNA of spores of laboratory-cultured V. corneae. Approximately, 41 discontinuous kilobases of V. corneae were cloned and sequenced to generate these GSS. Putative identities were assigned to 44 of the V. corneae GSS based on BLASTX searches, representing 21 discrete proteins. Of these 21 deduced V. corneae proteins, only two had been reported previously from other microsporidia (until the recent report of the Encephalitozoon cuniculi genome). Two of the V. corneae proteins were of particular interest, reverse transcriptase and topoisomerase IV (parC). Since the existence of transposable elements in microsporidia is controversial, the presence of reverse transcriptase in V. corneae will contribute to resolution of this debate. The presence of topoisomerase IV was remarkable because this enzyme previously had been identified only from prokaryotes. The 74 GSS included 26.7 kilobases of unique sequences from which two statistics were generated: GC content and codon usage. The GC content of the unique GSS was 42%, lower than that of another microsporidian, E. cuniculi (48% for protein-encoding regions), and substantially higher than that predicted for a third microsporidian, Spraguea lophii (28%). A comparison using the Pearson correlation coefficient showed that codon usage in V. corneae was similar to that in the yeasts, Saccharomyces cerevisiae (r = 0.79) and Shizosaccharomyces pombe (r = 0.70), but was markedly dissimilar to E. cuniculi (r = 0.19).

Amino Acid Sequence↗

Direct Fmoc/tert-Bu solid phase synthesis of octamannosyl polylysine dendrimer-peptide conjugates.

The mannose binding proteins on the surface of the dendritic cells are responsible for capture of pathogens in the early stages of immune response. Conjugation to mannose dendrimers is a rarely explored but potentially powerful strategy for enhancing immunogenicity of synthetic peptides relying on direct delivery to dendritic cells. We describe a general protocol for preparation of pure, monodisperse third-generation mannosylated poly-L-lysine dendrimer-peptide conjugates using direct, machine-assisted Fmoc/t-Bu solid phase peptide synthesis. The glycodendrons were elaborated onto the N- or C-terminus of sequences derived from HIV-1 gp41, SARS-CoV S2 protein, and Influenza Hemagglutinin (consisting of 15-44 residues). The products were obtained in a homogeneous state after cleavage from the resin, deprotection, and a single purification on semipreparative RP-HPLC.

Animals↗

Polymerase chain reaction based assessment of leukoreduction efficacy using a cytomegalovirus DNA transfected human T-cell line.

BACKGROUND AND OBJECTIVES: Detection of CMV DNA by PCR in seropositive blood donors is affected by nucleic acid extraction, amplification conditions and PCR product detection sensitivity. Clinical studies have shown that leukoreduced blood products are as effective as CMV-seronegative blood products in minimizing transfusion-associated CMV infection. We developed a PCR-based test model to assess the efficacy of leukoreduction on the removal of CMV DNA from blood donations. MATERIALS AND METHODS: Whole blood units were spiked with a human Jurkat T-lymphocyte cell line which had been stably transfected with a CMV DNA immediate early sequence. All blood units were either CMV seronegative or shown to be CMV PCR negative. The amount of CMV DNA by PCR before, during and after filtration of blood units with third-generation leukocyte reduction filters were determined using a semi-quantitative adaptation of the Digene SHARP Signal System Assay (DSSSA). RESULTS: Whole blood units spiked with 1-2 x 10(6) CMV transfected Jurkat cells contained no detectable CMV DNA after leukoreduction. With a higher inoculum of 2.9 x 10(8) transfected Jurkat cells per unit, CMV DNA was detected after leukoreduction; however, there was an approximate 3 log decrease in the amount of CMV DNA detected. CONCLUSION: This CMV DNA transfected human Jurkat cell line in conjunction with semi-quantitative CMV PCR may be used to model leukoreduction efficacy and provides evidence that leukoreduction of blood products can decrease the amount of cellular CMV DNA by approximately 3 logs.

Cytomegalovirus↗

Intron 4 mutation in APC gene results in splice defect and attenuated FAP phenotype.

The adenomatous polyposis coli (APC) protein is a tumor suppressor frequently involved in the development of inherited and sporadic colon cancers. Somatic mutations of the APC gene are found in 80% of all colon cancers. Inherited mutations result in familial adenomatous polyposis (FAP) as well as an attenuated form of this syndrome. FAP is characterized by the early age onset of hundreds to thousands of colonic adenomatous polyps and a virtual certainty of colon cancer unless the colon is removed. The attenuated form of FAP (AFAP) is characterized by fewer adenomas, later onset of adenomas and cancer, and a decreased lifetime cancer risk. We report a 37-year-old man with a history of more than 50 colonic adenomatous polyps, located predominately in the right colon. An insertion of a single thymidine between the second and third base pairs of intron 4 of the APC gene was identified (c.531+2_531+3insT). Monoallelic hybrid cells harboring a single copy of human chromosome 5 were generated from patient lymphoblasts. Sequencing of the APC cDNA product from these cells revealed a single RNA transcript with aberrant splicing in the mutant mRNA whereby exon 4 is deleted. The translational reading frame is shifted after codon 140 and a translational stop is generated predicting a truncated protein of 147 amino acids, thus indicating that the intronic mutation is disease causing. The lack of a secondary transcript from the mutant allele suggests that incomplete exon skipping is not the molecular mechanism behind the attenuated phenotype.

Adenomatous Polyposis Coli↗

Analysis of the herpes simplex virus type 1 OriS sequence: mapping of functional domains.

The herpes simplex virus type 1 (HSV-1) OriS region resides within a 90-bp sequence that contains two binding sites for the origin-binding protein (OBP), designated sites I and II. A third presumptive OBP-binding site (III) within OriS has strong sequence similarity to sites I and II, but no sequence-specific OBP binding has yet been demonstrated at this site. We have generated mutations in sites I, II, and III and determined their replication efficiencies in a transient in vivo assay in the presence of a helper virus. Mutations in any one of the sites reduced DNA replication significantly. To study the role of OriS sequence elements in site I and the presumptive site III in DNA replication, we have also generated a series of mutations that span from site I across the presumptive binding site III. These mutants were tested for their ability to replicate and for the ability to bind OBP by using gel shift analyses. The results indicate that mutations across site I drastically reduce DNA replication. Triple-base-pair substitution mutations that fall within the crucial OBP-binding domain, 5'-YGYTCGCACT-3' (where Y represents C or T), show a reduced level of OBP binding and DNA replication. Substitution mutations in site I that are outside this crucial binding sequence show a more detrimental effect on DNA replication than on OBP binding. This suggests that these sequences are required for initiation of DNA replication but are not critical for OBP binding. Mutations across the presumptive OBP-binding site III also resulted in a loss in efficiency of DNA replication. These mutations influenced OBP binding to OriS in gel shift assays, even though the mutated sequences are not contained within known OBP-binding sites. Replacement of the wild-type site III with a perfect OBP-binding site I results in a drastic reduction of DNA replication. Thus, our DNA replication assays and in vitro DNA-binding studies suggest that the binding of the origin sequence by OBP is not the only determining factor for initiation of DNA replication in vivo.

Animals↗

Immune response to ferredoxin; nonresponder status is not due to suppression.

Ferredoxin (Fd), a small protein from Clostridium pasteurianum, has been selected for immunologic studies because of its limited number (two) of antigenic determinants. Functionally (as determined by antibody binding), monodeterminant fragments of Fd can be generated enzymatically, leaving molecules only a few amino acids smaller than the native protein, with unaltered solid phase binding properties. These fragments were used to assess the immune response to each of the two determinants. Clear differences in immunologic properties can be assigned to sequences within Fd: the amino terminal tripeptide is responsible for inducing a proliferative response and limited antibody production, whereas the carboxy terminal dipeptide accounts for most of the antibody activity, yet little, if any, T-proliferative activity. Studies with the enzyme-generated fragments of Fd have unmasked a sequence proximal to the amino terminal that represents a second determinant for T cell proliferation but does not have any demonstrable antibody-inducing activity. This third determinant is shown to induce responsiveness to Fd in nonresponder animals after the removal of the amino terminal tripeptide. The results indicate that nonresponsiveness to this molecule in H-2d mice is not a direct effect of suppression.

Animals↗

An immunoglobulin heavy chain variable region gene is generated from three segments of DNA: VH, D and JH.

We have determined the sequences of separate germline genetic elements which encode two parts of a mouse immunglobulin heavy chain variable region. These elements, termed gene segments, are heavy chain counterparts of the variable (V) and joining (J) gene segments of immunoglobulin light chains. The VH gene segment encodes amino acids 1-101 and the JH gene segment encodes amino acids 107-123 of the S107 phosphorylcholine-binding VH region. This JH gene segment and two other JH gene segments are located 5' to the mu constant region gene (Cmu) in germline DNA. We have also determined the sequence of a rearranged VH gene encoding a complete VH region, M603, which is closely related to S107. In addition, we have partially determined the VH coding sequences of the S107 and M167 heavy chain mRNAs. By comparing these sequences to the germline gene segments, we conclude that the germline VH and JH gene segments do not contain at least 13 nucleotides which are present in the rearranged VH genes. In S107, these nucleotides encode amino acids 102-106, which form part of the third hypervariable region and consequently influence the antigen-binding specificity of the immunoglobulin molecule. This portion of the variable region may be encoded by a separate germline gene segment which can be joined to the VH and JH gene segments. We term this postulated genetic element the D gene segment, referring to its role in the generation of heavy chain diversity. Essentially the same noncoding sequences are found 3' to the VH gene segment and as inverse complements 5' to two JH gene segments. These are the same conserved nucleotides previously found adjacent to light chain V and J gene segments. Each conserved sequence consists of blocks of seven and ten conserved nucleotides which are separated by a spacer of either 11 or 22 nonconserved nucleotides. The highly conserved spacing, corresponding to one or two turns of the DNA helix, maintains precise spatial orientations between blocks of conserved nucleotides. Gene segments which can join to one another (VK and JK, for example) always have spacers of different lengths. Based on these observations, we propose a model for variable region gene rearrangement mediated by proteins which recognize the same conserved sequences adjacent to both light and heavy chain immunoglobulin gene segments.

Animals↗

RSDB: representative protein sequence databases have high information content.

MOTIVATION: Biological sequence databases are highly redundant for two main reasons: 1. various databanks keep redundant sequences with many identical and nearly identical sequences 2. natural sequences often have high sequence identities due to gene duplication. We wanted to know how many sequences can be removed before the databases start losing homology information. Can a database of sequences with mutual sequence identity of 50% or less provide us with the same amount of biological information as the original full database? RESULTS: Comparisons of nine representative sequence databases (RSDB) derived from full protein databanks showed that the information content of sequence databases is not linearly proportional to its size. An RSDB reduced to mutual sequence identity of around 50% (RSDB50) was equivalent to the original full database in terms of the effectiveness of homology searching. It was a third of the full database size which resulted in a six times faster iterative profile searching. The RSDBs are produced at different granularity for efficient homology searching. AVAILABILITY: All the RSDB files generated and the full analysis results are available through internet: ftp://ftp.ebi.ac. uk/pub/contrib/jong/RSDB/http://cyrah.e bi.ac.uk:1111/Proj/Bio/RSDB

Algorithms↗

The adipsin-acylation stimulating protein system and regulation of intracellular triglyceride synthesis.

We have previously characterized an activity from human plasma that markedly stimulates triglyceride synthesis in cultured human skin fibroblasts and human adipocytes. Based on its in vitro activity we named the active component acylation stimulating protein (ASP). The molecular identity of the active serum component has now been determined. NH2-terminal sequence analysis, ion spray ionization mass spectroscopy, and amino acid composition analysis all indicate that the active purified protein is a fragment of the third component of plasma complement, C3a-desArg. As well, reconstitution experiments with complement factors B, D, and complement C3, the components necessary to generate C3a, have confirmed the identity of ASP as C3a. ASP appears to be the final effector molecule generated by a novel regulatory system that modulates the rate of triglyceride synthesis in adipocytes.

Amino Acid Sequence↗

Shedding of somatic angiotensin-converting enzyme (ACE) is inefficient compared with testis ACE despite cleavage at identical stalk sites.

The somatic and testis isoforms of angiotensin-converting enzyme (ACE) are both C-terminally anchored ectoproteins that are shed by an unidentified secretase. Although testis and somatic ACE both share the same stalk and membrane domains the latter was reported to be shed inefficiently compared with testis ACE, and this was ascribed to cleavage at an alternative site [Beldent, Michaud, Bonnefoy, Chauvet and Corvol (1995) J. Biol. Chem. 270, 28962-28969]. These differences constitute a useful model system of the regulation and substrate preferences of the ACE secretase, and hence we investigated this further. In transfected Chinese hamster ovary cells, human somatic ACE (hsACE) was indeed shed less efficiently than human testis ACE, and shedding of somatic ACE responded poorly to phorbol ester activation. However, using several analytical techniques, we found no evidence that the somatic ACE cleavage site differed from that characterized in testis ACE. First, anti-peptide antibodies raised to specific sequences on either side of the reported cleavage site (Arg(1137)/Leu(1138)) clearly recognized soluble porcine somatic ACE, indicating that cleavage was C-terminal to Arg(1137). Second, a competitive ELISA gave superimposable curves for porcine plasma ACE, secretase-cleaved porcine somatic ACE (eACE), and trypsin-cleaved ACE, suggesting similar C-terminal sequences. Third, mass-spectral analyses of digests of released soluble hsACE or of eACE enabled precise assignments of the C-termini, in each case to Arg(1203). These data indicated that soluble human and porcine somatic ACE, whether generated in vivo or in vitro, have C-termini consistent with cleavage at a single site, the Arg(1203)/Ser(1204) bond, identical with the Arg(627)/Ser(628) site in testis ACE. In conclusion, the inefficient release of somatic ACE is not due to cleavage at an alternative stalk site, but instead supports the hypothesis that the testis ACE ectodomain contains a motif that activates shedding, which is occluded by the additional domain found in somatic ACE.

Amino Acid Sequence↗

Structural evidence for independent joining region gene in immunoglobulin heavy chains from anti-galactan myeloma proteins and its potential role in generating diversity in complementarity-determining regions.

We have determined the variable region sequences of four heavy chains from beta(1-6)D-galactan-binding myeloma proteins. Two of these proteins are identical to position 100 which is located in the third complementarity-determining region (CDR-3). The remaining two differ at a total of 8 positions over the first 100 amino acids, and all of the differences can be explained by single-base mutations at the DNA level. When an assessment is made of the protein segment following CDR-3, which has been termed "J segment" or "FR4," a completely different pattern of variation is observed. The J segments from the four proteins can be divided into two sets. Members of each set share a series of linked amino acids not found in members of the alternative set. The two proteins identical to position 100 have J segments from the two different sets, suggesting that recombination has occurred between V and J genes. An examination of the CDR-3 sequences from the four heavy chains reveals substitutions at positions 100 and 105. Gly is found at 100 in two of the proteins and His in the remaining two. In the two proteins with Gly-100, the following J sequence is limited to one of the two sets of J segments defined by linked amino acids. Similarly, the two heavy chains with His-100 have J segments from the second set. Thus, at the protein level an apparent association is seen between CDR-3 and J segment. If CDR-3 should be found linked to J segment at the DNA level, a new mechanism would be introduced for increasing antibody diversity by recombining various CDR-3 plus J genes with genes coding for the remainder of the variable region. Alternatively, if CDR-3 were coded for by the V gene, then the recombination of V with J may provide an opportunity to introduce mutations in CDR-3. In this case the linkage of amino acids in CDR-3 and the J segments would suggest that recognition signals are used such that certain V genes only pair with a given J gene.

Amino Acid Sequence↗

Transposition in Shigella dysenteriae: isolation and analysis of IS911, a new member of the IS3 group of insertion sequences.

Twenty-nine clear-plaque mutants of bacteriophage lambda were isolated from a Shigella dysenteriae lysogen. Three were associated with insertions in the cI gene: two were due to insertion of IS600, and the third resulted from insertion of a new element, IS911. IS911 is 1,250 base pairs (bp) long, carries 27-bp imperfect terminal inverted repeats, and generates 3-bp duplications of the target DNA on insertion. It was found in various copy numbers in all four species of Shigella tested and in Escherichia coli K-12 but not in E. coli W. Analysis of IS911-mediated cointegrate molecules indicated that the majority were generated without duplication of IS911. They appeared to result from direct insertion via one end of the element and the neighboring region of DNA, which resembles a terminal inverted repeat of IS911. Nucleotide sequence analysis revealed that IS911 carries two consecutive open reading frames which code for potential proteins showing similarities to those of the IS3 group of elements.

Amino Acid Sequence↗

Identification and characterization of two alternative splice variants of human interleukin-2.

Our previous work showed that alternative splicing is used to make an inhibitory variant of human interleukin (IL)-4. Because of homology between IL-4 and IL-2 proteins and receptors, we tested whether alternative splicing is used to generate similar inhibitory variants of human IL-2. Messenger RNA from peripheral blood mononuclear cells was subjected to reverse transcription-polymerase chain reaction using IL-2 exon 1- and exon 4-specific primers. Two amplification products, named IL-2delta2 and IL-2delta3, were found in addition to the native IL-2 product. The IL-2delta2 cDNA sequence was identical to IL-2 cDNA throughout the entire coding region, except exon 2 was omitted by alternative splicing. In IL-2delta3 cDNA, the third exon of IL-2 was omitted by alternative splicing. Unlike IL-2, IL-2delta2 and IL-2delta3 did not stimulate T cell proliferation. However, both inhibited IL-2 costimulation of T cell proliferation, and both inhibited cellular binding of rhIL-2 to high affinity IL-2 receptors. Thus, IL-2 is the second cytokine that uses alternative splicing to generate variants that are competitive inhibitors.

Alternative Splicing↗

Autosomal dominant cone dystrophy caused by a novel mutation in the GCAP1 gene (GUCA1A).

PURPOSE: To describe the clinical features and genetic analysis of a family with an autosomal dominant cone dystrophy (adCD). METHODS: Selected members of a family with an autosomal dominant cone dystrophy underwent ophthalmic evaluation. Blood samples were obtained, genomic DNA was isolated, and genomic fragments were amplified by PCR. Linkage to locus D6S1017 was established. DHPLC mutational analysis and direct sequencing were used to identify a mutation in GUCA1A, the gene encoding the guanylate cyclase activating protein 1 (GCAP1). RESULTS: Of 24 individuals who are at risk of the disease in a five generation family, 11 members were affected. Clinical presentations included photophobia, color vision defects, central acuity loss, and legal blindness with advanced age. The disease phenotype was observed in the second and third decades of life and segregated in an autosomal dominant fashion. An electroretinogram performed on one proband revealed profoundly subnormal and prolonged photopic and flicker responses, but preserved scotopic ERGs, consistent with a cone dystrophy. Mutational analysis and direct sequencing revealed a C451T transition in GUCA1A, corresponding to a novel L151F mutation in GCAP1. Like the E155G mutation, this mutation occurs in the EF4 hand domain, a region of GCAP1 critical in conferring calcium sensitivity to the protein. The leucine at this position is highly conserved among vertebrate guanylate cyclase activating proteins. CONCLUSIONS: A novel L151F missense mutation in the EF4 high affinity Ca2+ binding site of GCAP1 is linked to adCD in a large pedigree. The cone dystrophy in this family shares clinical and electrophysiologic characteristics with other previously described adCD caused by mutations in GUCA1A.

Adult↗

An additional promoter functions in the human aldolase A gene, but not in rat.

The aldolase A gene was isolated from a human DNA library, mapped and sequenced. This gene comprises 12 exons and spans 6.5 kb. From the genomic DNA sequence and from the previous sequence analysis of the cDNA, it was revealed that the first exon L1 and the second exon encode the 5' non-coding sequence of mRNA L1, while the third and forth exons (corresponding to exons M and L2) encode different mRNA, mRNA M and L2, respectively; the following eight exons (exons 5-12) are shared commonly by all the mRNA species. These results indicate that the mRNA species are generated from a single aldolase A gene from one of exons L1, M or L2, in addition to exons 5-12, and also that the usage of a leader exon is similar but clearly distinct from that of rat aldolase A gene which we analyzed [Joh, K., Arai, Y., Mukai, T. & Hori, K. (1986) J. Mol. Biol. 190, 401-410]. By comparing the promoter regions in the human and rat aldolase A genes, we found similar sequences in the rat genome corresponding to those of the human L1, M and L2 promoter. We could not, however, detect any transcripts starting from sequences corresponding to the human L1 promoter in the rat genome, although the products corresponding to human M and L2 were detected. Thus, we conclude that the L1 promoter was either acquired by the human genome or deleted from the rat genome after human and rat diverged during evolution.

Amino Acid Sequence↗

Rat hepatic cytosolic phosphoenolpyruvate carboxykinase (GTP). Structures of the protein, messenger RNA, and gene.

The primary structure of the messenger RNA coding for cytosolic phosphoenolpyruvate carboxykinase was determined by sequencing cDNA and genomic DNA and by primer extension of the mRNA. The molecule is 2624 nucleotides in length; this includes 143 nontranslated nucleotides at the 5' end and 615 nontranslated nucleotides at the 3' end. The 3' nontranslated sequence contains a 102-base pair region of alternating purine-pyrimidine nucleotides (the majority of which are UpG dinucleotides), several direct repeats and palindromic sequences, and 8 CpG dinucleotides. The corresponding segment of the phosphoenolpyruvate carboxykinase gene thus has characteristics which favor the formation of Z-DNA. The amino acid sequence of phosphoenolpyruvate carboxykinase was deduced from the mRNA sequence and confirmed by fast atom bombardment mass spectrometric analysis of peptides generated with trypsin and Staphylococcus aureus V8 protease. The protein consists of 621 amino acids and has a molecular weight of 69,289. Charon 4A lambda bacteriophage clones containing genomic DNA coding for phosphoenolpyruvate carboxykinase were isolated from a library of partial HaeIII digests of rat liver DNA. Two clones, lambda PC112 and lambda PC103, contained the entire coding region in 15-kilobase inserts and were used to subclone the gene into pBR322 as EcoRI, BamHI, or SstI-KpnI fragments. Using these subclones, the structure of the phosphoenolpyruvate carboxykinase gene was determined by S1 nuclease mapping, R-loop analysis, and DNA sequencing. The gene is composed of 10 exons and 9 introns with a total length of 6.0 kilobases. The transcription initiation site of the gene was determined by a combination of in vitro transcription in a HeLa cell lysate system, primer extension of mRNAPEPCK, and S1 nuclease mapping. In vitro transcription of purified DNA templates revealed three RNA polymerase II-dependent start sites. Two sites were separated by 600 base pairs on the coding strand and the third site was on the noncoding strand. The products of S1 nuclease mapping and primer extension from a BglII site were compared in order to determine which of the coding strand initiation sites was expressed in vivo. In both cases a 69-base pair fragment was generated and the 5' end of this corresponded to a thymidine residue identified in a sequence ladder of the genomic DNA coding strand. We conclude that mRNAPEPCK synthesis initiates with an adenine residue 69 base pairs 5' of the BglII site; this corresponds to the 3' most transcription initiation site determined in vitro.

Amino Acid Sequence↗