Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Staphylococcal phosphoenolpyruvate-dependent phosphotransferase system. Purification and protein sequencing of the Staphylococcus carnosus histidine-containing protein, and cloning and DNA sequencing of the ptsH gene.

The histidine-containing protein (HPr) of the bacterial phosphoenolpyruvate-dependent phosphotransferase system (PTS) was isolated from Staphylococcus carnosus and purified to homogeneity. The protein sequence was determined by Edman degradation of peptides obtained by proteolytic digestion with proteases V8, trypsin and chemical cleavage with BrCN. Furthermore, immunological screening of a chromosomal S. carnosus DNA gene library in pUC19 vector enabled us to isolate S. carnosus HPr-expressing colonies. The nucleotide sequence of this ptsH gene and its flanking regions was determined by the dideoxy-chain-termination technique. Upstream, the 264-bp open reading frame of the ptsH gene is flanked by a putative S. carnosus promoter structure and a putative ptsI gene downstream suggesting that ptsH gene is the first gene in the PTS operon of S. carnosus. Comparison of the amino acid sequence of S. carnosus HPr with the HPr sequence of Staphylococcus aureus (derived from peptide sequencing) showed a high degree of similarity.

Amino Acid Sequence↗

Complementary DNA sequencing: expressed sequence tags and human genome project.

Automated partial DNA sequencing was conducted on more than 600 randomly selected human brain complementary DNA (cDNA) clones to generate expressed sequence tags (ESTs). ESTs have applications in the discovery of new human genes, mapping of the human genome, and identification of coding regions in genomic sequences. Of the sequences generated, 337 represent new genes, including 48 with significant similarity to genes from other organisms, such as a yeast RNA polymerase II subunit; Drosophila kinesin, Notch, and Enhancer of split; and a murine tyrosine kinase receptor. Forty-six ESTs were mapped to chromosomes after amplification by the polymerase chain reaction. This fast approach to cDNA characterization will facilitate the tagging of most human genes in a few years at a fraction of the cost of complete genomic sequencing, provide new genetic markers, and serve as a resource in diverse biological research fields.

Amino Acid Sequence↗

Identification, DNA sequence, and distribution of IS981, a new, high-copy-number insertion sequence in lactococci.

An insertion in the lactococcal plasmid pGBK17, which inactivated the gene(s) encoding resistance to the prolate-headed phage c2, was cloned, sequenced, and identified as a new lactococcal insertion sequence (IS). IS981 was 1,222 bp in size and contained two open reading frames, one large enough to encode a transposase. IS981 ended in imperfect inverted repeats of 26 of 40 bp and generated a 5-bp direct repeat of target DNA at the site of insertion. IS981 was present on the chromosome of Lactococcus lactis subsp. lactis LM0230 from where it transposed to pGBK17 during transformation. Twenty-three strains of lactococci examined for the presence of IS981 by Southern hybridization showed 4 to 26 copies per genome, with L. lactis subsp. cremoris strains containing the highest number of copies. Comparison of the DNA sequence and the amino acid sequence of the long open reading frame to other known sequences showed that IS981 is related to a family of IS elements that includes IS2, IS3, IS51, IS150, IS600, IS629, IS861, IS904, and ISL1.

Amino Acid Sequence↗

Identification of a 35-kilodalton serovar-cross-reactive flagellar protein, FlaB, from Leptospira interrogans by N-terminal sequencing, gene cloning, and sequence analysis.

During the screening of antibodies to pathogenic leptospires, a murine monoclonal antibody (designated M138) was found to react with various serovars. An antigen of approximately 35 kDa from Leptospira interrogans serovar pomona, which reacted strongly with M138, was characterized by N-terminal amino acid sequencing and identified as a flagellin, a class B polypeptide subunit (FlaB) of the periplasmic flagella. The gene encoding the FlaB protein, flaB, was amplified from the genomic DNA of several pathogenic serovars by PCR with a single pair of oligonucleotide primers, suggesting that FlaB is highly conserved among these serovars. Cloning and sequence analysis of flaB from serovar pomona revealed that it contains an 849-bp open reading frame with a G + C content of 46.88% which encodes a 283-amino-acid protein with a calculated molecular mass of 31.297 kDa and a predicted pI of 9.065. A sequence comparison of flagellin proteins revealed that the amino acid sequence is most variable in the central portion of the serovar pomona FlaB, which is believed to contain specific sequence information and which may thus be useful in the design of DNA or synthetic peptide probes suitable for the detection of infection with pathogenic leptospires.

Amino Acid Sequence↗

Flagellar switch of Salmonella typhimurium: gene sequences and deduced protein sequences.

The fliG, fliM, and fliN genes of Salmonella typhimurium encode flagellar components that participate in energy transduction and switching. We have cloned these genes and determined their sequences. The deduced amino acid sequences correspond to proteins with molecular masses of 36,809, 37,815, and 14,772 daltons, respectively. None of the protein sequences are especially hydrophobic or look as though they correspond to integral membrane proteins, a result consistent with other evidence suggesting that the proteins may be peripheral to the membrane, possibly mounted onto the basal body M ring. The fliL gene, which immediately precedes fliM, is of unknown function; it encodes a protein with a deduced molecular mass of 17,082 daltons. The hydropathy profile of FliL indicates that it is likely to be an integral membrane protein with at least one spanning segment, near its N terminus. None of the four proteins exhibit consensus N-terminal signal sequences. Comparison of the fliL, fliM, and fliN sequences with the homologous ones in Escherichia coli reveals ranges of similarities of 77 to 95% at the amino acid level and 75 to 86% at the nucleotide level, with the majority (58 to 89%) of codon changes being synonymous ones.

Amino Acid Sequence↗

Nucleotide sequencing of the Proteus mirabilis calcium-independent hemolysin genes (hpmA and hpmB) reveals sequence similarity with the Serratia marcescens hemolysin genes (shlA and shlB).

We cloned a 13.5-kilobase EcoRI fragment containing the calcium-independent hemolysin determinant (pWPM110) from a clinical isolate of Proteus mirabilis (477-12). The DNA sequence of a 7,191-base-pair region of pWPM110 was determined. Two polypeptides are encoded in this region, HpmB and HpmA (in that transcriptional order), with predicted molecular masses of 63,204 and 165,868 daltons, respectively. A putative Fur-binding site was identified upstream of hpmB overlapping the -35 region of the proposed hpm promoter. In vitro transcription-translation of pWPM110 DNA and other subclones confirmed the assignment of molecular masses for the predicted polypeptides. These polypeptides are predicted to have NH2-terminal leader peptides of 17 and 29 amino acids, respectively. NH2-terminal amino acid sequence analysis of purified extracellular hemolysin (HpmA) confirmed the cleavage of the 29-amino-acid leader peptide in the secreted form of HpmA. Hemolysis assays and immunoblot analysis of Escherichia coli containing subclones expressing hpmA, hpmB, or both indicated that HpmB is necessary for the extracellular secretion and activation of HpmA. Significant nucleotide identity (52.1%) was seen between hpm and the shl hemolysin gene sequences of Serratia marcescens despite differences in the G+C contents of these genes (hpm, 38%; shl, 65%). The predicted amino acid sequences of HpmB and HpmA are also similar to those of ShlB and ShlA, the respective sequence identities being 55.4 and 46.7%. Predicted cysteine residues and major hydrophobic and amphipathic domains have been strongly conserved in both proteins. Thus, we have identified a new hemolysin gene family among gram-negative opportunistic pathogens.

Amino Acid Sequence↗

Expressed sequence tags for the chicken genome from a normalized 10-day-old white leghorn whole-embryo cDNA library. 3. DNA sequence analysis of genetic variation in commercial chicken populations.

Single nucleotide polymorphisms (SNPs) have emerged as a major class of DNA markers with the advantage of permitting the development of high-density genetic maps adequate for quantitative trait loci (QTL) identification by linkage-disequilibrium analysis. Here we describe results of a relatively high-depth survey of chicken broiler and layer populations for SNPs in targeted genomic regions of chicken expressed sequence tag (EST) sites. The sequences scanned, representing the composite sequence of 12 amplified fragments for a total of 6489 bp, were randomly distributed, occurring on six different chromosomes or linkage groups in the chicken genome. Although one of the genomic DNA sequences did not match the reference cDNA sequence, another contained an intron that separated two putative exons. The number of SNPs observed within each of the 12 EST-targeted genomic regions ranged from 0 to 10 for a total of 44 and a frequency of 0.7%. About 70% of the polymorphisms were shared between layer and broiler populations. The average heterozygosity within the populations ranged from 0.15 to 0.48, with the layer populations showing the higher heterozygosity. SNPs and oligonucleotides described will provide a resource for genetic analysis in commercial chicken populations. The data appear to indicate that the relative frequency of SNPs in the targeted regions scanned is higher than the frequency reported for any of the other regions scanned to date in other eukaryotic genomes. Additionally, the results suggest that the use of DNA pools may offer an efficient approach to SNP detection in chickens, as has been shown in other vertebrates.

Animals↗

The amino acid sequence of neocarzinostatin apoprotein deduced from the base sequence of the gene.

A segment of the neocarzinostatin apoprotein gene corresponding to T30 to A91 of the protein was amplified using a polymerase chain reaction (PCR) with total DNA from Streptomyces carzinostaticus subsp. neocarzinostaticus E-793 (ATCC 15944) as the template and with 5'- and 3'-primers synthesized in consideration of the codon usage of streptomyces. The PCR product was cloned, sequenced and confirmed to direct an amino acid sequence reasonably well matching that reported. Using the PCR product as a probe, we cloned a DNA segment (2580 bp) spanning an open reading frame (ORF) for preapoprotein (leader peptide plus apoprotein) and its upstream and downstream flanking regions. The amino acid sequence deduced from the base sequence of the DNA clearly identified those amino acid residues which had remained inconsistent among different research groups. The base sequence homology with other apoprotein genes of related antibiotics was analyzed and was found to be limited within the structural gene.

Amino Acid Sequence↗

Numerical characterization and similarity analysis of DNA sequences based on 2-D graphical representation of the characteristic sequences.

Based on the classifications of the four nucleic acid bases, He and Wang reduced a DNA sequence to three binary sequences, which are called the characteristic sequences (J. Chem. Inf. Comput. Sci. 42 (2002) 1080). In this paper, we associate each characteristic sequence with a (b)L / (b)L matrix by giving a 2-D 'two horizontal lines' graphical representation, and thus obtain a 3-component vector with entries being the sums of the maximal and minimal eigenvalues of the (b)L / (b)L matrices. The introduced vector results in more simple characterizations and comparisons among the coding sequences of exon 1 of beta-globin gene of eleven different species.

Animals↗

Bovine and feline gastrin cDNA sequences and the amino acid and nucleotide sequence homologies among mammalian species.

The complete nucleotide sequences of cDNAs encoding bovine and feline preprogastrins have been cloned from the antral mucosa mRNA. The gastrin mRNA of each animal encodes a preprogastrin of 104 amino acids consisting of a signal peptide, a prosegment of 37 amino acids, and a gastrin 34 sequence, followed by a glycine (the amide donor). The cleavage following a pair of lysine residues yields gastrin 17. We found that pairs of arginine residues flanking gastrin 34, the typical processing site sequence of all other preprogastrins and many peptide hormones, were arginines in the bovine preprogastrin, but the first basic amino acid pair had changed to Arg-Trp (57-58 residues) instead of Arg-Arg in the feline preprogastrin. Comparison of these amino acid and nucleotide sequences with published mammalian sequences showed extensive homology in the coding (63 to 73% amino acid identity) and in the untranslated regions (67 to 89% identity). Prosequence, the most variable region, shows greater amino acid difference between bovine and human preprogastrin (54% identity), and between bovine and rat preprogastrin (54% identity) than between other species (62 to 82% identity).

Amino Acid Sequence↗

Rapid assignment of nucleotide sequence data to allele types for multi-locus sequence analysis (MLSA) of bacteria using an adapted database and modified alignment program.

A novel database and modified alignment program is described which provides a fast and accurate procedure for assigning nucleotide sequences to allele types for multi-locus sequence analysis (MLSA). The database has between 40 and 160 alleles per organism including Neisseria meningitidis, Streptococcus pneumoniae, Staphylococcus aureus and Haemophilus influenzae. The database directly compares the query nucleotide sequence against all alleles within the database and this system reduces the time taken for the analysis of nucleotide sequence data and assignment of alleles for subsequent sequence analysis.

Alleles↗

Bacillus subtilis alkaline phosphatases III and IV. Cloning, sequencing, and comparisons of deduced amino acid sequence with Escherichia coli alkaline phosphatase three-dimensional structure.

Bacillus subtilis has an alkaline phosphatase multigene family. Two members of this gene family, phoAIII and phoAIV, were cloned, taking advantage of in vitro constructed strains containing a plasmid insertion within one or the other of the structural genes. The DNA sequences of the two genes showed approximately 64% identity at the DNA level and 63% identity in the deduced primary amino acid sequences. The phoAIII and phoAIV genes code for predicted proteins of 47,149 and 45,935 Da, respectively. Comparison of the deduced primary amino acid sequence of the mature proteins with other sequenced alkaline phosphatases from Escherichia coli, yeast, and humans shows 25-30% identity. Based on the refined crystal structure of E. coli alkaline phosphatase, it appears that the active site and the core of the structure are retained in both Bacillus alkaline phosphatases. However, both proteins are truncated at the amino terminus compared with other mature alkaline phosphatases, three sizable surface loops of E. coli are deleted, and a minidomain is replaced with a larger domain in the model. Neither Bacillus alkaline phosphatase sequenced contains any cysteine residues, an amino acid implicated in intrachain disulfide bond formation in other alkaline phosphatases.

Alkaline Phosphatase↗

Nucleotide sequence of Escherichia coli asnB and deduced amino acid sequence of asparagine synthetase B.

The Escherichia coli asparagine synthetase B gene (asnB) has been cloned into a temperature-sensitive, low copy plasmid, pOU71, as shown by the complementation of an E. coli asparagine auxotroph, E. coli JE6279. The nucleotide sequence of asnB and the flanking sequences were determined. The proposed coding region for the gene is 1662 nucleotides in length, and the deduced amino acid sequence of the coding region results in a protein that has a molecular weight of 62,666 and contains 554 amino acids. A promoter region is identified based on the transcription start site that was determined by primer extension experiments. Homology studies of the asnB protein sequence with the human asparagine synthetase and E. coli asparagine synthetase A protein show that there is a high degree of homology with only the human asparagine synthetase. A purF type glutamine amide transfer domain was identified upon inspection of the amino-terminal amino acid sequence of the asparagine synthetase B protein.

Amino Acid Sequence↗

Characterization of the flavoprotein moieties of NADPH-sulfite reductase from Salmonella typhimurium and Escherichia coli. Physicochemical and catalytic properties, amino acid sequence deduced from DNA sequence of cysJ, and comparison with NADPH-cytochrome P-450 reductase.

NADPH-sulfite reductase flavoprotein (SiR-FP) was purified from a Salmonella typhimurium cysG strain that does not synthesize the hemoprotein component of the sulfite reductase holoenzyme. cysJ, which codes for SiR-FP, was cloned from S. typhimurium LT7 and Escherichia coli B, and both genes were sequenced. Physicochemical analyses and deduced amino acid sequences indicate that SiR-FP is an octamer of identical 66-kDa peptides and contains 4 FAD and 4 FMN per octamer. Potentiometric titrations of SiR holoenzyme, SiR-FP, and FMN-depleted SiR-FP yielded the following redox potentials for the prosthetic groups at pH 7.7: E'1 (FMNH./FMN) = -152 mV; E'2 (FMNH2/FMNH.) = -327 mV; E'3 (FADH./FAD) = -382 mV; E'4 (FADH2/FADH.) = -322 mV. Microcoulometric titration of SiR-FP at 25 degrees C yielded data which were in full agreement with these potentials. Spectroscopic and catalytic studies of native SiR-FP and of SiR-FP depleted of FMN support the following electron flow sequence: NADPH----FAD----FMN. FMN can then contribute electrons to the hemoprotein component of sulfite reductase, as well as to cytochrome c and various diaphorase acceptors. The FMN is postulated to cycle between the FMNH2 and FMNH. oxidation states during catalysis; in this sense SiR-FP shares a catalytic mechanism with NADPH-cytochrome P-450 oxidoreductase. SiR-FP domains involved in binding FMN, FAD, and NADPH are proposed from amino acid sequence homologies with Desulfovibrio vulgaris flavodoxin (Dubourdieu, M., and Fox, J.L. (1977) J. Biol. Chem. 252, 1453-1463) and spinach ferredoxin-NADP+ oxidoreductase (Karplus, P.A., Walsh, K.A., and Herriott, J. R. (1984) Biochemistry 23, 6576-6583). Comparison of the deduced amino acid sequences of SiR-FP and NADPH-cytochrome P-450 oxidoreductase (Porter, T. D., and Kasper, C.B. (1985) Proc. Natl. Acad. Sci. U. S.A. 82, 973-977) also showed identities that suggest these two proteins are descended from a common precursor, which contained binding regions for both FMN and FAD.

Amino Acid Sequence↗

Human gastric cathepsin E. Predicted sequence, localization to chromosome 1, and sequence homology with other aspartic proteinases.

The predicted sequence of human gastric cathepsin E (CTSE) was determined by analysis of cDNA clones isolated from a library constructed with poly(A+) RNA from a gastric adenocarcinoma cell line. The CTSE cDNA clones were identified using a set of complementary 18-base oligonucleotide probes specific for a 6-residue sequence surrounding the first active site of all previously characterized human aspartic proteinases. Sequence analysis of CTSE cDNA clones revealed a 1188-base pair open reading frame that exhibited 59% sequence identity with human pepsinogen A. The predicted CTSE amino acid sequence includes a 379-residue proenzyme (Mr = 40,883) and a 17-residue signal peptide. The predicted CTSE amino acid composition was consistent with that of purified material from gastric mucosa and gastric adenocarcinoma cell lines. Additional evidence for the identification of the CTSE cDNA clones was obtained by analysis of poly(A+) RNA isolated from CTSE-producing and -nonproducing gastric adenocarcinoma cell subclones. Three RNA transcripts (3.6, 2.6, and 2.1 kilobases) were identified in poly(A+) RNA isolated from a gastric adenocarcinoma cell line that produced CTSE that were absent from nonproducing subclones. CTSE contains 7 cysteine residues, of which 6 were localized by comparative maximal alignment analysis with pepsinogen A to conserved residues that form intrachain disulfide bonds. The seventh cysteine residue of CTSE is located within the activation peptide region of the proenzyme. We suspect that this residue forms an interchain disulfide bond and thereby determines the dimerization of CTSE proenzyme molecules that is observed under native conditions. The CTSE gene was localized to human chromosome 1 by concurrent cytogenetic and cDNA probe analyses of a panel of human x mouse somatic cell hybrids.

Amino Acid Sequence↗

Molecular cloning and nucleotide sequence of cDNAs encoding the precursors of rat long chain acyl-coenzyme A, short chain acyl-coenzyme A, and isovaleryl-coenzyme A dehydrogenases. Sequence homology of four enzymes of the acyl-CoA dehydrogenase family.

cDNAs encoding the entire coding regions of the precursors (p) of rat long chain acyl-CoA (LCAD), short chain acyl-CoA (SCAD) and isovaleryl-CoA dehydrogenase (IVD) have been cloned and sequenced. Three cDNAs for rat liver LCAD together cover a 1440-base pair region. These cDNAs encode the entire 430-amino acid sequence of pLCAD, including the 30-amino acid leader peptide and the 400-amino acid mature LCAD. A single 1773 base pair cDNA for rat SCAD covers the entire coding region (414 amino acids), including the 26-amino acid leader peptide and the 388-amino acid mature peptide. Four identified IVD cDNAs, when combined, encompass a 2104 base region, and encode 424 amino acids including a 30-amino acid leader peptide and the 394-amino acid mature peptide. The identities of all cDNA clones have been confirmed by matching the amino acid sequences predicted from the respective cDNAs to the amino-terminal and tryptic peptide sequences derived from the corresponding purified rat enzyme. Comparison of the sequences of four rat acyl-CoA dehydrogenases, including LCAD, MCAD, SCAD, and IVD, and two of their human counterparts (MCAD and SCAD) reveals a high degree of homology (57 invariant and 92 near invariant residues: 30.6-35.4% of identical residues in pairwise comparisons), suggesting that these enzymes belong to a gene family and have evolved from a common ancestral gene.

Acyl-CoA Dehydrogenase↗

Isolation and sequence analysis of human cadherin-6 complementary DNA for the full coding sequence and its expression in human carcinoma cells.

The expression pattern of E- and P-cadherin in human carcinomas has been reported by many laboratories. However, little is known about the involvement of other cadherin types in human carcinomas. cDNA clones for a cadherin molecule were isolated from a cDNA library of human hepatocellular carcinoma cells which lacked E- and P-cadherin expression but exhibited cell aggregation activity mediated by an unknown cadherin, and they were subjected to sequence analysis. The overlapped clones covered 4315 nucleotides and were found to encode a typical cadherin molecule consisting of 790 amino acids. Since the deduced amino acid sequence was identical to a partially available human cadherin-6 sequence except for two amino acid residues, the clones were considered to be human cadherin-6 cDNAs encoding the entire open reading frame. The deduced amino acid sequence also showed extremely high homology with recently reported rat K-cadherin, 97% for the putative mature protein, suggesting that cadherin-6 is the human counterpart of rat K-cadherin. Expression of cadherin-6 in various human normal tissues and carcinoma cells was examined by Northern blot analysis using a specific probe corresponding to the signal and precursor sequence. Among normal tissues examined, brain, cerebellum, and kidney showed strong expression of cadherin-6, whereas lung, pancreas, and gastric mucosa showed weak expression. Transcripts of cadherin-6 were not detected in normal liver, whereas four of six hepatocellular carcinoma cell lines examined expressed cadherin-6 abundantly. As reported for rat K-cadherin, three renal carcinoma cell lines also expressed cadherin-6 strongly. The most interesting finding was obtained for small cell lung carcinoma lines. Among 15 of such cell lines examined, all of 11 cadherin-6-positive lines were classified into the classic type, whereas the negative cell lines were all of the variant type. The present results suggest that besides E- and P-cadherin, other cadherin molecules are expressed in human cancers and are responsible for additional biological properties of the carcinoma cells.

Base Sequence↗

Molecular cloning, sequencing, and expression of the 36 kDa protein present in pars planitis. Sequence homology with yeast nucleopore complex protein.

PURPOSE: Patients with active pars planitis have increased levels of a 36 kDa protein (p-36) in their circulation. The current studies were undertaken to determine the primary structure of this protein. METHODS: A degenerate oligonucleotide probe based on the amino terminal sequence of p-36 was used to identify a clone from a human spleen cDNA library. The cDNA insert was subcloned into the EcoR1 site of pUC-19, and both strands were sequenced. Southern blot analysis was used to study the genomic hybridization pattern. p-36 cDNA was subcloned in a pSG5 expression vector, and the construct was used to transfect COS-7 cells. RESULTS: The cDNA sequence contained an open reading frame of 966 base pairs encoding a protein of 322 amino acids, an untranslated region of 322 base pairs, and 2693 base pairs at the 5' and 3' ends, respectively. The deduced amino acid sequence showed 96.8% identity with the carboxy-terminal region of a yeast nucleopore complex protein, nup 100. Southern blot analysis of human genomic DNA revealed a simple hybridization pattern. Transfection of p-36 cDNA in COS-7 cells resulted in the presence of p-36 mRNA and expression of protein. CONCLUSIONS: The 36 kDa protein (p-36) detected at increased levels in the blood of patients with active pars planitis was cloned from a human spleen cDNA library. Its deduced amino acid sequence is homologous with the carboxy-terminal region of a nucleopore complex protein. Thus, we refer to this protein as nup36.

Amino Acid Sequence↗