Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The amino acid sequence of the red kidney bean Fe(III)-Zn(II) purple acid phosphatase. Determination of the amino acid sequence by a combination of matrix-assisted laser desorption/ionization mass spectrometry and automated Edman sequencing.

Purple acid phosphatase of the common bean Phaseolus vulgaris is a homodimeric 110-kDa glycoprotein with a Fe(III)-Zn(II) center in the active site of each monomer. After exchange of Zn(II) for Fe(II), the enzyme spectroscopically and kinetically resembles the mammalian purple acid phosphatases with Fe(III)-Fe(II) centers in monomeric 35-kDa proteins. The kidney bean enzyme consists of 432 amino acids/monomer with five N-glycosylated asparagine residues. The complete amino acid sequence was determined by a combination of matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) and classical sequencing methods. Our strategy involved mass determination and sequence analysis of all cyanogen-bromide-generated fragments by automated Edman degradation. Limited cleavages with cyanogen bromide were performed to obtain fragments containing still uncleaved Met-Xaa linkages. MALDI mass spectra of these products allowed the characterization of each fragment and the determination of the order of the cyanogen bromide fragments in the intact protein without producing overlapping peptides. For one large 30-kDa methionine-free fragment, the alignment of the Edman-degraded tryptic peptides was obtained by MALDI-MS analysis and enzymic microscale peptide laddering of overlapping Glu-C-generated fragments. The employed strategy shows that the classical method, in combination with modern mass spectrometry, is an attractive approach for primary structure determination in addition to the DNA sequencing method.

Acid Phosphatase↗

Nucleotide sequence of the right early region of Bacillus subtilis phage PZA completes the 19366-bp sequence of PZA genome. Comparison with the homologous sequence of phage phi 29.

We have sequenced the rightmost 2079 bp of the Bacillus subtilis phage PZA genome. This region encompasses the right early region. We compared it with the homologous region of phage phi 29. Six open reading frames (ORFs) were found in this region of PZA and one of them was assigned to gene 17. Analysis of putative ribosome-binding sites and comparison with phi 29 ORFs indicate that at least some of the remaining ORFs could encode proteins. Corresponding genes were not identified so far by genetic methods. Promoter candidates in the right early region of PZA were found and compared to phi 29 promoters. The sequenced region together with previously determined sequences [Paces et al., Gene 38 (1985) 45-56 and 44 (1986) 107-114] completes the entire 19,366-bp sequence of phage PZA genome.

Bacillus subtilis↗

Nucleotide sequence of the late region of Bacillus phage phi 29 completes the 19,285-bp sequence of phi 29 genome. Comparison with the homologous sequence of phage PZA.

The 12,177-bp nucleotide (nt) sequence of the late region of Bacillus phage luminal diameter 29 genome was determined. This sequence completes the entire 19,285-bp sequence of phage luminal diameter 29 DNA. Eleven open reading frames were found in this region, and these were assigned to eleven late genes. Ribosome-binding sites and a potential transcriptional promoter and terminator are considered. The nt sequence was compared to the homologous region of the closely related phage PZA and tolerated variations at the nt and amino acid (aa) level were evaluated. The most frequent changes are silent nt substitutions in the third position of codons, but aa substitutions are also found.

Bacteriophages↗

Amino acid sequence of the Bb fragment from complement Factor B. Sequence of the major cyanogen bromide-cleavage peptide (CB-II) and completion of the sequence of the Bb fragment.

The amino acid sequence of peptide CB-II, the major product (mol.wt. 30 000) of CNBr cleavage of fragment Bb from human complement Factor B, is given. The sequence was obtained from peptides derived by trypsin cleavage of peptide CB-II and clostripain digestion of fragment Bb. Cleavage of two Asn-Gly bonds in peptide CB-II was also found useful. These results, along with those presented in the preceding paper [Gagnon & Christie (1983) Biochem. J. 209, 51-60], yield the complete sequence of the 505 amino acid residues of fragment Bb. The C-terminal half of the molecule shows strong homology of sequence with serine proteinases. Factor B has a catalytic chain (fragment Bb) with a molecular weight twice that of proteinases previously described, suggesting that it is a novel type of serine proteinase, probably with a different activation mechanism.

Amino Acid Sequence↗

Sequence of the A-protein of coliphage MS2. I. Isolation of A-protein, determination of the NH2- and COOH-terminal sequences, isolation and amino acid sequence of the tryptic peptides.

The A-protein of coliphage MS2 was purified to a state of sufficient homogeneity to study its primary structure. The NH2-terminal sequence was determined for the first 8 residues. Comparison with the reported sequence of R17 protein (Weiner, A. M., Platt, T., and Weber, K. (1972) J. Biol. Chem. 247, 3242-3251) shows a difference at position 6 where alanine in R17 is replaced by threonine in MS2. The COOH-terminal sequence was shown to be -Arg-Leu-Ser-Arg, confirming the existence of UAG as the termination codon of the maturation protein (Comtreras, R., Ysebaert, M., Min Jou, W., and Fiers, W. (19731 Nature New Biol. 241, 99-101; Vandekerckhove, J., Nolf, F., and Van Montagu, M. C. (1973) Nature New Biol. 241, 102; Remaut E., and Fiers, W. (1972) J. Mol. Biol. 71, 243-261). Peptides obtained by enzymatic hydrolysis with trypsin were fractionated by a combination of gel filtration and paper electrophoresis and chromatography. Thirty-eight peptides were analyzed for amino acid composition and sequence. They provide information for 312 of the 393 residues of the A-protein polypeptide chain.

Amino Acid Sequence↗

Complete nucleotide sequence of a gene coding for heat- and pH-stable alpha-amylase of Bacillus licheniformis: comparison of the amino acid sequences of three bacterial liquefying alpha-amylases deduced from the DNA sequences.

The gene coding for the heat-stable and pH-stable alpha-amylase of Bacillus licheniformis 584 (ATCC 27811) was cloned in Escherichia coli and the nucleotide sequence of a DNA fragment of 1,948 base pairs containing the entire amylase gene was determined. As inferred from the DNA sequence, the B. licheniformis alpha-amylase had a signal peptide of 29 amino acid residues and the mature enzyme comprised 483 amino acid residues, giving a molecular weight of 55,200. The amino acid sequence of B. licheniformis alpha-amylase showed 65.4% and 80.3% homology with those of heat-stable Bacillus stearothermophilus alpha-amylase and relatively heat-unstable Bacillus amyloliquefaciens alpha-amylase, respectively. Nevertheless, several regions of the alpha-amylases appeared to be clearly distinct from one another when their hydropathy profiles were compared.

Amino Acid Sequence↗

Nucleotide sequence of the long terminal repeat and flanking cellular sequences of avian endogenous retrovirus ev-2: variation in Rous-associated virus-0 expression cannot be explained by differences in primary sequence.

A fragment of chicken DNA containing the left long terminal repeat of endogenous retrovirus ev-2 and flanking cellular sequences has been molecularly cloned and analyzed. Comparison with sequence data from the analogous regions of ev-1 and Rous-associated virus-0 viral DNA reveals similarities among flanking regions of the integrated proviruses and among all three long terminal repeats. From the latter finding, we conclude that the difference in level of expression of ev-2 and its progeny Rous-associated virus-0 provirus cannot be due to sequence differences in their upstream long terminal repeats.

Animals↗

Amino acid sequence studies on the alpha chain of human fibrinogen. Overlapping sequences providing the complete sequence.

The complete amino acid sequence of the alpha chain of human fibrinogen has been determined. It contains 610 amino acid residues and has a calculated molecular weight of 66,124. The chain has 10 methionines, and fragmentation with cyanogen bromide yields 11 peptides [Doolittle, R.F., Cassman, K.G., Cottrell, B.A., Friezner, S.J., Hucko, J.T., & Takagi, T. (1977) Biochemistry 16, 1703]. The arrangement of the 11 fragments was determined by the isolation of peptide overlaps from plasmic and staphylococcal protease digests of fibrinogen and/or alpha chains. In addition, certain of the cyanogen bromide fragments, preliminary reports of whose sequences have appeared previously, have been reexamined in order to resolve several discrepancies. The alpha chain is homologous with the beta and gamma chains of fibrinogen, although a large repetitive segment of unusual composition is absent from the latter two chains. The existence of this unusual segment divides the sequence of the alpha chain into three zones of about 200 residues each that are readily distinguishable on the basis of amino acid composition alone.

Amino Acid Sequence↗

Direct sequence analysis of human herpesvirus 6 (HHV-6) sequences from infants and comparison of HHV-6 sequences from mother/infant pairs.

Direct sequence analysis of polymerase chain reaction-amplified DNA fragments from the large tegument protein (LTP) gene of human herpesvirus 6 (HHV-6) was performed with use of uncultured peripheral blood mononuclear cells (PBMCs) from four mother/infant pairs. In two cases, LTP gene sequences were identical in paired mother/infant specimens, thus suggesting that mother-to-infant transmission of HHV-6 may have occurred. The genetic stability of HHV-6 strains was confirmed by the fact that there was no difference between amplified DNA fragments from sequential PBMC samples from two of two infants analyzed. In contrast, a change in the amplified viral strain was detected in an infant who had reinfection with HHV-6 variant B (HHV-6B). Furthermore, HHV-6B strains concurrently amplified from saliva and PBMCs from an adult were found to be different. The data suggest that HHV-6 may be frequently transmitted from mother-to-infant and that reinfection with HHV-6B may occur.

Adult↗

Comparison of sample sequences of the Salmonella typhi genome to the sequence of the complete Escherichia coli K-12 genome.

Raw sequence data representing the majority of a bacterial genome can be obtained at a tiny fraction of the cost of a completed sequence. To demonstrate the utility of such a resource, 870 single-stranded M13 clones were sequenced from a shotgun library of the Salmonella typhi Ty2 genome. The sequence reads averaged over 400 bases and sampled the genome with an average spacing of once every 5,000 bases. A total of 339,243 bases of unique sequence was generated (approximately 7% representation). The sample of 870 sequences was compared to the complete Escherichia coli K-12 genome and to the rest of the GenBank database, which can also be considered a collection of sampled sequences. Despite the incomplete S. typhi data set, interesting categories could easily be discerned. Sixteen percent of the sequences determined from S. typhi had close homologs among known Salmonella sequences (P < 1e-40 in BlastX or BlastN), reflecting the proportion of these genomes that have been sequenced previously; 277 sequences (32%) had no apparent orthologs in the complete E. coli K-12 genome (P > 1e-20), of which 155 sequences (18%) had no close similarities to any sequence in the database (P > 1e-5). Eight of the 277 sequences had similarities to genes in other strains of E. coli or plasmids, and six sequences showed evidence of novel phage lysogens or sequence remnants of phage integrations, including a member of the lambda family (P < 1e-15). Twenty-three sample sequences had a significantly closer similarity a sequence in the database from organisms other than the E. coli/Salmonella clade (which includes Shigella and Citrobacter). These sequences are new candidate lateral transfer events to the S. typhi lineage or deletions on the E. coli K-12 lineage. Eleven putative junctions of insertion/deletion events greater than 100 bp were observed in the sample, indicating that well over 150 such events may distinguish S. typhi from E. coli K-12. The need for automatic methods to more effectively exploit sample sequences is discussed.

Bacteriophages↗

The beta globin gene cluster of the prosimian primate Galago crassicaudatus: nucleotide sequence determination of the 41-kb cluster and comparative sequence analyses.

The nucleotide sequence of the beta globin gene cluster of the prosimian Galago crassicaudatus has been determined. A total sequence spanning 41,101 bp contains and links together previously published sequences of the five galago beta-like globin genes (5'-epsilon-gamma-psi eta-delta-beta-3'). A computer-aided search for middle interspersed repetitive sequences identified 10 LINE (L1) elements, including a 5' truncated repeat that is orthologous to the full-length L1 element found in the human epsilon-gamma intergenic region. SINE elements that were identified included one Alu type I repeat, four Alu type II repeats, and two methionine tRNA-derived Monomer (type III) elements. Alu type II and Monomer sequences are unique to the galago genome. Structural analyses of the cluster sequence reveals that it is relatively A+T rich (about 62%) and regions with high G+C content are associated primarily with globin coding regions. Comparative analyses with the beta globin cluster sequences of human, rabbit, and mouse reveal extensive sequence homologies in their genic regions, but only human, galago, and rabbit sequences share extensive intergenic sequence homologies. Divergence analyses of aligned intergenic and flanking sequences from orthologous human, galago, and rabbit sequences show a gradation in the rate of nucleotide sequence evolution along the cluster where sequences 5' of the epsilon globin gene region show the least sequence divergence and sequences just 5' of the beta globin gene region show the greatest sequence divergence.

Amino Acid Sequence↗

Analysis of sequence-tagged-connector strategies for DNA sequencing.

The BAC-end sequencing, or sequence-tagged-connector (STC), approach to genome sequencing involves sequencing the ends of BAC inserts to scatter sequence tags (STCs) randomly across the genome. Once any BAC or other large segment of DNA is sequenced to completion by conventional shotgun approaches, these STC tags can be used to identify a minimum tiling path of BAC clones overlapping the nucleation sequence for sequence extension. Here, we explore the properties of STC-sequencing strategies within a mathematical model of a random target with homologous repeats and imperfect sequencing technology to understand the consequences of varying various parameters on the incidence of problem clones and the cost of the sequencing project. Problem clones are defined as clones for which either (A) there is no identifiable overlapping STC to extend the sequence in a particular direction or (B) the identified STC with minimum overlap comes from a nonoverlapping clone, either owing to random false matches or repeat-family homology. Based on the minimum overlap, we estimate the number of clones to be entirely sequenced and, then, using cost estimates, identify the decision rule (the degree of sequence similarity required before a match is declared between an STC and a clone) to minimize overall sequencing cost. A method to optimize the overlap decision rule is highly desirable, because both the total cost and the number of problem clones are shown to be highly sensitive to this choice. For a target of 3 Gb containing approximately 800 Mb of repeats with 85%-90% identity, we expect <10 problem clones with 15 times coverage by 150-kb clones. We derive the optimal redundancy and insert sizes of clone libraries for sequencing genomes of various sizes, from microbial to human. We estimate that establishing the resource of STCs as a means of identifying minimally overlapping clones represents only 1%-3% of the total cost of sequencing the human genome, and, up to a point of diminishing returns, a larger STC resource is associated with a smaller total sequencing cost.

Genome, Human↗

Prediction of protein structure by evaluation of sequence-structure fitness. Aligning sequences to contact profiles derived from three-dimensional structures.

The problem of protein structure prediction is formulated here as that of evaluating how well an amino acid sequence fits a hypothetical structure. The simplest and most complicated approaches, secondary structure prediction and all-atom free energy calculations, can be viewed as sequence-structure fitness problems. Here, an approach of intermediate complexity is described, which involves; (1) description of a protein structure in terms of contact interface vectors, with both intra-protein and protein-solvent contacts counted, (2) derivation of sequence preferences for 2 up to 29 contact interface types, (3) generation of numerous hypothetical model structures by placing the input sequence into a large set of known three-dimensional structures in all possible alignments, (4) evaluation of these models by summing the sequence preferences over all structural positions and (5) choice of predicted three-dimensional structure as that with the best sequence-structure fitness. Evolutionary information is incorporated by using position-dependent core weights derived from multiple sequence alignments. A number of tests of the method are performed: (1) evaluation of cyclic shifts of a sequence in its native structure; (2) alignment of a sequence in its native structure, allowing gaps; (3) alignment search with a sequence or sequence fragment in a database of structures; and (4) alignment search with a structure in a database of sequences. The main results are: (1) a native sequence can very well find its native structure among a large number of alternatives, in correct alignment; (2) substructures, such as (beta alpha)n units, can be detected in spite of very low sequence similarity; (3) remote homologous can be detected, with some dependence on the set of parameters used; (4) contact interface parameters are clearly superior to classical secondary structure parameters; (5) a simple interface description in terms of just two states, protein-protein and protein-water contacts, performs surprisingly well; (6) the use of core weights considerably improves accuracy in detection of remote homologues; (7) based on a sequence database search with a myoglobin contact profile, the C-terminal domain of a viral origin of replication binding protein is predicted to have an all-helical fold. The sequence-structure fitness concept is sufficiently general to accommodate a large variety of protein structure prediction methods, including new models of intermediate complexity currently being developed.

Amino Acid Sequence↗

Effect of intergenic consensus sequence flanking sequences on coronavirus transcription.

Insertion of a region, including the 18-nucleotide-long intergenic sequence between genes 6 and 7 of mouse hepatitis virus (MHV) genomic RNA, into an MHV defective interfering (DI) RNA leads to transcription of subgenomic DI RNA in helper virus-infected cells (S. Makino, M. Joo, and J. K. Makino, J. Virol. 66:6031-6041, 1991). In this study, the subgenomic DI RNA system was used to determine how sequences flanking the intergenic region affect MHV RNA transcription and to identify the minimum intergenic sequence required for MHV transcription. DI cDNAs containing the intergenic region between genes 6 and 7, but with different lengths of upstream or downstream flanking sequences, were constructed. All DI cDNAs had an 18-nucleotide-long intergenic region that was identical to the 3' region of the genomic leader sequence, which contains two UCUAA repeat sequences. These constructs included 0 to 1,440 nucleotides of upstream flanking sequence and 0 to 1,671 nucleotides of downstream flanking sequence. An analysis of intracellular genomic DI RNA and subgenomic DI RNA species revealed that there were no significant differences in the ratios of subgenomic to genomic DI RNA for any of the DI RNA constructs. DI cDNAs which lacked the intergenic region flanking sequences and contained a series of deletions within the 18-nucleotide-long intergenic sequence were constructed to determine the minimum sequence necessary for subgenomic DI RNA transcription. Small amounts of subgenomic DI RNA were synthesized from genomic DI RNAs with the intergenic consensus sequences UCUAAAC and GCUAAAC, whereas no subgenomic DI RNA transcription was observed from DI RNAs containing UCUAAAG and GCTAAAG sequences. These analyses demonstrated that the sequences flanking the intergenic sequence between genes 6 and 7 did not play a role in subgenomic DI RNA transcription regulation and that the UCUAAAC consensus sequence was sufficient for subgenomic DI RNA transcription.

Animals↗

Mechanisms for the generation of src-deletion mutants and recovered sarcoma viruses: identification of viral sequences involved in src deletions and in recombination with c-src sequences.

The precise src deletions in six transformation-defective (td) deletion mutants derived from the Schmidt-Ruppin strain of Rous sarcoma virus were determined by sequence analysis. Examination of the parental viral sequences neighboring the junctions of deletions in these td mutants revealed that these regions contained either directly repeated or inverted complementary sequences ranging from 9 to 28 nucleotides. Five td mutants were found to contain deletions flanked by directly repeated sequences, of which the 3' direct repeat was retained whereas the 5' direct repeat was deleted in the resulting td viral RNA. In the deletions of two td mutants where inverted complementary sequences were present at junctions of the deletions, both copies of the inverted complementary sequence were deleted in the td viral RNA. It is proposed from these observations that deletions of these mutants have been generated during the synthesis of minus-strand viral DNA by reverse transcriptase by jumping over a sequence flanked by direct repeats or by skipping a stem-and-loop structure formed via inverted complementary sequences on the viral RNA template. Data provide further information on the sequences in the td viral genome that are required for the generation of recovered sarcoma viruses (rASVs) by recombination with c-src. Sequence data of td viruses revealed that retaining as few as 82 nucleotides of the 3' src coding sequence is sufficient, whereas retaining as much as one-third of the 5' src but none of the 3' src coding sequences is not sufficient, for the generation of rASVs. Those that generate replication-competent rASVs retain, in addition to the 3' src region, a portion of the 5' src and/or its immediate upstream sequence that is homologous to exon 1 of the c-src DNA. These two sequence domains apparently provided 5' and 3' homologous regions for recombination between td viral genome and c-src DNA resulting in nondefective rASVs. Td109, which was shown previously to generate only replication-defective rASVs, retains 296 nucleotides of the 3' src sequence but lacks all the 5' src and 316 nucleotides of its immediate upstream region. It is concluded that the 5' src coding sequence and its immediate upstream region are not essential for the generation of rASVs. However, retaining a portion of those sequences is required for the generation of replication-competent rASVs.

Animals↗

Sequence database searches via de novo peptide sequencing by tandem mass spectrometry.

A method is described for searching protein sequence databases using tandem mass spectra of tryptic peptides. The approach uses a de novo sequencing algorithm to derive a short list of possible sequence candidates which serve as query sequences in a subsequent homology-based database search routine. The sequencing algorithm employs a graph theory approach similar to previously described sequencing programs. In addition, amino acid composition, peptide sequence tags and incomplete or ambiguous Edman sequence data can be used to aid in the sequence determinations. Although sequencing of peptides from tandem mass spectra is possible, one of the frequently encountered difficulties is that several alternative sequences can be deduced from one spectrum. Most of the alternative sequences, however, are sufficiently similar for a homology-based sequence database search to be possible. Unfortunately, the available protein sequence database search algorithms (e.g. Blast or FASTA) require a single unambiguous sequence as input. Here we describe how the publicly available FASTA computer program was modified in order to search protein databases more effectively in spite of the ambiguities intrinsic in de novo peptide sequencing algorithms.

Algorithms↗

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition↗

Sequenase sequence profiles used for HLA-DPB1 sequencing-based typing.

Sequencing-based HLA typing (SBT) is a PCR based high resolution HLA typing method in which polymorphic regions of the gene are sequenced and directly used for typing. Currently, for class II SBT, alleles are identified by comparison of the exon 2 sequence with their corresponding allele sequence library. Routine SBT requires reliable identification of heterozygosity, and automated assignment of the alleles. In sequencing strategies different enzymes can be used for primer extension. The most characteristic difference between sequences obtained by two protocols using Sequenaseregistered, or Taq-cycle sequencing, respectively, is a difference in incorporation of nucleotides in the primer extension leading to different sequence profiles. In Taq-cycling sequencing variable nucleotide incorporation results in irregular, but reproducible peak patterns, whereas Sequenase incorporates nucleotides in nearly equal amounts, resulting in more even peak patterns. In a previously published multi-center study we evaluated HLA-DPB1 SBT using Taq-cycle sequencing, and showed that typing can reliably be performed, considering the specific sequence profiles. In this study the applicability of Sequenase for HLA-DPB1 SBT was tested. A panel of samples were typed by SBT at five test sites which participate in the Sequencing Based Typing component of the 12th International Histocompatibility Workshop. The panel represents the existing polymorphism at all known polymorphic positions of exon 2, both in homozygous and heterozygous combinations. The assignment of homozygosity and heterozygosity was validated by Multi-Sequence Analysis, performing cluster analysis of chromatographic data of all sequences at each position. Sequence characteristics were examined and considered for appropriate assignment. Data reveals that Sequenase sequencing can also reliably be used for HLA-DPB1 typing.

Bacteriophage T7↗