Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Sequence of 2,617 nucleotides from the 3' end of Newcastle disease virus genome RNA and the predicted amino acid sequence of viral NP protein.

DNA fragments complementary to the Newcastle disease virus genome (strain D26) were cloned and sequenced. The sequence of 2,617 nucleotides from the 3' end of the genome was determined and an open reading frame (OP-1) consisting of 1,467 nucleotides, most likely encoding NP protein, was found in this region. This was followed by a second unfinished open reading frame (OP-2) of at least 729 nucleotides which continued beyond the 2,617th nucleotide. Another relatively short (312 nucleotides long) open reading frame (OP-2') was found overlapping with OP-2, but its significance is still unclear. The amino acid sequence deduced from the nucleotide sequence of OP-1 showed a moderate homology to that of the NP protein of Sendai virus in the central portion of the peptide. The leader sequence of 53 nucleotides was also identified. The 5' end of mRNAs synthesized in the infected cells was analyzed and found to be m7GpppA, suggesting that the transcription of viral mRNAs starts with A, but not with G residue.

Amino Acid Sequence↗

Alignment of the amino terminal amino acid sequence of human cytochrome c oxidase subunits I and II with the sequence of their putative mRNAs.

Thirteen of the first fifteen amino acids from the NH2-terminus of the primary sequence of human cytochrome c oxidase subunit I and eleven of the first twelve amino acids of subunit II have been identified by microsequencing procedures. These sequences have been compared with the recently determined 5'-end proximal sequences of the HeLa cell mitochondrial mRNAs and unambiguously aligned with two of them. This alignment has allowed the identification of the putative mRNA for subunit I, and has shown that the initiator codon for this subunit is only three nucleotides away from the 5'-end of its mRNA; furthermore, the results have substantiated the idea that the translation of human cytochrome c oxidase subunit II starts directly at the 5'-end of its putative mRNA, as had been previously inferred on the basis of the sequence homology of human mitochondrial DNA with the primary sequences of the bovine subunit.

Amino Acid Sequence↗

Sequence tagged microsatellite profiling (STMP): improved isolation of DNA sequence flanking target SSRs.

Sequence tagged microsatellite profiling (STMP) enables the rapid development of large numbers of co-dominant DNA markers, known as sequence tagged microsatellites (STMs). Each STM is amplified by PCR using a single primer specific to the conserved DNA sequence flanking the microsatellite repeat in combination with a universal primer that anchors to the 5'-ends of the microsatellites. It is also possible to convert STMs into conventional microsatellite, or simple sequence repeat (SSR), markers that are amplified using a pair of primers flanking the repeat sequence. Here, we describe a modification of the STMP procedure to significantly improve the capacity to convert STMs into conventional SSRs and, therefore, facilitate the development of highly specific DNA markers for purposes such as marker-assisted breeding. The usefulness of this technique was demonstrated in bread wheat.

Conserved Sequence↗

Multi-priming sequencing: a DNA sequencing method involving restriction enzyme-digested DNA fragments as primers.

An improved strategy for fluorescence-labeled dideoxy chain termination sequencing involving restriction enzyme-digested DNA fragments as primers, which are prepared from the DNA to be sequenced, is described. By using modified nucleoside triphosphates for strand protection in chain termination reactions, newly synthesized chains were detached from a primer at the regenerated recognition site by means of suitable restriction enzyme digestion. The digests could be analyzed with commercial automated DNA sequencers. Thus, by using restriction DNA fragments (double-stranded) as primers, sequence information was obtained from both "minus" and "plus" single-stranded DNA templates without subcloning. Nor is the synthesis of oligonucleotide primers needed. This method, named "Multi-Priming Sequencing," was proven to be time-saving, economical, and effective compared to conventional methods.

Base Sequence↗

The primary structure of phosphoenolpyruvate carboxylase of Escherichia coli. Nucleotide sequence of the ppc gene and deduced amino acid sequence.

The nucleotide sequence of the ppc gene, the structural gene for phosphoenolpyruvate carboxylase [EC 4.1.1.31], of Escherichia coli K-12 was determined. The gene codes for a polypeptide comprising 883 amino acid residues with a calculated molecular weight of 99,061. The amino acid sequence deduced from the nucleotide sequence was entirely consistent with the protein chemical data obtained with the purified enzyme, including the NH2- and COOH-terminal sequences and amino acid composition. The coding region is preceded by two putative ribosome binding sites, and is followed closely by a good representative of rho-independent terminator. The codon usage in the ppc gene suggests a moderate expression of the gene. The secondary structure of the enzyme was predicted from the deduced amino acid sequence.

Amino Acid Sequence↗

Automated sequence preprocessing in a large-scale sequencing environment.

A software system for transforming fragments from four-color fluorescence-based gel electrophoresis experiments into assembled sequence is described. It has been developed for large-scale processing of all trace data, including shotgun and finishing reads, regardless of clone origin. Design considerations are discussed in detail, as are programming implementation and graphic tools. The importance of input validation, record tracking, and use of base quality values is emphasized. Several quality analysis metrics are proposed and applied to sample results from recently sequenced clones. Such quantities prove to be a valuable aid in evaluating modifications of sequencing protocol. The system is in full production use at both the Genome Sequencing Center and the Sanger Centre, for which combined weekly production is approximately 100, 000 sequencing reads per week.

Automation↗

Comparison of sequence profiles. Strategies for structural predictions using sequence information.

Distant homologies between proteins are often discovered only after three-dimensional structures of both proteins are solved. The sequence divergence for such proteins can be so large that simple comparison of their sequences fails to identify any similarity. New generation of sensitive alignment tools use averaged sequences of entire homologous families (profiles) to detect such homologies. Several algorithms, including the newest generation of BLAST algorithms and BASIC, an algorithm used in our group to assign fold predictions for proteins from several genomes, are compared to each other on the large set of structurally similar proteins with little sequence similarity. Proteins in the benchmark are classified according to the level of their similarity, which allows us to demonstrate that most of the improvement of the new algorithms is achieved for proteins with strong functional similarities, with almost no progress in recognizing distant fold similarities. It is also shown that details of profile calculation strongly influence its sensitivity in recognizing distant homologies. The most important choice is how to include information from diverging members of the family, avoiding generating false predictions, while accounting for entire sequence divergence within a family. PSI-BLAST takes a conservative approach, deriving a profile from core members of the family, providing a solid improvement without almost any false predictions. BASIC strives for better sensitivity by increasing the weight of divergent family members and paying the price in lower reliability. A new FFAS algorithm introduced here uses a new procedure for profile generation that takes into account all the relations within the family and matches BASIC sensitivity with PSI-BLAST like reliability.

Algorithms↗

Solid-phase sequence analysis of polypeptides eluted from polyacrylamide gels. An aid to interpretation of DNA sequences exemplified by the Escherichia coli unc operon and bacteriophage lambda.

An approach to sequencing proteins by the solid-phase method combined with isolation of proteins and polypeptides by gel electrophoresis is described. Mixtures of proteins or polypeptides resulting from digests are fractionated in the presence of dodecylsulphate in polyacrylamide gels. They are detected with Coomassie blue, eluted, selectively reacted with porous glass derivatives and sequenced in their amino-terminal regions with the aid of a new microsequencer. Alternatively they can be analysed or digested with enzymes and fingerprinted. It is a relatively rapid method of purifying proteins for sequence analysis which we have used to provide partial protein sequence data to complement DNA sequences. Nine genes, four from the unc operon of Escherichia coli encoding the alpha, beta, gamma and epsilon subunits of ATP synthase and five for capsid proteins of bacteriophage lambda, have been identified by this method.

Adenosine Triphosphatases↗

Amino acid sequences of hemoglobins I and II from root nodules of the non-leguminous Parasponia rigida-rhizobium symbiosis, and a correction of the sequence of hemoglobin I from Parasponia andersonii.

The amino acid sequence of hemoglobins I (pI 6.15 as oxyhemoglobin) and II (pI 5.64 as oxyhemoglobin) from the nitrogen-fixing root nodules of Parasponia rigida have been determined by protein sequencing. The sequence of hemoglobin I (pI 6.16, as oxyhemoglobin) from Parasponia andersonii was re-examined and the corrected primary structure, now in agreement with that predicted from the DNA sequence, is reported. The three Parasponia hemoglobins contain 161 amino acid residues (Mr approximately equal to 18,700 including the heme) with a single cysteine residue and five methionine residues. The N-terminal serine is blocked by an acetyl group. The primary structure of the Parasponia hemoglobins is highly conserved. Hemoglobins I from the two species of Parasponia are identical; both show microheterogeneity at position 30 (Asp/Glu substitution) and hemoglobin I from P. rigida shows microheterogeneity at position 150 (Ala/Val) while hemoglobin I from P. andersonii has only an Ala at 150. P. rigida hemoglobin II shows no microheterogeneity at these positions, having Asp and Val residues respectively, and it contains a single amino acid change of a Gln for an Arg at position 85, which accounts for the 0.5 unit difference in isoelectric point observed between hemoglobins I and II. The sequence data are consistent with allelic heterogeneity at a single locus rather than different genes.

Amino Acid Sequence↗

CTBP1/RBP1, a Saccharomyces cerevisiae protein which binds to T-rich single-stranded DNA containing the 11-bp core sequence of autonomously replicating sequence, is a poly(deoxypyrimidine)-binding protein.

South-Western screening of a glutathione-S-transferase fusion protein library constructed from the yeast Saccharomyces cerevisiae genomic DNA lead to isolation of core T-rich-strand-binding protein (CTBP) clones that bound to single-stranded DNA containing the T-rich-strand of the 11-bp core sequence of autonomously replicating sequences. One of these clones, CTBP1, contains a portion of previously described RBP1 which is an RNA-binding and single-stranded DNA-binding protein of S. cerevisiae. GST-CTBP1 as well as the full-length fusion protein with RBP1 (GST-RBP1) bind exclusively to the T-rich strand of the core sequence with an apparent dissociation constant of 5 nM, but not to the A-rich strand or double strand of the same sequence. Mutations within the core which reduce the number of T or C residues decrease the affinity of this protein. In keeping with this, binding of GST-CTBP1 to the core sequence is efficiently completed by poly(dT), poly(dT-dC) or poly(dC), but not by poly(dA) or poly(dG) to significant extents. Among polyribonucleic acids, GST-CTBP1 binds to poly(U) and poly(I) with greatest affinity, whereas GST-RBP1 binds to RNA in a rather non-specific manner. In no cases was affinity for RNA greater than that for DNA. Our results indicate that CTBP1/RBP1 is a polydeoxypyrimidine-binding protein of S. cerevisiae. CTBP1 contains two sets of an RNA-recognition motif (RRM) and a glutamine stretch. The binding affinity of the N-terminal or C-terminal set containing one RRM and one glutamine stretch is nearly two orders of magnitude lower than that of the wild-type CTBP1 containing both sets. The isolated N-terminal or C-terminal RRM alone (RRM1 and RRM2, respectively) is sufficient for binding nucleic acids with the binding specificity similar to that of the wild-type RRM, although the binding affinity of the isolated RRM2 is nearly two orders of magnitude lower than that of RRM1. Our results indicate that the two RRMs present in CTBP1/RBP1 have differential binding affinities and that the high affinity of RRM for polydeoxypyrimidine results from synergy between two lower-affinity RRMs.

Alcohol Oxidoreductases↗

Identification of mycobacteria from animals by restriction enzyme analysis and direct DNA cycle sequencing of polymerase chain reaction-amplified 16S rRNA gene sequences.

Two methods, based on analysis of the polymerase chain reaction-amplified 16S rRNA gene by restriction enzyme analysis (REA) or direct cycle sequencing, were developed for rapid identification of mycobacteria isolated from animals and were compared to traditional phenotypic typing. BACTEC 7H12 cultures of the specimens were examined for "cording," and specific polymerase chain reaction amplification was performed to identify the presence of tubercle complex mycobacteria. Combined results of separate REAs with HhaI, MspI, MboI, and ThaI differentiated 12 of 15 mycobacterial species tested. HhaI, MspI, and ThaI restriction enzyme profiles differentiated Actinobacillus species from mycobacterial species. Mycobacterium bovis could not be differentiated from M. bovis BCG or Mycobacterium tuberculosis. Similarly, Mycobacterium avium and Mycobacterium paratuberculosis could not be distinguished from each other by REA but were differentiated by cycle sequencing. Compared with traditional typing, both methods allowed rapid and more accurate identification of acid-fast organisms recovered from 21 specimens of bovine and badger origin. Two groups of isolates were not typed definitively by either molecular method. One group of four isolates may constitute a new species phylogenetically very closely related to Mycobacterium simiae. The remaining unidentified isolates (three badger and one bovine) had identical restriction enzyme profiles and shared 100% nucleotide identify over the sequenced signature region. This nucleotide sequence most closely resembled the data base sequence of Mycobacterium senegalense.

Animals↗

Nucleotide sequence analysis of the long terminal repeat of murine virus-like DNA (VL30) and its adjacent sequences: resemblance to retrovirus proviruses.

VL30 DNA represents a retrovirus-like multigene family of mice whose genetic origin is unknown. We have now determined the primary nucleotide sequences and the adjacent sequences of the long terminal direct repeats (LTRs) possessed by a randomly selected VL30 unit. The LTR of the VL30 unit comprised 435 nucleotide base pairs and had an inverted repeat of five bases at its 5' and 3' termini. At the joints with flanking mouse DNA was the VL30 sequence (5')TG . . . CA(3') and a tetranucleotide direct repeat of flanking sequences. At the inner boundary of the 5' LTR was an 18-base sequence that is complementary to tRNApro, and at the inner boundary of the 3' LTR was a purine-rich tract ending with AATG. These results suggested that VL30 DNA used the same integration strategy that is exercised by retrovirus proviruses and transposable elements and that the VL30 LTR is synthesized in a similar way that the LTR of retroviruses is synthesized. The data thus reinforce the retrovirus-like nature of VL30 genetic information.

Animals↗

Envelope gene sequence of two in vitro-generated mink cell focus-forming murine leukemia viruses which contain the entire gp70 sequence of the endogenous nonecotropic parent.

The mink cell focus-forming (MCF) class of recombinant murine leukemia viruses (CI-1 to 4) were isolated from iododeoxyuridine-induced C3H/MCA 5 cells in culture and molecularly cloned. These genomes included infectious (CI-3) and defective (CI-4) recombinants. A total of 2,408 nucleotides of CI-3 virus DNA, including the MCF envelope gene, were sequenced and compared with ecotropic, dual-tropic, and xenotropic sequences. The extent of recombinational exchange in CI-3 was from 145 nucleotides 3' of the splice acceptor site for the envelope mRNA to nucleotide 1,722, between the end of gp70 and the beginning of Prp15E. Thus, the entire gp70 sequence of the endogenous nonecotropic parent was present in this recombinant. The nature and location of the recombinant junctions were consistent with a mechanism involving DNA exchange during reverse transcription. Comparison of the substituted sequence in CI-3 with that of Moloney MCF virus suggests a very close relationship, if not identity, between the endogenous dual-tropic proviruses from which they were derived. A nonidentity of xenotropic and MCF gp70s was observed, suggesting that xenotropic murine leukemia viruses are not the nonecotropic parent of the env gene of MCF murine leukemia viruses. The replication-defective virus CI-4 had a 684-nucleotide deletion present in the env gene, eliminating the hydrophobic regions within the gp70 carboxy end and the p15E amino end. This sequence was bordered by an 11-nucleotide direct repeat in CI-3 viral DNA.

Amino Acid Sequence↗

Cloning and sequencing of defective particles derived from the autonomous parvovirus minute virus of mice for the construction of vectors with minimal cis-acting sequences.

The production of wild-type-free stocks of recombinant parvovirus minute virus of mice [MVM(p)] is difficult due to the presence of homologous sequences in vector and helper genomes that cannot easily be eliminated from the overlapping coding sequences. We have therefore cloned and sequenced spontaneously occurring defective particles of MVM(p) with very small genomes to identify the minimal cis-acting sequences required for DNA amplification and virus production. One of them has lost all capsid-coding sequences but is still able to replicate in permissive cells when nonstructural proteins are provided in trans by a helper plasmid. Vectors derived from this particle produce stocks with no detectable wild-type MVM after cotransfection with new, matched, helper plasmids that present no homology downstream from the transgene.

Base Sequence↗

RNA replication from the simian virus 5 antigenomic promoter requires three sequence-dependent elements separated by sequence-independent spacer regions.

We have previously shown for the paramyxovirus simian virus 5 (SV5) that a functional promoter for RNA replication requires proper spacing between two discontinuous elements: a 19-base segment at the 3' terminus (conserved region I [CRI]) and an 18-base internal region (CRII) that is contained within the coding region of the L protein gene. In the work described here, we have used a reverse-genetics system to determine if the 53-base segment between CRI and CRII contains additional sequence-specific signals required for optimal replication or if this segment functions solely as a sequence-independent spacer region. A series of copyback defective interfering minigenome analogs were constructed to contain substitutions of nonviral sequences in place of bases 21 to 72 of the antigenomic promoter, and the relative level of RNA replication was measured by Northern blot analysis. The results from our mutational analysis indicate that in addition to CRI and CRII, optimal replication from the SV5 antigenomic promoter requires a third sequence-dependent element located 51 to 66 bases from the 3' end of the RNA. Minigenome RNA replication was not affected by changes in the either the position of this element in relation to CRI and CRII or the predicted hexamer phase of NP encapsidation. Thus, optimal RNA replication from the SV5 antigenomic promoter requires three sequence-dependent elements, CRI, CRII and bases 51 to 66.

Base Sequence↗

DNA sequence and structure requirements for cleavage of V(D)J recombination signal sequences.

Purified RAG1 and RAG2 proteins can cleave DNA at V(D)J recombination signals. In dissecting the DNA sequence and structural requirements for cleavage, we find that the heptamer and nonamer motifs of the recombination signal sequence can independently direct both steps of the cleavage reaction. Proper helical spacing between these two elements greatly enhances the efficiency of cleavage, whereas improper spacing can lead to interference between the two elements. The signal sequences are surprisingly tolerant of structural variation and function efficiently when nicks, gaps, and mismatched bases are introduced or even when the signal sequence is completely single stranded. Sequence alterations that facilitate unpairing of the bases at the signal/coding border activate the cleavage reaction, suggesting that DNA distortion is critical for V(D)J recombination.

Animals↗

Sequencing of selected regions of the human immunoglobulin heavy-chain gene locus that completes the sequence from JH through the delta constant region.

Much of the nucleotide sequence between the start of the joining region and the end of the immunoglobulin heavy chain delta gene has already been determined. However, two gaps existed in potentially functionally important regions in this sequence: the region between the 3' end of the joining region and the heavy chain enhancer region and that between the enhancer and the mu constant region. We have determined the nucleotide sequences of these regions. The 734 bp between the joining and enhancer regions contained no additional joining regions. The 4525 bp region between the heavy chain enhancer and the mu constant region contains the mu switch region, which consists of pentameric repeats. Approximately 60% of these repeats are GGGCT and GAGCT. With the determination of these sequences, the entire region of the heavy chain locus starting upstream of the joining region to downstream of the last exon of the delta constant region (a total of more than 29 kb) has now been sequenced.

Base Sequence↗

Sequence accuracy of large DNA sequencing projects.

Very little information has been accumulated regarding the likely accuracy of final or consensus DNA sequence data. With the large-scale efforts anticipated for the Human Genome Project, the subjective determination of final sequence must eventually be replaced with more objective, automatic methods. This will require a much better understanding of the nature of error in raw sequencing data and its impact on the determination of the final sequence. This paper describes a start at defining the error model of large-scale sequencing efforts based on random subcloning strategies.

Consensus Sequence↗