Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

MRI of orofacial tumors and paragangliomas with 2D GE sequences: indications and optimal sequence parameters.

The aim of this paper is to determine to what extent and in which cases 2D gradient echo (2D GE) sequences can be applied alternatively or additionally to spin echo (SE) sequences for improved diagnostic evaluation. Imaging with SE sequences is the most frequently used MR technique in the assessment of ear, nose and throat (ENT) tumors. In literature there are only a few reports on the contrast behaviour of 2D GE sequences using different sequence parameters and their application in ENT tumors. This paper set out to establish the most suitable sequence type and sequence parameters. Measurements were performed with a Magnetom SP 63 MR system (Siemens) with a field strength of 1.5 T, using head and Helmholtz coils. One-hundred twenty-eight volunteers and 369 patients were examined with 2D GE sequences. In order to find the best MR technique for the examination of orofacial tumors and paragangliomas, FLASH, FISP and PSIF sequences with different sequence parameters (TR, TE, flip angle, bandwidth) were applied. The results of these examinations confirmed the superiority of 2D GE sequences over SE sequences. In conclusion, contrast enhanced or unenhanced T1 weighted SE sequences should be replaced by 2D GE FLASH 70 degrees in the examination of orofacial tumors. In the case of paragangliomas of the jugular bulb and the parapharyngeal space T1- and T2-weighted SE sequences should be replaced by 2D GE FLASH 40 degrees.

Head and Neck Neoplasms↗

Effect of Downstream Sequence on the Cleavage of Envelop Protein 1 Signal Sequence in Hepatitis C Virus.

The RNA genome of hepatitis C virus encodes a polyprotein of 3 000 amino acids, which is processed into 10 viral proteins by proteases provided by host cells and virus itself. Multiple precursors are produced due to inefficient processing. Here, the study of E1 signal sequence (C/E1 site) processing in eukaryotic vaccinia virus/T7 system is reported. Differently truncated HCV structural proteins were expressed in this system. It was found that the efficient cleavage of E1 signal sequence was affected by downstream envelope protein sequences. When the lacZ gene encoding a product with similar size was engineered downstream to the E1 signal sequence, the inefficient cleavage of signal sequence was also observed, suggesting that the effect of downstream sequence on the cleavage was due to the presence of the envelop protein sequences. Computer-aided analysis clearly showed that E1 signal sequences was a typical signal sequence. The influence of downstream sequences to signal sequence cleavage demonstrated here was uncommon. To date, similar observations were only reported for the processing of IL-12 signal sequence and the C/prM site of flavivirus. As both flavivirus and HCV are classified into the same Flaviviridae family, this downstream-sequence-related cleavage of signal sequence worths further studying.

Journal Article↗

[Turbo STIR sequence: optimization and comparison with conventional STIR sequence in bone diseases].

INTRODUCTION: The most common fat-suppressed sequence used to study skeletal conditions is the STIR sequence which has shown high sensitivity in the detection of skeletal lesions and whose main drawback is its long acquisition time. Currently, Turbo-STIR (T-STIR) sequences can shorten the acquisition time. The purpose of this study was therefore to compare the conventional STIR sequence with the new T-STIR sequence in the study of skeletal conditions to compare their diagnostic yield. MATERIAL AND METHODS: Twenty patients with different types of skeletal lesions were examined. MR examinations were performed with a Philips Gyroscan S15/ACS II unit (1.5 T). All the patients underwent a STIR sequence (TR/TE = 1500/20, TI = 180 ms, matrix = 204 x 256, NEX = 2, slice thickness = 5 mm, acquisition time = 9 min 24 s) and a T-STIR sequence (TR/TE = 1500/20, TI = 180 ms, matrix = 204 x 256, NEX = 2, slice thickness = 5 mm, TFL = 3, acquisition time = 3 min 33 s). The images were evaluated by measuring both quantitative parameters--percent contrast (%C), contrast to noise ratio (C/N), signal to noise ratio (S/N)--and qualitative parameters--lesion conspicuity, margins and extension, motion artifacts, image quality. RESULTS: The only statistically significant difference between the two sequences was image quality, which was superior in the conventional STIR sequence (p < .05). No statistically significant difference was demonstrated with the quantitative evaluation. DISCUSSION: In this study, T-STIR sequences were performed with low-high acquisition profile to acquire an actual echo time of 20 ms which permits to obtain optimal S/N with good spatial resolution. Therefore, T-STIR sequences with low-high acquisition profile provides better results than T-STIR sequences with linear acquisition profile which permits to obtain an actual echo time of 40 ms. CONCLUSION: This work shows that T-STIR sequences can replace conventional STIR sequences in the study of skeletal conditions reducing the acquisition time by 60%. This result can be obtained only by an accurate optimization of acquisition parameters.

Adolescent↗

Sequence determination and analysis of the 3' region of chicken pro-alpha 1(I) and pro-alpha 2(I) collagen messenger ribonucleic acids including the carboxy-terminal propeptide sequences.

Three pro-alpha 1 collagen cDNA clones, pCg1, pCg26, and pCg54, and two pro-alpha 2 collagen cDNA clones, pCg 13 and pCg45, were subjected to extensive DNA sequence determination. The combined sequences specified the amino acid sequences for chicken pro-alpha 1 and pro-alpha 2 type I collagens starting at residue 814 in the collagen triple-helical region and continuing to the procollagen C-termini as determined by the first in-phase termination codon. Thus, the sequences of 272 pro-alpha 1 C-terminal, 260 pro-alpha 2 C-terminal, 201 pro-alpha 1 helical, and 201 pro-alpha 2 helical amino acids were established. In addition, the sequences of several hundred nucleotides corresponding to noncoding regions of both procollagen mRNAs were determined. In total, 1589 pro-alpha 1 base pairs and 1691 pro-alpha 2 base pairs were sequenced, corresponding to approximately one-third of the total length of each mRNA. Both procollagen mRNA sequences have a high G+C content. The pro-alpha 1 mRNA is 75% G+C in the helical coding region sequenced and 61% G&C in the C-terminal coding region while the pro-alpha 2 mRNA is 60% and 48% G+C, respectively, in these regions. The dinucleotide sequence pCG occurs at a higher frequence in both sequences than is normally found in vertebrate DNAs and is approximately 5 times more frequent in the pro-alpha 1 sequence than in the pro-alpha 2 sequence. Nucleotide homology in the helical coding regions is very limited given that these sequences code for the repeating Gly-X-Y tripeptide in a region where X and Y residues are 50% conserved. These differences are clearly reflected in the preferred codon usages of the two mRNAs.

Amino Acid Sequence↗

Detection of sequence variants in the gene for human type II procollagen (COL2A1) by direct sequencing of polymerase chain reaction-amplified genomic DNA.

The direct sequencing of the human type II procollagen (COL2A1) gene from polymerase chain reaction (PCR)-amplified genomic DNA is described. Thirty-two regions of the COL2A1 gene were asymmetrically amplified with intron primers which were specifically chosen to amplify a region spanning 500 to 800 bp of sequence encoding one or more exons and their accompanying intervening sequences. Primers for dideoxynucleotide sequencing of the PCR products were then designed to provide complete exon sequence information and to insure that intron:exon splice junction sequence data would be obtained. Amplification and sequencing reactions were performed on an automated workstation to facilitate the handling of multiple DNA templates. The procedure allowed efficient sequencing of over 25,000 bp of each allele of the COL2A1 gene per diploid genome. We used this method for the comparative analyses of COL2A1 sequences in DNA isolated from the blood of 42 unrelated individuals and we identified 21 neutral sequence variants in the gene. The sequence variations were confirmed by independent assays, including restriction enzyme digestion. The sequence variants described here will be important for identifying haplotypes of the type II procollagen gene that will be useful in defining a genetic etiology for diseases of cartilaginous tissues.

Alleles↗

A reassessment of mammalian alpha A-crystallin sequences using DNA sequencing: implications for anthropoid affinities of tarsier.

alpha A-crystallin, a major structural protein in the ocular lenses of all vertebrates, has been a valuable tool for molecular phylogenetic studies. This paper presents the complete sequence for human alpha A-crystallin derived from cDNA and genomic clones. The deduced amino acid sequence differs at two phylogenetically informative positions from that previously inferred from peptide composition. This led us to examine the same region of the alpha A-crystallin gene in 12 other mammalian species using direct sequencing of PCR-amplified genomic DNA. New sequences were added to the database, and corrections were made to all anthropoid sequences, defining clear synapomorphies for anthropoids as a clade distinct from prosimians. Within the anthropoids there are further synapomorphies delineating hominoids, Old World monkeys, and New World monkeys. Significantly, sequence revisions and the addition of new sequence for a prosimian, the sifaka, eliminate the previous support for the proposed anthropoid affinities of the tarsier inferred from alpha A-crystallin protein sequences. In addition, DNA sequences provide greater resolution of certain relationships. For example, although they are identical in protein sequence, comparison of DNA sequences clearly separates mouse and the common tree shrew, grouping the tree shrew closer to prosimians. These results show that adding DNA sequences to the existing alpha A-crystallin database can enhance its value in resolving phylogenetic relationships.

Amino Acid Sequence↗

ProtEST: protein multiple sequence alignments from expressed sequence tags.

MOTIVATION: An automatic sequence searching method (ProtEST) is described which constructs multiple protein sequence alignments from protein sequences and translated expressed sequence tags (ESTs). ProtEST is more effective than a simple TBLASTN search of the query against the EST database, as the sequences are automatically clustered, assembled, made non-redundant, checked for sequence errors, translated into protein and then aligned and displayed. RESULTS: A ProtEST search found a non-redundant, translated, error- and length-corrected EST sequence for > 58% of sequences when single sequences from 1407 Pfam-A seed alignments were used as the probe. The average family size of the resulting alignments of translated EST sequences contained > 10 sequences. In a cross-validated test of protein secondary structure prediction, alignments from the new procedure led to an improvement of 3.4% average Q3 prediction accuracy over single sequences. AVAILABILITY: The ProtEST method is available as an Internet World Wide Web service http://barton.ebi.ac.uk/servers/protest.html+ ++ The Wise2 package for protein and genomic comparisons and the ProtESTWise script can be found at http://www.sanger.ac.uk/Software/Wise2 CONTACT: geoff@ebi.ac.uk

Amino Acid Sequence↗

Sequence evolution and phylogenetic signal in control-region and cytochrome b sequences of rainbow fishes (Melanotaeniidae).

The nucleotide sequences of segments of the cytochrome b gene (351 bp), the tRNA(Pro) gene (49 bp), and the control region (approximately 313 bp) of mitochondrial DNA were obtained from 26 fish representing different populations and species of Melanotaenia and one species of Glossolepis, freshwater rainbow fishes confined to Australia and New Guinea. The purpose was to investigate relative rates and patterns of sequence evolution. Overall levels of divergence were similar for the cytochrome b and tRNA control-region sequences, both ranging from < 1% within subspecies to 15%-19% between genera. However, the patterns of sequence evolution differed. For the cytochrome b gene, transitions consistently exceeded transversions, the bias ranging from 4.2:1 to 2:1, depending on the level of sequence divergence. However, in the control-region sequence, a bias toward transitions (2:1) was observed only in comparisons between very similar sequences, and transversions outnumbered transitions in comparisons of divergent sequences. Graphic comparisons suggested that the control region was saturated for transitions at relatively low levels of sequence divergence but accumulated transversions at a greater rate than did the cytochrome b sequence. These distinct patterns of base substitution are associated with differences in A+T content, which is 70% for the tRNA control-region segment versus 50% for cytochrome b. A test for skewness in the distribution of lengths of random trees indicated that both segments contained phylogenetic signal. Parsimony analyses of the data from the two regions, with or without weighting schemes appropriate to the respective patterns of sequence evolution, identified the same five groupings of sequences, but the relationships among the groups differed. However, in most cases the branches uniting different combinations of groups were poorly supported, and the differences among topologies were insignificant. Considering the observed patterns of base substitution and the results of the phylogenetic analyses, we deduce that both the control region and cytochrome b are appropriate for population genetic studies but that the control region is less effective than cytochrome b for resolving relationships among divergent lineages of rainbow fishes.

Animals↗

The carboxylesterase family exhibits C-terminal sequence diversity reflecting the presence or absence of endoplasmic-reticulum-retention sequences.

Resident proteins of the endoplasmic reticulum lumen are continuously retrieved from an early Golgi compartment by a receptor-mediated mechanism. The sorting or retention sequence on the endoplasmic reticulum proteins is located at the C-terminus and was initially shown to be the tetrapeptide KDEL in mammalian cells and HDEL in Saccharomyces cerevisiae. The carboxylesterases are a large family of enzymes primarily localized to the lumen of the endoplasmic reticulum. Retention sequences in these proteins have been difficult to identify due to atypical and heterogeneous C-terminal sequences. Utilizing the polymerase chain reaction with degenerate primers, we have identified and characterized the C-termini of four members of the carboxylesterase family from rat liver. Three of the carboxylesterases sequences contained C-terminal sequences (HVEL, HNEL or HTEL) resembling the yeast sorting signal which were reported to be non-functional in mammalian cells. A fourth carboxylesterase contained a distinct C-terminal sequence, TEHT. A full-length esterase cDNA clone, terminating in the sequence HVEL, was isolated and was used to assess the retention capabilities of the various esterase C-terminal sequences. This esterase was retained in COS-1 cells, but was secreted when its C-terminal tetrapeptide, HVEL, was deleted. Addition of C-terminal sequences containing HNEL and HTEL resulted in efficient retention. However, the C-terminal sequence containing TEHT was not a functional retention signal. Both HDEL, the authentic yeast retention signal, and KDEL were efficient retention sequences for the esterase. These studies show that some members of the rat liver carboxylesterase family contain novel C-terminal retention sequences that resemble the yeast signal. At least one member of the family does not contain a C-terminal retention signal and probably represents a secretory form.

Amino Acid Sequence↗

Pairwise end sequencing: a unified approach to genomic mapping and sequencing.

Strategies for large-scale genomic DNA sequencing currently require physical mapping, followed by detailed mapping, and finally sequencing. The level of mapping detail determines the amount of effort, or sequence redundancy, required to finish a project. Current strategies attempt to find a balance between mapping and sequencing efforts. One such approach is to employ strategies that use sequence data to build physical maps. Such maps alleviate the need for prior mapping and reduce the final required sequence redundancy. To this end, the utility of correlating pairs of sequence data derived from both ends of subcloned templates is well recognized. However, optimal strategies employing such pairwise data have not been established. In the present work, we simulate and analyze the parameters of pairwise sequencing projects including template length, sequence read length, and total sequence redundancy. One pairwise strategy based on sequencing both ends of plasmid subclones is recommended and illustrated with raw data simulations. We find that pairwise strategies are effective with both small (cosmid) and large (megaYAC) targets and produce ordered sequence data with a high level of mapping completeness. They are ideal for finescale mapping and gene finding and as initial steps for either a high- or a low-redundancy sequencing effort. Such strategies are highly automatable.

Base Composition↗

Sequence-tagged connectors: a sequence approach to mapping and scanning the human genome.

The sequence-tagged connector (STC) strategy proposes to generate sequence tags densely scattered (every 3.3 kilobases) across the human genome by arraying 450,000 bacterial artificial chromosomes (BACs) with randomly cleaved inserts, sequencing both ends of each, and preparing a restriction enzyme fingerprint of each. The STC resource, containing end sequences, fingerprints, and arrayed BACs, creates a map where the interrelationships of the individual BAC clones are resolved through their STCs as overlapping BAC clones are sequenced. Once a seed or initiation BAC clone is sequenced, the minimum overlapping 5' and 3' BAC clones can be identified computationally and sequenced. By reiterating this "sequence-then-map by computer analysis against the STC database" strategy, a minimum tiling path of clones can be sequenced at a rate that is primarily limited by the sequencing throughput of individual genome centers. As of February 1999, we had deposited, together with The Institute for Genomic Research (TIGR), into GenBank 314,000 STCs ( approximately 135 megabases), or 4.5% of human genomic DNA. This genome survey reveals numerous genes, genome-wide repeats, simple sequence repeats (potential genetic markers), and CpG islands (potential gene initiation sites). It also illustrates the power of the STC strategy for creating minimum tiling paths of BAC clones for large-scale genomic sequencing. Because the STC resource permits the easy integration of genetic, physical, gene, and sequence maps for chromosomes, it will be a powerful tool for the initial analysis of the human genome and other complex genomes.

Chromosome Mapping↗

Sequence arrangement in herpes simplex virus type 1 DNA: identification of terminal fragments in restriction endonuclease digests and evidence for inversions in redundant and unique sequences.

It has been proposed by Sheldrick and Berthelot (1974) that the terminal sequences of herpes simplex virus type 1 (HSV-1) DNA are repeated in an internal inverted form and that the inverted redundant sequences delimit and separate two unique sequences, S and L. In this study the sequence arrangement in HSV-1 DNA has been investigated with restriction endonuclease cleavage, end-labeling studies, and molecular hybridization experiments. The terminal fragments in digests with restriction endonucleases Hind III, Hpa-1, EcoRI and Bum were identified and shown to be consistent with the Sheldrick and Berthelot model. Inverted fragments which contain unique sequences as well as redundant sequences, and which the model predicts, were identified by DNA-DNA hybridization studies. Further cleavage of Bum fragments with Hpa-1 also revealed inversions of the terminal sequences that contained unique sequences. The results obtained showed that the unique sequences S and L are relatively inverted in different DNA molecules in the population, resulting in the presence of four related genomes with rearranged sequences in apparently equal amounts. The redundant sequences bounding S do not share complete sequence homology with those bounding L, but hybridization studies are presented which show that the terminal 0.3% of the genome is repeated in every redundant sequence.

Base Sequence↗

The structure of the human apolipoprotein C-II gene. Electron microscopic analysis of RNA:DNA hybrids, complete nucleotide sequence, and identification of 5' homologous sequences among apolipoprotein genes.

Cloned human apo-C-II cDNA was used as a hybridization probe to identify the human apo-C-II gene in a genomic library constructed in our laboratory. The isolated apo-C-II DNA was studied both by electron microscopy and by direct sequence analysis. Ultrastructural morphological analysis of RNA-DNA hybrids revealed that the apo-C-II gene had complex structures because of regions of inverted complementary sequences in and around the gene forming stem-and-loop structures which interfere with the formation of stable RNA:DNA hybrids. Extensive morphological analysis revealed a minimum of 3 intervening sequences (IVS), and their lengths were measured. Direct sequence analysis of the cloned gene confirmed the presence of 3 IVS. There are 4 Alu type sequences in IVS-I. We sequenced 4340 nucleotides which include 545 nucleotides in the 5' flanking region, the entire gene which spans 3320 nucleotides, and 475 nucleotides in the 3' flanking region which also encompasses an additional Alu sequence. The 5' end of the gene was identified by primer extension and sequencing of the primer extended cDNA. Apo-C-II mRNA structure was deduced from the cDNA sequence, the primer extension experiments, and the genomic sequence. It is 494 nucleotides in length. Its sequence differs from previously published sequences in that there are 7 additional nucleotides before the polyadenylate tail. In the 5' flanking region, nucleotides -234 to -213 encompass a GC-rich region which exhibits high homology (greater than 70%) to the 5' flanking regions of the genes of all the apolipoproteins published to date, namely, apo-A-II (-497 to -471), apo-A-I (approximately -196 to -179), apo-E (-409 to -391), and apo-C-III (approximately -116 to -103). This highly conserved region might represent some evolutionarily conserved sequences from these related genes and/or might represent a region with regulatory function.

Apolipoprotein C-II↗

Accuracy of automated DNA sequencing: a multi-laboratory comparison of sequencing results.

A double-stranded (ds)DNA template of "unknown" sequence was distributed to approximately 80 core DNA sequencing laboratories by the Association of Biomolecular Resource Facilities (ABRF) for automated DNA sequence analysis. Forty-four different facilities responded with 83 usable sequence submissions. These sequences were grouped by both sequencing protocol (dye-primer or dye-terminator) and whether manually edited or not. The sequences were aligned with the known sequence, and the number of correct base calls, insertions, deletions, no-calls and miscalls were determined for each group. The dye-primer sequencing protocol provided the longest and most accurate sequence. The edited dye-primer data were > 95% accurate out to 400-450 bp, while the edited dye-terminator data could call only 300-350 bases at this accuracy. However, 75% of the laboratories in this sampling preferred the dye-terminator protocol, presumably because of its versatility and convenience. Laboratories that manually edited the automatically called data were able to obtain an additional 100 bases of good sequence when the dye-primer protocol was used. Surprisingly though, editing of dye-terminator results did not increase the amount of good sequence, although the dye-terminator protocol had a superior base-calling ability within the first 100 bases of called sequence.

Autoanalysis↗

Method for prediction of protein function from sequence using the sequence-to-structure-to-function paradigm with application to glutaredoxins/thioredoxins and T1 ribonucleases.

The practical exploitation of the vast numbers of sequences in the genome sequence databases is crucially dependent on the ability to identify the function of each sequence. Unfortunately, current methods, including global sequence alignment and local sequence motif identification, are limited by the extent of sequence similarity between sequences of unknown and known function; these methods increasingly fail as the sequence identity diverges into and beyond the twilight zone of sequence identity. To address this problem, a novel method for identification of protein function based directly on the sequence-to-structure-to-function paradigm is described. Descriptors of protein active sites, termed "fuzzy functional forms" or FFFs, are created based on the geometry and conformation of the active site. By way of illustration, the active sites responsible for the disulfide oxidoreductase activity of the glutaredoxin/thioredoxin family and the RNA hydrolytic activity of the T1 ribonuclease family are presented. First, the FFFs are shown to correctly identify their corresponding active sites in a library of exact protein models produced by crystallography or NMR spectroscopy, most of which lack the specified activity. Next, these FFFs are used to screen for active sites in low-to-moderate resolution models produced by ab initio folding or threading prediction algorithms. Again, the FFFs can specifically identify the functional sites of these proteins from their predicted structures. The results demonstrate that low-to-moderate resolution models as produced by state-of-the-art tertiary structure prediction algorithms are sufficient to identify protein active sites. Prediction of a novel function for the gamma subunit of a yeast glycosyl transferase and prediction of the function of two hypothetical yeast proteins whose models were produced via threading are presented. This work suggests a means for the large-scale functional screening of genomic sequence databases based on the prediction of structure from sequence, then on the identification of functional active sites in the predicted structure.

Algorithms↗

Sequence analysis of homogeneous peptides of shark immunoglobulin light chains by tandem mass spectrometry: correlation with gene sequence and homologies among variable and constant region peptides of sharks and mammals.

Morphologically, sharks are living fossils that are remarkably similar to their Devonian ancestors of ca. 400 million years ago. If a parallel conservation in biochemical properties characterizes shark evolution, knowledge of the properties of shark immunoglobulins should provide information on the structure of primordial immunoglobulins and their genes. The problem of polyclonality of shark immunoglobulins has precluded detailed analysis of shark immunoglobulin light polypeptide chains. Here, we approach the problem of obtaining direct sequence information on polyclonal light chains of shark immunoglobulins by isolating homogeneous peptides from tryptic digests of shark light chains and sequencing these by tandem mass spectrometry. To confirm the location of the peptides, we isolated a complementary DNA (cDNA) clone from a sandbar shark cDNA library in the expression vector lambda gt11, identifying the clone by its ability to produce a peptide serologically detectable using rabbit antibody to purified shark light chain. The correspondence between peptide sequence and that derived from gene sequence provided direct proof that the gene studied was that of a major expressed serum light chain. Using this combined approach, we isolated homogeneous peptides from both constant and variable regions. The variable region peptides showed homology to corresponding sequences of mammalian V lambda and V kappa sequences. The constant region gene sequence we obtained was homologous to mammalian C lambda sequence. The four constant region tryptic peptides we sequenced corresponded exactly to stretches of the C lambda sequence derived from the DNA sequence. The combined approach described here shows that shark light chains exhibit heterogeneity at both the protein and gene level, but that the constant regions of these chains can be identified as homologs of mammalian lambda chains and that evolutionary conservation has occurred in V region sequences ranging from elasmobranchs to man.

Amino Acid Sequence↗

A method of sequencing without subcloning and its application to the identification of a novel ORF with a sequence suggestive of a transcriptional regulator in the water mold Achlya ambisexualis.

Genomic amplification with transcript sequencing (GAWTS) is a method of direct sequencing that involves amplification with PCR using primers containing phage promoters, transcription of the amplified product, and sequencing with reverse transcriptase. GAWTS requires the generation of PCR primers that are specific for the sequences on both sides of a region. Here we describe promoter ligation and transcript sequencing (PLATS), a direct method for rapidly obtaining novel sequences that utilizes generic primers and only requires knowledge of the sequence on one side of a region. PLATS involves restriction digestion of the amplified vector insert, ligation with a phage promoter, and then GAWTS using phage promoter sequences as the PCR primers. The method is rapid and economical because it uses a limited set of oligonucleotides, and it is potentially amenable to automation because it does not require in vivo manipulations. PLATS facilitates the determination of a genomic sequence responsible for cross-hybridization in a Southern blot. Using PLATS, sequence has been obtained from a 1.1-kb segment in Achlya ambisexualis, which cross-hybridizes to the DNA-binding region of the chicken and Xenopus estrogen receptors. To our knowledge, this represents the first sequence reported from the Oomycetes, a large and widely distributed group of fungi. The sequence reveals a large, transcribed open reading frame that is markedly deficient in the dinucleotide TpA. A putative zinc finger containing three cysteines and one histidine (C-X2-C-X12-H-X3-C) and an acidic segment hint that this clone may be a member of a novel class of transcriptional regulators.

Amino Acid Sequence↗

Heterogeneity of hepatitis B virus C-gene sequences: implications for amplification and sequencing.

Occasionally direct sequencing of amplified hepatitis B virus DNA leads to weak signals on autoradiograms. Using amplified C-gene sequences we investigated whether this is due to sequence heterogeneity of virus populations and use of inappropriate primers for direct sequencing. High C-gene sequence heterogeneity (point mutations, stop codons and a one codon deletion) was observed in HBV genomes from serum of a chronic carrier who underwent interferon treatment. The type of C-gene mutations detected by direct sequencing depended on the type of primers used. Cloning and sequencing of amplified C-gene sequences demonstrated that this was due to mutations in the region complementary to the sequencing primer. These data demonstrate the existence of novel HBV C-gene mutants and imply that multiple or degenerate sequencing and amplification primers are essential for accurate evaluation of the extent of HBV C-gene heterogeneity. Based on comparative sequence analysis of all available completely or incompletely sequenced C-genes, guidelines for optimal primer design are proposed for similar studies.

Adult↗