Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Extensive sequence-specific information throughout the CAR/RRE, the target sequence of the human immunodeficiency virus type 1 Rev protein.

The significance and location of sequence-specific information in the CAR/RRE, the target sequence for the Rev protein of the human immunodeficiency virus type 1 (HIV-1), have been controversial. We present here a comprehensive experimental and computational approach combining mutational analysis, phylogenetic comparison, and thermodynamic structure calculations with a systematic strategy for distinguishing sequence-specific information from secondary structural information. A target sequence analog was designed to have a secondary structure identical to that of the wild type but a sequence that differs from that of the wild type at every position. This analog was inactive. By exchanging fragments between the wild-type sequence and the inactive analog, we were able to detect an unexpectedly extensive distribution of sequence specificity throughout the CAR/RRE. The analysis enabled us to identify a critically important sequence-specific region, region IIb in the Rev-binding domain, strongly supports a proposed base-pairing interaction in this location, and places forceful constraints on mechanisms of Rev action. The generalized approach presented can be applied to other systems.

Base Sequence↗

Fitness of a turnip crinkle virus satellite RNA correlates with a sequence-nonspecific hairpin and flanking sequences that enhance replication and repress the accumulation of virions.

satC, a satellite RNA associated with Turnip crinkle virus (TCV), enhances the ability of the virus to colonize plants by interfering with stable virion accumulation (F. Zhang and A. E. Simon, unpublished data). Previous results suggested that the motif1-hairpin (M1H), a replication enhancer on minus strands, forms a plus-strand hairpin flanked by CA-rich sequence that may be involved in enhancing systemic infection (G. Zhang and A. E. Simon, J. Mol. Biol. 326:35-48, 2003). In this study, sequence and structural requirements of the M1H were further assayed by replacing the 28-base M1H with 10 random bases and then subjecting the pool of satellite RNA to functional selection in plants. Unlike previous results with 28-base replacement sequences (G. Zhang and A. E. Simon, J. Mol. Biol. 326:35-48, 2003), only a few of the 10-base SELEX (systematic evolution of ligands by exponential enrichment) assay winners contained short motifs in their minus-sense orientation that were similar to TCV replication elements. However, all second- and third-round winning replacement sequences folded into hairpins flanked by CA-rich sequence predicted to be more stable on plus strands than minus strands. Plus strands of several of the most fit satellite RNAs contained insertions of CA-rich sequence at the base of their hairpins whose presence correlated with enhanced replication and reduced detection of virions. Deletion of the M1H resulted in no detectable virions despite very low satellite accumulation. These results support the hypothesis that a sequence-nonspecific plus-strand hairpin brings together flanking CA-rich sequences in the M1H region that confers fitness to satC by reducing the accumulation of stable virions.

Base Sequence↗

Finding new human minisatellite sequences in the vicinity of long CA-rich sequences.

Microsatellites and minisatellites are two classes of tandem repeat sequences differing in their size, mutation processes, and chromosomal distribution. The boundary between the two classes is not defined. We have developed a convenient, hybridization-based human library screening procedure able to detect long CA-rich sequences. Analysis of cosmid clones derived from a chromosome 1 library show that cross-hybridizing sequences tested are imperfect CA-rich sequences, some of them showing a minisatellite organization. All but one of the 13 positive chromosome 1 clones studied are localized in chromosomal bands to which minisatellites have previously been assigned, such as the 1pter cluster. To test the applicability of the procedure to minisatellite detection on a larger scale, we then used a large-insert whole-genome PAC library. Altogether, 22 new minisatellites have been identified in positive PAC and cosmid clones and 20 of them are telomeric. Among the 42 positive PAC clones localized within the human genome by FISH and/or linkage analysis, 25 (60%) are assigned to a terminal band of the karyotype, 4 (9%) are juxtacentromeric, and 13 (31%) are interstitial. The localization of at least two of the interstitial PAC clones corresponds to previously characterized minisatellite-containing regions and/or ancestrally telomeric bands, in agreement with this minisatellite-like distribution. The data obtained are in close agreement with the parallel investigation of human genome sequence data and suggest that long human (CA)s are imperfect CA repeats belonging to the minisatellite class of sequences. This approach provides a new tool to efficiently target genomic clones originating from subtelomeric domains, from which minisatellite sequences can readily be obtained. [The sequence data described in this paper have been submitted to the EMBL data library under accession nos. AJ000377-AJ000383.]

Bacteriophage P1↗

Detection of DNA sequence polymorphisms by enzymatic amplification and direct genomic sequencing.

The discovery of RFLPs and their utilization as genetic markers has revolutionized research in human molecular genetics. However, only a fraction of the DNA sequence polymorphisms in the human genome affect the length of a restriction fragment and hence result in an RFLP. Polymorphisms that are not detected as RFLPs are typically passed over in the screening process though they represent a potentially important source of informative genetic markers. We have used a rapid method for the detection of naturally occurring DNA sequence variations that is based on enzymatic amplification and direct sequencing of genomic DNA. This approach can detect essentially all useful sequence variations within the region screened. We demonstrate the feasibility of the technique by applying it to the human retinoblastoma susceptibility locus. We screened 3,712 bp of genomic DNA from each of nine individuals and found four DNA sequence polymorphisms. At least one of these DNA sequence polymorphisms was informative in each of three families with hereditary retinoblastoma that were not informative with any of the known RFLPs at this locus. We believe that direct sequencing is a reasonable alternative to other methods of screening for DNA sequence polymorphisms and that it represents a step forward for obtaining informative markers at well-characterized loci that have been minimally informative in the past.

Base Sequence↗

Complete nucleotide sequence of cDNA and deduced amino acid sequence of rat liver arginase.

Arginase (EC 3.5.3.1) catalyzes the last step of urea synthesis in the liver of ureotelic animals. The nucleotide sequence of rat liver arginase cDNA, which was isolated previously (Kawamoto, S., Amaya, Y., Oda, T., Kuzumi, T., Saheki, T., Kimura, S., and Mori, M. (1986) Biochem. Biophys. Res. Commun. 136, 955-961) was determined. An open reading frame was identified and was found to encode a polypeptide of 323 amino acid residues with a predicted molecular weight of 34,925. The cDNA included 26 base pairs of 5'-untranslated sequence and 403 base pairs of 3'-untranslated sequence, including 12 base pairs of poly(A) tract. The NH2-terminal amino acid sequence, and the sequences of two internal peptide fragments, determined by amino acid sequencing, were identical to the sequences predicted from the cDNA. Comparison of the deduced amino acid sequence of the rat liver arginase with that of the yeast enzyme revealed a 40% homology.

Amino Acid Sequence↗

Purification, cloning and nucleotide sequence determination of cynomolgus monkey apolipoprotein C-II: comparison to the human sequence.

We have purified apolipoprotein C-II (apo C-II) from cynomolgus monkey plasma, prepared antibody against it and used the antibody to isolate a cDNA containing the complete coding sequence for cynomolgus monkey apo C-II. Sequence analysis indicated that the monkey apo C-II cDNA was 200 bp longer than the human and the difference in size was all in the 5 degrees untranslated region of mRNA. This was confirmed by Northern analysis of human and monkey RNA. There was an open reading frame in the monkey apo C-II cDNA sequence encoding a preprotein of 101 amino acids - identical in size to the human protein. The carboxyl terminal 44 amino acids of the protein were 100% homologous to the human apo C-II amino acid sequence indicating evolutionary conservation of both structure and function. However, the amino terminal 35 amino acids of the protein were only 75% homologous and the amino terminal 19 amino acids were only 58% homologous to the human sequence. The amino acid sequence derived from the nucleotide sequence predicts a more basic protein than the human apo C-II and this is confirmed by isoelectric focusing and immunoblotting.

Amino Acid Sequence↗

Gene sequence for the 9 kDa component of Photosystem II from the cyanobacterium Phormidium laminosum indicates similarities between cyanobacterial and other leader sequences.

A 9 kDa polypeptide which is loosely attached to the inner surface of the thylakoid membrane and is important for the oxygen-evolving activity of Photosystem II in the thermophilic cyanobacterium Phormidium laminosum has been purified, a partial amino acid sequence obtained and its gene cloned and sequenced. The derived amino acid sequence indicates that the 9 kDa polypeptide is initially synthesised with an N-terminal leader sequence of 44 amino acids to direct it across the thylakoid membrane. The leader sequence consists of a positively charged N-terminal region, a long hydrophobic region and a typical cleavage site. These features have analogous counterparts in the "thylakoid-transfer domain" of lumenal polypeptides from chloroplasts of higher plants. These findings support the view of the proposed function of this domain in the two-stage processing model for import of lumenal, nuclear-encoded polypeptides. In addition, there is striking primary sequence homology between the leader sequences of the 9 kDa polypeptide and those of alkaline phosphatase (from the periplasmic space of Escherichia coli) and, particularly in the region of the cleavage site, the 16 kDa polypeptide of the oxygen-evolving apparatus in the thylakoid lumen of spinach chloroplasts.

Amino Acid Sequence↗

Heptad motifs within the distal subdomain of the coiled-coil rod region of M protein from rheumatic fever and nephritis associated serotypes of group A streptococci are distinct from each other: nucleotide sequence of the M57 gene and relation of the deduced amino acid sequence to other M proteins.

Streptococcal M protein, a dimeric alpha helical coiled-coil molecule, is an antigenically variable virulence factor on the surface of the bacteria. Our recent conformational analysis of the complete sequence of the M6 protein led us to propose a basic model for the M protein consisting of an extended central coiled-coil rod domain flanked by a variable N-terminal and a conserved C-terminal end domains. The central coiled-coil rod domain of M protein, which constitutes the major part of the M molecule, is made up of repeating heptads of the generalized sequence a-b-c-d-e-f-g, wherein "a" and "d" are predominantly apolar residues. Based on the differences in the heptad pattern of apolar residues and internal sequence homology, the central coiled-coil rod domain of M protein could be further divided into three subdomains I, II, and III. The streptococcal sequelae rheumatic fever (RF) and acute glomerulonephritis (AGN) have been known to be associated with distinct serotypes. Consistent with this, we observed that the AGN associated M49 protein exhibits a heptad motif that is distinct from the RF associated M5 and M6 proteins. Asn and Leu predominated in the "a" and "d" positions, respectively, in subdomain I of the M5 and M6 proteins, whereas apolar residues predominated in both these positions in the M49 protein. To establish whether the heptad motif of M49 is unique to this protein, or is a general characteristic of nephritis-associated serotypes, the amino acid sequence of M57, another nephritis-associated serotype, has now been examined. The gene encoding M57 was amplified by PCR, cloned into pUC19 vector, and sequenced. The C-terminal half of M57 is highly homologous to other M proteins (conserved region). In contrast, its N-terminal half (variable region) revealed no significant homology with any of the M proteins. Heptad periodicity analysis of the M57 sequence revealed that the basic design principles, consisting of distinct domains observed in the M6 protein, are also conserved in the M57 molecule. However, the heptad motif within the coiled-coil subdomain I of M57 was distinct from M5 and M6 but similar to M49. Similar analyses of the heptad characteristics within the reported sequences of M1, M12, and M24 proteins further confirmed the conservation of the overall architectural design of sequentially distinct M proteins.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Derivation of the sequence of the signal peptide in human C4b-binding protein and interspecies cross-hybridisation of the C4bp cDNA sequence.

A 5' cDNA clone coding for human C4b-binding protein (C4bp) was isolated, characterised and sequenced to complete the cDNA sequence coding for residues 1-32 thus confirming the protein sequence data of Chung et al. [(1985) Biochem. J. 230, 133-141]. The sequence extended to allow derivation of the putative leader peptide sequence which was 32 residues in length and showed a high of hydrophobicity typical of other documented leader sequences. Cross hybridisation was detected between the human C4bp cDNA probes and genomic DNA isolated from various species on Southern blots suggesting that genomic sequence homologous to that coding for C4bp has been conserved during evolution.

Amino Acid Sequence↗

The epsilon-subunit of ATP synthase from bovine heart mitochondria. Complementary DNA sequence, expression in bovine tissues and evidence of homologous sequences in man and rat.

The epsilon-subunit of ATP synthase from bovine heart mitochondria is assembled into the extrinsic membrane sector, F1-ATPase. The mature protein is 50 amino acid residues in length and its function is unknown. It is a nuclear gene product that is imported into the organelle. A mixture of 64 oligonucleotides 17 bases long, designed on the basis of the known protein sequence, was synthesized and used as a hybridization probe to isolate a cognate cDNA clone from a bovine library. The DNA sequence of this clone was determined, and the protein sequence of the epsilon-subunit deduced from it agrees exactly with that determined by direct sequence analysis of the protein isolated from bovine hearts. The bovine cDNA was used as a hybridization probe to examine the expression of the epsilon-subunit in various bovine tissues. mRNAs related to the cDNA are found in all of these tissues, and no evidence was obtained of the presence of mRNAs for the epsilon-subunit with similar coding sequences and dissimilar 3' non-coding regions. By hybridization experiments with digests of DNA from cow, man and rat it has been shown that sequences related to the bovine cDNA are present in the genomes of all three species. More than one related sequence was detected in all cases, indicating the presence in all three genomes of more than one gene and/or pseudogenes.

Amino Acid Sequence↗

Cloning and sequencing of Octopus dofleini hemocyanin cDNA: derived sequences of functional units Ode and Odf.

A number of additional cDNA clones coding for portions of the very large polypeptide chain of Octopus dofleini hemocyanin were isolated and sequenced. These data reveal two very similar coding sequences, which we have denoted "A-type" and "G-type." We have obtained complete A-type sequences coding for functional units Ode and Odf; consequently a total of three such unit sequences are now known from a single subunit of one molluscan hemocyanin. This presents the opportunity to make sequence comparisons within one hemocyanin subunit. Domains within one subunit show on the average 42% identity in amino acid residues; corresponding functional units from hemocyanins of different species show degrees of identity of 53-75%. Therefore, molluscan hemocyanins already existed before the individual molluscan classes diverged in the early Cambrian. Sequence comparisons of molluscan hemocyanins with arthropodan hemocyanins and tyrosinases allow us to identify the ligands of the "Copper B" site with high probability. Possible ligands for the "Copper A" site are proposed, based on sequence comparisons between molluscan hemocyanins and tyrosinases. Besides two histidine side chains, a methionine side chain might be involved in binding of Copper A, a result not in conflict with spectroscopic studies.

Amino Acid Sequence↗

Precise sequence complementarity between yeast chromosome ends and two classes of just-subtelomeric sequences.

The terminal regions (last 20 kb) of Saccharomyces cerevisiae chromosomes universally contain blocks of precise sequence similarity to other chromosome terminal regions. The left and right terminal regions are distinct in the sense that the sequence similarities between them are reverse complements. Direct sequence similarity occurs between the left terminal regions and also between the right terminal regions, but not between any left ends and right ends. With minor exceptions the relationships range from 80% to 100% match within blocks. The regions of similarity are composites of familiar and unfamiliar repeated sequences as well as what could be considered "single-copy" (or better "two-copy") sequences. All terminal regions were compared with all other chromosomes, forward and reverse complement, and 768 comparisons are diagrammed. It appears there has been an extensive history of sequence exchange or copying between terminal regions. The subtelomeric sequences fall into two classes. Seventeen of the chromosome ends terminate with the Y' repeat, while 15 end with the 800-nt "X2" repeats just adjacent to the telomerase simple repeats. The just-subterminal repeats are very similar to each other except that chromosome 1 right end is more divergent.

Base Sequence↗

Relationship between P-box amino acid sequence and DNA binding specificity of the thyroid hormone receptor. The effects of sequences flanking half-sites in thyroid hormone response elements.

The three P-box amino acids in the DNA recognition alpha-helix of steroid/thyroid hormone receptors participate in the discrimination of the central base pairs of the hexameric half-sites of receptor response elements in DNA. Using a series of variant receptors incorporating all 19 possible substitutions for each individual P-box amino acid of the human thyroid hormone receptor (hT3R beta), we demonstrated that the first P-box position must have a glutamate, and the second P-box position must have either an alanine or a glycine for high affinity binding to everted repeat elements with half-site sequences of AGGNCA. In the present study, the influence of half-site flanking sequence on the compatibility of P-box amino acids in hT3R beta with DNA binding was investigated. When a 5' sequence of CTG flanked AGGNCA half-sites in an everted repeat, several additional P-box variant receptors were able to bind to the DNA that were not able to bind when the half-sites were flanked with the 5' sequence CAG. Flanking sequence had the most dramatic effects on amino acid substitutions at the first P-box position, with smaller effects observed at the second P-box position and only subtle effects observed at the third P-box position. Expansion of the number of P-box sequences compatible with binding of hT3R beta to thyroid hormone response elements required the thymidine in the CTG flanking sequence, an everted repeat of the AGGNCA half-sites, and an intermolecular interaction in the C terminus of the receptor.

Amino Acid Sequence↗

Nucleotide sequences of coat protein genes for three isolates of barley yellow dwarf virus and their relationships to other luteovirus coat protein sequences.

Barley yellow dwarf virus (BYDV) can be separated into two groups based on, among other criteria, serological relationships that are presumably governed by the viral capsid structure. Nucleotide sequences for the coding regions of coat proteins of approximately 22 K were identified for the MAV-PS1, P-PAV (group 1) and NY-RPV (group 2) isolates of BYDV. The MAV-PS1 and P-PAV coat protein sequences shared 71% deduced amino acid similarity whereas that of the NY-RPV isolate shared no more than 51% similarity with either the MAV-PS1 or the P-PAV sequence. Other comparisons showed that these and other BYDV coat protein sequences examined to date share a high degree of identity with those identified from other luteoviruses. Among luteovirus coat protein sequences in general, several highly conserved domains were identified whereas other domains differentiate MAV-PS1 and PAV isolates from NY-RPV and other luteoviruses. Sequence similarities and differences among BYDV coat proteins (approx. 22K) are consistent with the serological relationships exhibited by these viruses. Amino acid sequence comparisons between BYDV isolates that share common aphid vectors indicate that it is unlikely that these coat proteins are involved in aphid specificity.

Amino Acid Sequence↗

The nucleotide sequence of potato virus A genomic RNA and its sequence similarities with other potyviruses.

The complete nucleotide sequence of potato virus A (PVA) was obtained from six independent cDNA clones. The RNA genome of PVA is 9565 nucleotides long and contains one open reading frame (ORF) of 9177 bases encoding a large polyprotein of 3059 amino acids with a calculated M(r) of 340K. Seven potential proteinase NIa, one HC-pro and one P1 proteinase recognition sites were found in PVA polyprotein by searching for cleavage site consensus sequences amongst the potyvirus group. The non-coding region preceding the ORF is 161 nucleotides long. The termination codon is followed by a 227-nucleotide sequence. Overall nucleotide sequence identity compared with several completely sequenced potyvirus genomes is between 53 and 58%, with overall amino acid sequence identity between 65 and 71%. When the putative amino acid sequences of individual proteins of PVA were compared with the corresponding proteins of other potyviruses, P1 and P3 appeared the least conserved (34 to 53%) whereas the other proteins were in most cases from 63 to 80% identical to each other.

Amino Acid Sequence↗

Nucleotide sequence and characterization of a repetitive DNA element from the genome of Bordetella pertussis with characteristics of an insertion sequence.

A repeating element of DNA has been isolated and sequenced from the genome of Bordetella pertussis. Restriction map analysis of this element shows single internal ClaI, SphI, BstEII and SalI sites. Over 40 DNA fragments are seen in ClaI digests of B. pertussis genomic DNA to which the repetitive DNA sequence hybridizes. Sequence analysis of the repeat reveals that it has properties consistent with bacterial insertion sequence (IS) elements. These properties include its length of 1053 bp, multiple copy number and presence of 28 bp of near-perfect inverted repeats at its termini. Unlike most IS elements, the presence of this element in the B. pertussis genome is not associated with a short duplication in the target DNA sequence. This repeating element is not found in the genomes of B. parapertussis or B. bronchiseptica. Analysis of a DNA fragment adjacent to one copy of the repetitive DNA sequence has identified a different repeating element which is found in nine copies in B. parapertussis and four copies in B. pertussis, suggesting that there may be other repeating DNA elements in the different Bordetella species. Computer analysis of the B. pertussis repetitive DNA element has revealed no significant nucleotide homology between it and any other bacterial transposable elements, suggesting that this repetitive sequence is specific for B. pertussis.

Base Sequence↗

Analysis of the lacZ sequences from two Streptococcus thermophilus strains: comparison with the Escherichia coli and Lactobacillus bulgaricus beta-galactosidase sequences.

The lacZ gene from Streptococcus thermophilus A054, a commercial yogurt strain, was cloned on a 7.2 kb PstI fragment in Escherichia coli and compared with the previously cloned lacZ gene from S. thermophilus ATCC 19258. Using the dideoxy chain termination method, the DNA sequences of both lacZ structural genes were determined and found to be 3071 bp in length. When the two sequences were more closely analysed, 21 nucleotide differences were detected, of which only nine resulted in amino acid changes in the proteins, the remainder occurring in wobble positions of the respective codons. Only three bases separated the termination codon for the lacS gene from the initiation codon for lacZ, suggesting that the lactose utilization genes are organized as an operon. The amino acid sequence of the beta-galactosidase, derived from the DNA sequence, corresponds to a protein with a molecular mass of 116860 Da. Comparison of the S. thermophilus amino acid sequences with those from Lactobacillus bulgaricus, E. coli and Klebsiella pneumoniae showed 48, 35 and 32.5% identity respectively. Although little sequence homology was observed at the DNA level, many regions conserved in the amino acid sequence were identified when the beta-galactosidase proteins from S. thermophilus, E. coli and L. bulgaricus were compared.

Amino Acid Sequence↗

Nucleotide sequence of cDNA and predicted amino acid sequence of rat liver uricase.

A cDNA clone for rat liver uricase (EC 1.7.3.3), which is localized in the core of peroxisomes, was isolated from a rat liver cDNA library in lambda gt11. The clone, referred to as lambda rURC-1, induced formation of a fusion protein (molecular mass 140 kDa) in the presence of isopropyl beta-D-thiogalactoside. Immunoglobulin G reactive with the fusion protein recognized only a protein with molecular mass of 33 kDa, corresponding to the molecular mass of uricase. The nucleotide sequence of the isolated cDNA was determined and the amino acid sequence was predicted. An open reading frame was identified and found to encode a polypeptide of 280 amino acids with a molecular mass of 32226 Da. The cDNA contained 14 base pairs of 5'-untranslated sequence and 192 base pairs of 3'-untranslated sequence. The sequences of four internal peptide fragments, determined by Edman degradation, were identical to parts of the sequence predicted from the cDNA. The complete amino acid sequence predicted for rat liver uricase was compared with that of soybean nodulin uricase and nine highly homologous regions of the two enzymes were found.

Amino Acid Sequence↗