Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Allele-specific PCR amplification due to sequence identity between a PCR primer and an amplicon: is direct sequencing so reliable?

PCR-direct sequencing (DS) is thought to be a very reliable method of determining DNA sequence and genotyping. Under certain conditions, however, DS can generate inaccurate results. Here we report a case of erroneous DS, in which a single nucleotide polymorphism (SNP) in the human PAX9 gene was mistyped due to allele-dependent PCR amplification. Examination of the amplified region showed that the 5' eight bases of one of the PCR primers were identical to the eight bases of the reverse strand downstream of the SNP, and the ninth base matched one of the alleles. Altering the primer so that it matched the other allele reversed the allele-specific inhibition. Reducing the base-pairing abolished the inhibition. Thus, the SNP was responsible for the difference in annealing efficacy of the primer and was therefore critical for the allele dependency. The allele-specific inhibition presented here can occur with any PCR primer sequence that encompasses a site that is polymorphic in the gene sequence. This phenomenon needs to be considered as a possibility when interpreting results from all PCR-based experiments. Sequence similarity between PCR primers and internal amplified regions should be considered for all methods for mutation detection and genotyping using PCR.

Alleles↗

Complete amino acid sequence of the large subunit of the low-Ca2+-requiring form of human Ca2+-activated neutral protease (muCANP) deduced from its cDNA sequence.

The complete amino acid sequence of the large subunit (catalytic subunit) of human low-Ca2+-requiring-calcium-activated neutral protease (muCANP) was deduced from its cDNA base sequence. It is composed of 714 amino acid residues and its sequence is highly homologous to the chicken CANP sequence determined previously. Human muCANP, like chicken CANP, has a clear 4-domain structure, and their fundamental structures are essentially the same, although their Ca2+ sensitivities are significantly different. The role of each domain in the Ca2+ sensitivity and protease activity of CANP is discussed on the basis of sequence comparison.

Amino Acid Sequence↗

Isolation of sequences from a random-sequence expression library that mimic viral epitopes.

We describe the use of random peptide sequences for the mapping of antigenic determinants. An oligonucleotide with a completely degenerate sequence of 17 or 23 nucleotides was inserted into a bacterial expression vector. This resulted in an expression library producing random hexa- or octapeptides attached to a beta-galactosidase hybrid protein. Mimotopes, or antigenic sequences that mimic an epitope, were selected by immunoscreening of colonies with monoclonal antibodies, which were specific for antigenic sites on the spike protein of the coronavirus transmissible gastroenteritis virus. We report one mimotope for antigenic site II, eight for site III and one for site IV. The site III and site IV mimotopes were closely similar to the corresponding linear epitopes, localized previously in the amino acid sequence of the S protein. An alignment of the site II mimotope and the sequence of the S protein around Trp97, which is substituted in escape mutants, suggests that this mimotope mimics a conformational epitope located around residues 97-103. Applications of mimotopes to epitope mapping, serodiagnosis and vaccine development are discussed.

Amino Acid Sequence↗

Bovine beta-crystallin complementary DNA clones. Alternating proline/alanine sequence of beta B1 subunit originates from a repetitive DNA sequence.

A library of recombinant plasmids carrying complementary DNA sequences synthesized from bovine lens messenger RNAs was constructed. Clones coding for five different beta-crystallin subunits: beta B1, beta B3, beta Bp, beta s, beta A3 (and beta A1), were identified by means of hybridization selection, followed by one- and two-dimensional gel electrophoresis of the translational products. Under rather stringent conditions each of these clones hybridizes with its corresponding mRNA and does not show significant cross-hybridization with mRNAs coding for other beta-crystallins, except in the case of the homologous beta A3 and beta A1-crystallins. The beta A3 and beta A1 subunits seem to be encoded by one mRNA using two different AUG codons as start position for translation. We have also determined the nucleotide sequence of a beta B1-crystallin cDNA (pBL beta B1) which enabled us to deduce the complete amino acid sequence of the protein. The beta B1-crystallin, a characteristic component of the high molecular weight crystallin aggregate (beta H), is internally homologous both at DNA and protein level as has been reported for gamma- and other beta-crystallins. This is in agreement with the idea that these proteins had a common ancestral precursor gene that internally duplicated. The G + C content of the coding sequence of beta B1 is very high: 67% overall and even 84.2% for the first 170 nucleotides, due to a remarkable non-random codon usage. A proline/alanine repetition in the N-terminal domain of the protein is encoded by a repetitive "simple" DNA sequence.

Alanine↗

Sequence organization of repetitive sequences enriched in small polydisperse circular DNAs from HeLa cells.

A total of 36 clones were randomly selected from a recombinant DNA library of small polydisperse circular DNA (spcDNA) molecules from HeLa cells and were shown to contain repetitive sequences of different reiteration frequencies that ranged from several hundred to several hundred thousand per genome. Sequencing of representative clones revealed tandem repeats of alphoid (alpha) satellite DNA, clustered repeats of the Alu family, KpnI family sequences, tandem repeats of an alpha satellite DNA specific to the X chromosome (alpha X), and A + T-rich segments carrying short stretches of poly(A) or poly(T). DNA rearrangement was frequently found in the repetitive sequences enriched in these spcDNA clones. Short regions of homology that were patchy and inverted were often found, especially at the novel joint where spcDNA sequences are circularized. The presence of these inverted repeats suggests that HeLa spcDNAs are formed by a mechanism that involves looping out of the spcDNA region and joining of the flanking DNA by illegitimate recombination.

Base Sequence↗

Sequence specificity of 125I-labelled Hoechst 33258 damage in six closely related DNA sequences.

The sequence selectivity of [125I]Hoechst 33258 in six 340 base-pair DNA sequences has been investigated. [125I]Hoechst 33258, which is a bis-benzimidazole and binds to the minor groove of B-DNA, preferentially binds to A + T-rich regions of DNA. Six out of nine strong binding sites contained four or more consecutive A.T base-pairs, while the other three strong binding sites were AAGGATT, TATAGAAA (the peak of damage was in the run of 3 A residues) and AAA. One of the six weak binding sites had five consecutive A.T base-pairs, two of the weak binding sites had three, and three did not have any. In addition to genomic 340 base-pair alpha RI-DNA (which is a tandem repeat in human cells), five 340 base-pair alpha RI-DNA clones were generated that differed from the genomic "consensus" sequence by a number of random base alterations. The effect of these base changes on the sequence specificity of [125I]Hoechst 33258 damage indicated that of the base changes that interrupted 14 binding sites, six decreased and eight did not change the extent of damage, while two sites changed position. Of the base alterations that augmented 17 binding sites, five increased, two decreased and ten did not alter the degree of cleavage, while ten sites changed position. It was concluded from the data that, while runs of consecutive A.T base-pairs was the most important parameter that determines [125I]Hoechst 33258 binding, other factors including position in the DNA sequence, nearest neighbour and long-range interactions were also important.

Autoradiography↗

Nucleotide sequence of the Lassa virus (Josiah strain) S genome RNA and amino acid sequence comparison of the N and GPC proteins to other arenaviruses.

The complete nucleotide sequence of the S genome RNA of the Josiah strain of Lassa virus was determined from cloned cDNA. The S RNA is 3402 nucleotides long with a calculated molecular weight of 1.09 x 10(6) Da. The nucleotide base composition is 26.84% adenine, 21.40% guanine, 22.75% cytosine, and 29.01% uridine. The 5' and 3' terminal nucleotide sequences are conserved and complimentary for 19 nucleotides, the nucleoprotein and glycoprotein genes are arranged in ambisense coding strategy, and the intergenic region contains an inverted complimentary sequence, as do all other arenavirus S RNAs characterized to date. Amino acid sequence comparisons between the nucleoproteins and glycoproteins of the Josiah and Nigerian (N sequences only) strains of Lassa virus, the WE and ARM strains of lymphocytic choriomeningitis virus (LCMV), Tacaribe, and Pichinde viruses are presented. These findings reveal that the G2 envelope glycoprotein is more conserved among different arenaviruses than the internal nucleoprotein.

Amino Acid Sequence↗

The nucleotide sequence of the gene encoding the F protein of canine distemper virus: a comparison of the deduced amino acid sequence with other paramyxoviruses.

The nucleotide sequence of the gene encoding the fusion protein of canine distemper virus was determined from cDNA clones derived from virus genome RNA and poly(A)+ RNA extracted from infected cells. The mRNA encoding the F protein is about 2300 nucleotides in length including the 3' poly(A) tail. There is a large open reading frame from nucleotides 86 to 2071 which begins at the first AUG codon in the F mRNA. This reading frame encodes a protein of 662 amino acid residues with a calculated mol. wt. of 73001. The first major hydrophobic domain in the amino acid sequence of the deduced protein (residues 104 to 130) may represent all or part of a signal sequence for cleavage of the N terminal part of the F2 protein. There are four potential N glycosylation sites in the F protein located within the F2 part of the molecule or the putative signal sequence, and one in the F1 portion. A second hydrophobic region corresponds to the proteolytic cleavage site which generates the F2 and F1 subunits. This stretches from residue 225 to 262 and the N terminal part of the F1 protein shows sequence conservation with the other paramyxoviruses. A third major hydrophobic domain near the C terminus of the F protein probably represents the membrane anchor for the F protein (residues 602 to 630). The F1 proteins of six paramyxoviruses are compared and shown to have substantial conservation of those residues important in the maintenance of tertiary structure of this protein.

Amino Acid Sequence↗

Sequence analysis of HLA-DR4B1 subtypes: additional first domain variability is detected by oligonucleotide hybridization and nucleotide sequencing.

The stimulating capacity of HLA-DR4 variants in mixed leukocyte culture correlates with variation in the polymorphic regions of the first domains of their DR beta 1 chains. Variability between amino acids 67 and 86 appears largely to determine HLA-DR4,Dw type. We have used a combination of a DR4B1-specific primer set in the polymerase chain reaction and specific oligonucleotide probes to examine DR4,Dw-associated nucleotide polymorphisms. Phenotype and gene frequencies are reported among 44 normal DR4 Caucasoids. Oligonucleotide probes were selected which enabled definition of Dw4-, w14-, w10-, w13-, and w15-associated sequences. Unexpectedly, several subjects were positive for Dw15 sequences which are usually characteristic of Oriental populations. Dw15 assignment was confirmed by nucleotide sequencing of DR4B1 polymerase chain reaction products. A pair of oligonucleotides informative for the glycine or valine dimorphism at position 86 was used to identify two novel DR4B1 alleles, designated 13.2 and 14.2. Nucleotide sequencing shows that these represent recombinants between third hypervariable regions associated with Dw13 and Dw14 and a codon for glycine at position 86 which is usually found associated with Dw4 and Dw15 sequences.

Alleles↗

Nucleotide sequence of the chicken cardiac alpha actin gene: absence of strong homologies in the promoter and 3'-untranslated regions with the skeletal alpha actin sequence.

The entire nucleotide sequence of the chicken cardiac alpha-actin (CC alpha A) gene has been determined. This is the first complete sequence of a cardiac actin gene that includes the promoter region, cap site, all the introns, and the polyadenylation site. The gene contains six introns, five of which interrupt the coding region at amino acids (aa) 41, 150, 204, 267, and 327. The first intron is in the 5'-noncoding region and is 438 bp in length. The CC alpha A gene encodes an mRNA of approx. 1400 bp with 5'- and 3'-untranslated region of 59 and 184 nucleotides (nt), respectively. Like the chicken skeletal alpha-actin gene, the CC alpha A gene has the codon for the aa cysteine between the initiator ATG and the codon for the N-terminal aspartic acid residue of the mature protein. There are no strong homologies (less than 13 consecutive nt) in the promoter or 3'-untranslated regions between the CC alpha A and chicken skeletal alpha-actin genes even though both are expressed in skeletal muscle during development. However, the 3'-untranslated region of the CC alpha A gene demonstrates significant sequence homology (76% over a 200-nt region) with the same region in the partial sequence of the human cardiac gene. The conservation of these sequence homologies between identical isoforms rather than the different alpha actin genes suggests these conserved regions may have a role in regulation rather than tissue-specific expression, as previously proposed.

Actins↗

Sequence of the immunoregulatory early region 3 and flanking sequences of adenovirus type 35.

Adenovirus type 35 (Ad35) is an important pathogen in immunosuppressed individuals such as AIDS patients and bone marrow transplant recipients. Ad35, a member of Ad subgroup B, differs with respect to pathogenic properties from the more fully characterized subgroup C Ad, such as Ad2 and Ad5. One region of human Ad which varies between subgroups and which may influence Ad pathogenesis is early region 3 (E3), a region which appears to modulate the immune response to Ad infection. In order to begin to characterize the differences between the Ad35 E3 and the E3 of other Ad, the complete DNA sequence of the Ad35 E3 promoter and coding sequence along with two flanking structural proteins, pVIII and fiber, has been determined. Ad35 contains open reading frames which are unique to the subgroup B Ad in addition to the four characterized immunoregulatory proteins encoded by the subgroup C Ad. Further evaluation of the sequence of one of these proteins, 18.5K, which is the class-I major histocompatibility complex (MHC) binding protein of 18.5 kDa, demonstrates that the amino acid sequence of this Ad2 gp19K homologue fits a proposed model of gp19K-MHC interaction. Analysis of promoter sequences demonstrates that an NF-kappa B site found in the subgroup C E3 promoter is absent from the Ad35 E3 promoter. In addition, the fiber genes of Ad35 and other subgroup B Ad have been shown to diverge in an unexpected way, yielding three clusters of fiber homology.

Adenovirus E3 Proteins↗

Identification and sequence analysis of IS1297, an ISS1-like insertion sequence in a Leuconostoc strain.

The insertion sequence (IS) ISS1 from Lactococcus lactis was amplified from lactococcal genomic DNA using a primer to the 18-bp inverted repeat sequence. The amplified product hybridized to a single EcoRI fragment in a total genomic DNA digest of Leuconostoc mesenteroides ssp. dextranicum NZDRI 2218. The DNA sequence of this ISS1-like element (IS1297) and the Le. mesenteroides sequences flanking the IS were determined and compared with other iso-ISS1 elements. No direct repeats were found immediately flanking IS1297; however, direct repeats were present approximately 60 bp on either side of the insertion site. IS1297 contained a major open reading frame (ORF) of 681 bp, encoding a putative 226-amino-acid protein with 96.5% homology to the presumed transposase of ISS1. An overlapping ORF of 174 bp in the same orientation was also present. A putative ORF in the opposite orientation to the transposase ORF, which has been shown in some iso-ISS1 elements, was not present in IS1297. IS1297 was shown to hybridize with other dairy Leuconostoc strains. This is the first sequence of an ISS1-like element from a genus other than Lactococcus; however, IS1297 has close similarity to the lactococcal iso-ISS1 elements, especially the iso-ISS1 element from the lactose plasmid, pTD1.

Base Sequence↗

Cloning, sequence analysis and confirmation of derived gene sequences for three epitope-mapped monoclonal antibodies against human phagocyte flavocytochrome b.

The integral membrane protein flavocytochrome b (Cyt b) is the catalytic core of the NADPH oxidase complex, a multicomponent enzyme system that initiates a cascade of reactive oxygen species that play a critical role in innate immunity and vascular physiology. Epitope-mapped, monoclonal antibodies (mAb) that recognize the large (gp91phox) and small (p22phox) subunits of Cyt b provide valuable reagents that have been used to examine structural and mechanistic aspects of oxidase function. In the present study, the heavy and light chain variable region genes of the Cyt b-specific mAbs 44.1, NS5, and NL7 have been amplified by RT-PCR, cloned and subject to DNA sequence analysis. Since the 5' degenerate primer sets used for mAb gene amplification were observed to introduce extensive heterogeneity into the heavy and light chain FR1 regions, N-terminal protein sequence analysis was also conducted to obtain the correct amino acid sequence of this region. In order to confirm the identity of the cloned genes, intact mAbs were resolved by two-dimensional electrophoresis and subject to in-gel tryptic digestion for analysis by both MALDI and nanospray LC-MS/MS. Databases searches using the derived mAb sequences predicted residues comprising CDR loops, identified candidate germline genes, and showed the respective germline genes to accurately predict the N-terminal amino acid residues for each variable region. The above studies report the amino acid sequence of Cyt b-specific mAb variable region genes with high confidence and provide essential information for future efforts at Cyt b structure analysis by resonance energy transfer and X-ray crystallography.

Amino Acid Sequence↗

Detecting localized repeats in genomic sequences: a new strategy and its application to Bacillus subtilis and Arabidopsis thaliana sequences.

A new method for the search of local repeats in long DNA sequences, such as complete genomes, is presented. It detects a large variety of repeats varying in length from one to several hundred bases, which may contain many mutations. By mutations we mean substitutions, insertions or deletions of one or more bases. The method is based on counting occurrences of short words (3-12 bases) in sequence fragments called windows. A score is computed for each window, based on calculating exact word occurrence probabilities for all the words of a given length in the window. The probabilities are defined using a Bernoulli model (independent letters) for the sequence, using the actual letter frequencies from each window. A plot of the probabilities along the sequence for high-scoring windows facilitates the identification of the repeated patterns. We applied the method to the 1.87 Mb sequence of chromosome 4 of Arabidopsis thaliana and to the complete genome of Bacillus subtilis (4.2 Mb). The repeats that we found were classified according to their size, number of occurrences, distance between occurrences, and location with respect to genes. The method proves particularly useful in detecting long, inexact repeats that are local, but not necessarily tandem. The method is implemented as a C program called EXCEP, which is available on request from the authors.

Arabidopsis↗

Sequence diversity in the 5'-UTR region of GB virus C/hepatitis G virus assessed using sequencing, heteroduplex mobility analysis and single-strand conformation polymorphism.

GB virus C/hepatitis G virus (GBV-C/HGV) is a positive-sense RNA virus belonging to the Flaviviridae family identified recently. Reverse transcription polymerase chain reaction (RT-PCR) was used to detect GBV-C/HGV RNA using nested primers designed to amplify 245 bp of the 5'-untranslated region (UTR). GBV-C/HGV RNA was detected in 20.7% of 101 HCV-RNA positive and 6.8% of 44 HCV-RNA negative specimens. Sequencing of the PCR products demonstrated they had between 84.3 and 100% nucleotide identity. Most of the diversity corresponded to two variable regions identified within the 5'-UTR. Phylogenetic analysis indicated that GBV-C/HGV subtypes present in Australia belonged to group 2 and were closest in evolutionary terms to isolates from the USA and Europe. All isolates were analysed using single-strand conformation polymorphism (SSCP) and heteroduplex mobility analysis (HMA) on 8% non-denaturing polyacrylamide gels. SSCP of the isolates identified a number of distinct conformation polymorphisms that corresponded with sequence-determined genetic diversity. HMA was developed to assess the amount of genetic diversity between isolates without the need for sequencing. The average difference between the predicted divergence of two isolates calculated from the mobility of the heteroduplex and the actual value (based on nucleotide sequence) was 2.3% in this sample of isolates, where the mean sequence divergence was 8.52%.

5' Untranslated Regions↗

Fluorescently labeled model DNA sequences for exonucleolytic sequencing.

We describe here the enzyme-catalyzed, low-density labeling of DNAs with fluorescent dyes. Firstly, for "natural" template DNAs, dNTPs were partially substituted in the labeling reactions by the respective fluorophore-bearing analogs. The DNAs were labeled by PCR using Taq DNA polymerase. The covalent incorporation of dye-dNTPs decreased in the following order: rhodamine-green-5-dUTP (Molecular Probes, the Netherlands), tetramethylrhodamine-4-dUTP (FluoroRed, Amersham Pharmacia Biotech), Cy5-dCTP (Amersham Pharmacia Biotech). Exonucleolytic degradation by the 3'-->5' exonuclease activity of T7 DNA polymerase (wild type) in the presence of excess reduced thioredoxin proceeded to complete breakdown of the labeled DNAs. The catalytic cleavage constants determined by fluorescence correlation spectroscopy were between 0.5 and 1.5 s(-1) at 16 degrees C, normalized for the covalently incorporated dye-nucleotides. Secondly, rhodamine-green-X-dUTP (Roche Diagnostics), tetramethylrhodamine-6-dUTP (Roche Diagnostics), and Cy5-dCTP were covalently incorporated into the antisense strand of "synthetic" 218-b DNA template constructs (master sequences) at well defined positions, starting from the primer binding site, by total substitution for the naturally occurring dNTPs. The 218-b DNA constructs were labeled by PCR with a thermostable 3'-->5' exonuclease deficient mutant of the Tgo DNA polymerase which we have selected. The advantage of the special, synthetic DNA constructs as compared to natural DNAs lies in the possibility of obtaining tailor-made nucleic acids, optimized for testing the performance of exonucleolytic sequencing. The number of incorporated fluorescent nucleotides determined by complete exonucleolytic degradation and fluorescence correlation spectroscopy were six out of six possible incorporations for rhodamine-green-X-dUTP and tetramethylrhodamine-6-dUTP, respectively. Their covalent and base-specific incorporations were confirmed by the novel analysis methodology of re-sequencing (i.e. mobility-shift gel electrophoresis, reversion-PCR and re-sequencing) first developed in the paper Földes-Papp et al. (2001) and in this paper. This methodology was then used by other groups within the whole sequencing project.

Base Sequence↗

Combining multiple structure and sequence alignments to improve sequence detection and alignment: application to the SH2 domains of Janus kinases.

In this paper, an approach is described that combines multiple structure alignments and multiple sequence alignments to generate sequence profiles for protein families. First, multiple sequence alignments are generated from sequences that are closely related to each sequence of known three-dimensional structure. These alignments then are merged through a multiple structure alignment of family members of known structure. The merged alignment is used to generate a Hidden Markov Model for the family in question. The Hidden Markov Model can be used to search for new family members or to improve alignments for distantly related family members that already have been identified. Application of a profile generated for SH2 domains indicates that the Janus family of nonreceptor protein tyrosine kinases contains SH2 domains. This conclusion is strongly supported by the results of secondary structure-prediction programs, threading calculations, and the analysis of comparative models generated for these domains. One of the Janus kinases, human TYK2, has an SH2 domain that contains a histidine instead of the conserved arginine at the key phosphotyrosine-binding position, betaB5. Calculations of the pK(a) values of the betaB5 arginines in a number of SH2 domains and of the betaB5 histidine in a homology model of TYK2 suggest that this histidine is likely to be neutral around pH 7, thus indicating that it may have lost the ability to bind phosphotyrosine. If this indeed is the case, TYK2 may contain a domain with an SH2 fold that has a modified binding specificity.

Amino Acid Sequence↗

Molecular cloning and deduced amino acid sequence of nonspecific lipid transfer protein (sterol carrier protein 2) of rat liver: a higher molecular mass (60 kDa) protein contains the primary sequence of nonspecific lipid transfer protein as its C-terminal part.

Two types of cDNA for nonspecific lipid transfer protein (nsLTP), identical to sterol carrier protein 2, of rat liver were cloned; one was 787 base pairs (bp) long containing a 429-bp open reading frame of 143 amino acids, with a mass of 15,303 Da (15-kDa protein). The cDNA from the other type was 1966 bp long, including a 1641-bp open reading frame of 547 amino acids, giving a mass of 59,002 Da (60-kDa protein). The deduced primary sequence for the 15-kDa protein was exactly the same as the published sequence of purified nsLTP, except for an extra N-terminal sequence of 20 amino acids, consistent with the finding that nsLTP is synthesized as a larger precursor and processed to a mature form. The sequence for the 60-kDa protein contained, at the 3' end, the full sequence of the 15-kDa protein, a larger precursor to nsLTP. The 15- and 60-kDa proteins, synthesized in vitro from the respective cDNAs, were both immunoprecipitated by rabbit anti-rat liver nsLTP antibody and comigrated in SDS/PAGE with the proteins made in vitro from total liver RNA. These results shed new light on the dispute among several groups of investigators about the crossreactivity of anti-nsLTP antibody with a higher molecular mass, 60-kDa protein. In Northern blot analysis, two major RNA bands, 0.85 and 2.2 kilobases (kb) long, were detected together with two minor bands of 1.6 and 2.9 kb. The 0.85- and 2.2-kb RNAs most likely encode the 15-and 60-kDa proteins, respectively.

Amino Acid Sequence↗