Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Sequence complexity for biological sequence analysis.

A new statistical model for DNA considers a sequence to be a mixture of regions with little structure and regions that are approximate repeats of other subsequences, i.e. instances of repeats do not need to match each other exactly. Both forward- and reverse-complementary repeats are allowed. The model has a small number of parameters which are fitted to the data. In general there are many explanations for a given sequence and how to compute the total probability of the data given the model is shown. Computer algorithms are described for these tasks. The model can be used to compute the information content of a sequence, either in total or base by base. This amounts to looking at sequences from a data-compression point of view and it is argued that this is a good way to tackle intelligent sequence analysis in general.

Algorithms↗

Identification of multiple genital HPV types and sequence variants by consensus and nested type-specific PCR coupled with cycle sequencing.

Consensus and type-specific HPV primers were employed for PCR and cycle sequencing of genital HPVs in scrapings and colposcopically directed biopsies of the cervix from a cohort of 188 female sex workers. A total of 27 individuals tested positive for a broad spectrum of HPV types, including HPVs 6b, 16, 18, 31, 33, 34, 35, 45, 56 and 58, as well as a new HPV type, with seven individuals displaying dual infections. Good correlation between the results of individually paired samples was observed. A HPV 16 primer biotinylated at the 5' end was also used as a probe, which could successfully detect amplified products of HPV 16 but not other HPV types tested by an automated ELISA detection system. DNA sequence analysis revealed several HPV sequence variants that harbored mutations, especially in the E6 gene, many of which culminated in non-conservative amino acid substitutions in the transforming E6 oncoprotein. Such an approach of coupling PCR with cycle sequencing permits the determination of many known and even novel HPV types associated with varying degrees of risk to cervical carcinogenesis, and enables the identification of HPV sequence variants of putative biological and clinical significance, thus justifying its utility as an adjunct tool to complement cervical cytology and colposcopy. This study also emphasises the need for educational, interventional and behavioral modification to minimise HPV transmission, such as through consistent condom usage among sex workers.

Cervix Uteri↗

High similarity sequence comparison in clustering large sequence databases.

We present a fast algorithm for sequence clustering and searching which works with large sequence databases. It uses a strictly defined similarity measure. The algorithm is faster than conventional EST clustering approaches because its complexity is directly related to the number of subwords shared by the sequences. Furthermore, the algorithm also works with proteic sequences and large sequences like entire chromosomes. We present a theoretical study of our approach and provide experimental results.

Algorithms↗

Nucleotide sequence analysis of human beta-globin gene by the quantification method: mutations in 3'-splice junction sequence and beta-thalassemia.

The nucleotide sequence at the intron-exon junction in the human beta-globin gene was analyzed by the quantification method (categorical discriminant analysis) proposed previously. Using the sample score of a 16-nucleotide sequence at a 3'-splice junction, we studied to what extent such a sequence contains the 3'-splice signal. To examine the applicability of our method, we further studied several mutants of beta-thalassemia, where nucleotide changes exist at 3'-splice junction sequences of the first and second introns. Other mutants involve point mutations which generate new 3'-splice signals within the first intron. Experimental results on the abnormal splicing in those mutants could be explained in terms of the sample scores of 16-nucleotide sequences and their locations relative to the branch point.

Base Sequence↗

Amino acid sequence of Escherichia coli glutamine synthetase deduced from the DNA nucleotide sequence.

Glutamine synthetase is encoded by the glnA gene of Escherichia coli and catalyzes the formation of glutamine from ATP, glutamate, and ammonia. A 1922-base pair fragment from a cDNA containing the glnA structural gene for E. coli glutamine synthetase has been sequenced. An open reading frame of 1404 base pairs encodes a protein of 468 amino acid residues with a calculated molecular weight of 51,814. With few exceptions, the amino acid sequence deduced from the DNA sequence agreed very well with the amino acid sequences of several peptides reported previously. The secondary structure predicted for the E. coli enzyme has approximately 36% of the residues in alpha-helices which is in agreement with calculations of approximately 39% based on optical rotatory dispersion data. Comparison of the amino acid sequences of glutamine synthetase from E. coli (468 amino acids) and Anabaena (473 amino acids) (Turner, N. E., Robinson, S. T., and Haselkorn, R. (1983) Nature 306, 337-342) indicates that 260 amino acids are identical and 80 are of the same type (polar or nonpolar) when aligned for maximum homology. Several homologous regions of these two enzymes exist, including the sites of adenylylation and oxidative modification, but the regulation of each enzyme is different.

Amino Acid Sequence↗

Complete amino acid sequence of mouse pro-opiomelanocortin derived from the nucleotide sequence of pro-opiomelanocortin cDNA.

Polyadenylated RNA was isolated from a mouse pituitary tumor cell line (AtT-20/D16v) which synthesizes and secretes adrenocorticotropic hormone and beta-endorphin. The RNA was used to construct a cDNA library by a double linker technique and the library was screened for pro-opiomelanocortin (POMC) sequences. One recombinant plasmid, pMKSU16, contained a 923-base pair insert comprising the entire POMC-coding sequence as well as 98 bases of 5' noncoding and 5 bases of 3' noncoding sequence. The protein sequence predicted by the cDNA shows mouse POMC to consist of 235 amino acids with seven potential tryptic cleavage sites consisting of pairs of basic amino acid residues. Comparison with the previously published bovine POMC sequence suggests certain regions of POMC are highly conserved between the two species, particularly the regions corresponding to the alpha-, beta-, and gamma-melanocyte-stimulating hormones.

Adrenocorticotropic Hormone↗

Survey of the hemagglutinin (HA) cleavage site sequence of H5 and H7 avian influenza viruses: amino acid sequence at the HA cleavage site as a marker of pathogenicity potential.

The deduced amino acid sequence at the hemagglutinin (HA) cleavage site of 76 avian influenza (AI) viruses, subtypes H5 and H7, was determined by reverse transcription-polymerase chain reaction and cycle sequencing techniques to assess pathogenicity. Eighteen of the 76 viruses were isolated in 1993 and 1994 from various sources in the United States. In addition, 34 H5 (4 highly pathogenic [HP] and 30 non-highly pathogenic [non-HP]) and 24 H7 (3 HP and 21 non-HP) repository viruses, isolated between 1927 and 1992, were sequenced and the sequences compared to those in recent isolates. All repository HP H5 and H7 viruses studied had multiple basic amino acids adjacent to the HA cleavage site and most had basic amino acids in excess of the proposed minimum motif B-X-B-R (B = basic amino acids arginine or lysine, X = nonbasic amino acid, R = arginine) that has been associated with high pathogenicity. Of the non-HP viruses studied, 35 of 38 for H5 and 30 of 31 for H7 conformed to the motif B-X-X-R and B-X-R, respectively. Two non-HP H5 viruses had the motif X-X-X-R at the cleavage site and a third had the motif B-X-X-K (K = basic amino acid lysine). One non-HP H7 (A/Pekin robin/CA/30412-5/94) had four basic amino acids (K-R-R-R) adjacent to the cleavage site. Although the Pekin robin isolate did not produce disease in chickens under the conditions of the study it did have the amino acid sequence compatible with that in HP AI viruses and, therefore, is considered potentially HP. This is the first account of an H7 virus that is non-HP in chickens but meets the molecular criterion to be classified as HP.

Amino Acid Sequence↗

Finding common sequence and structure motifs in a set of RNA sequences.

We present a computational scheme to search for the most common motif, composed of a combination of sequence and structure constraints, among a collection of RNA sequences. The method uses a simplified version of the Sankoff algorithm for simultaneous folding and alignment of RNA sequences, but maintains tractability by constructing multi-sequence alignments from pairwise comparisons. The overall method has similarities to both CLUSTAL and CONSENSUS, but the core algorithm assures that the pairwise alignments are optimized for both sequence and structure conservation. Example solutions, and comparisons with other approaches, are provided. The solutions include finding consensus structures identical to published ones.

Algorithms↗

Two antipeptide monoclonal antibodies that recognize adhesive sequences in fibrinogen: identification of antigenic determinants and unrelated sequences using synthetic combinatorial libraries.

The fine specificity of two different monoclonal antibodies raised against synthetic peptides, each representing one of the two Arg-Gly-Asp (RGD) sequences in fibrinogen, was examined using synthetic combinatorial libraries (SCLs). The monoclonal antibodies (mAb), mAb LJ-134B/29 and mAb LJ-155B/16, recognize both the immunogenic peptide and native fibrinogen. The specificity of mAb LJ-134B29 was mapped using hexa- and decapeptide positional scanning SCLs (PS-SCLs) and competitive ELISA. The most active amino acids at each position of the two libraries were identified from a single screening. Individual hexa- and decapeptides were synthesized and assayed to determine their binding affinities. The 16 individual hexapeptides represented single and multiple substitutions of the antigenic determinant sequence, -GDSTFE-, eight of which had affinities less than 10nM. Four of the twelve individual decapeptides were found to have binding affinities of approximately 300nM, or nearly three-fold less than the peptide immunogen. A dual-defined hexapeptide library was screened against mAb LJ-155B/16, and individual peptides were obtained through an iterative selection and synthesis process. Surprisingly, one of the most active sequences was Ac-WWYESW-NH2 (IC50 = 40nM), which showed no similarity to the sequence of the immunizing peptide. Further mapping of the specificity of this antibody revealed that the antigenic determinant within the peptide immunogen was not completely linear. Recognition of this unrelated sequence by mAb LJ-155B/16 was confirmed in a direct binding assay using biotinylated peptide. The use of SCLs for the elucidation of high affinity peptides recognized by these two antibodies may provide additional information on the molecular mechanisms of fibrinogen binding to different integrin receptors.

Amino Acid Sequence↗

Nucleotide sequencing of S-RNA segment and sequence analysis of the nucleocapsid protein gene of the newly isolated Akabane virus PT-17 strain.

The nucleotide sequences of the S-RNA of Akabane viruses JaGAr-39, OBE-1, Iriki and the newly isolated PT-17 strains and the Aino virus were determined and compared. The results reveal that the S-RNAs of the four Akabane strains share 96.9% homology in nucleotide sequences. Only one amino acid difference out of the 233 amino acids of the nucleocapsid protein (N) and three amino acid differences in the 91 amino acids of the nonstructural protein (NSs) were found among the Akabane viruses. Amino acid sequences of N and NSs proteins of the Aino virus have approximately 80% identity as compared with the Akabane viruses. The results also demonstrate that the four Akabane viruses and the Aino virus can be clearly differentiated by RFLP (restriction fragments length polymorphism) analysis using RT-PCR generated nucleocapsid protein genes and digested with HaeIII and HindIII. The phylogenetic tree based on the UPGMA (Unweighted Pair Group Method with Arithmetic Mean) analysis of the sequences of nucleocapsid protein genes and the S-DNAs revealed that the newly isolated PT-17 strain is most closely related to Iriki strain, than the JaGAr-39 or OBE-1 strains.

Amino Acid Sequence↗

Functional analysis of the cysteine residues and the repetitive sequence of Saccharomyces cerevisiae Pir4/Cis3: the repetitive sequence is needed for binding to the cell wall beta-1,3-glucan.

Identification of PIR/CIS3 gene was carried out by amino-terminal sequencing of a protein band released by beta-mercaptoethanol (beta-ME) from S. cerevisiae mnn9 cell walls. The protein was released also by digestion with beta-1,3-glucanases (laminarinase or zymolyase) or by mild alkaline solutions. Deletion of the two carboxyterminal Cys residues (Cys(214)-12aa-Cys(227)-COOH), reduced but did not eliminate incorporation of Pir4 (protein with internal repeats) by disulphide bridges. Similarly, site-directed mutation of two other cysteine amino acids (Cys(130)Ser or Cys(197)Ser) failed to block incorporation of Pir4; the second mutation produced the appearance of Kex2-unprocessed Pir4. Therefore, it seems that deletion or mutation of individual cysteine molecules does not seem enough to inhibit incorporation of Pir4 by disulphide bridges. In fks1Delta and gsc2/fks2Delta cells, defective in beta-1,3-glucan synthesis, modification of the protein pattern found in the supernatant of the growth medium, as well as the material released by beta-ME or laminarinase, was evident. However, incorporation of Pir4 by both disulphide bridges and to the beta-1,3-glucan of the cell wall continued. Deletion of the repetitive sequence (QIGDGQVQA) resulted in the secretion and incorporation by disulphide bridges of Pir4 in reduced amounts together with substantial quantities of the Kex2-unprocessed Pir4 form. Pir4 failed to be incorporated in alkali-sensitive linkages involving beta-1,3-glucan when the first repetitive sequence was deleted. Therefore, this suggests that this sequence is needed in binding Pir4 to the beta-1,3-glucan.

Amino Acid Sequence↗

Sequence analysis of a 10 kb fragment of yeast chromosome XI identifies the SMY1 locus and reveals sequences related to a pre-mRNA splicing factor and vacuolar ATPase subunit C plus a number of unidentified open reading frames.

We report the DNA sequence analysis of a region on the left arm of chromosome XI of Saccharomyces cerevisiae extending over 10 kb. The region contains five open reading frames (ORFs) of greater than 100 amino acids which do not show significant overlap with other ORFs. YKL408 contains a sequence with strong similarity to the RNA helicase pre-mRNA splicing factors PRP2, PRP16 and PRP22 (Burgess et al., 1990; Company et al., 1991; Ruby et al., 1991). YKL409 corresponds to the gene SMY1, the sequence of which was previously reported by Lillie and Brown (1992). YKL410 is identical to ATPase subunit C (Beltran et al., 1992) except for an N-terminal extension. YKL406 and YKL407 show no significant identity with any sequences in the databases searched.

Adenosine Triphosphatases↗

The gamma-globin genes and their flanking sequences in primates: findings with nucleotide sequences of capuchin monkey and tarsier.

By sequencing extensive regions of the beta-globin gene cluster from capuchin monkey (New World monkey) and tarsier (prosimian) we confirmed that capuchin monkey and tarsier have two and one gamma-globin gene(s), respectively. These findings indicate that the ancestral anthropoid gamma-globin gene duplicated after anthropoids diverged from tarsier, but before they diverged into platyrrhines (New World monkeys) and catarrhines (Old World monkeys, apes, and human). The capuchin monkey gamma 1-globin gene promoter region accumulated many nucleotide substitutions, including a T to C substitution in the proximal CCAAT element. This adverse mutation, along with the previous finding that the gamma 1 locus in spider monkey is a pseudogene, suggests that in platyrrhines the gamma 2-globin gene may be the primary fetal beta-like globin gene. The aligned gamma gene sequences contain several conserved sequence elements (phylogenetic footprint) of 6 bp or longer in the 5' flanking region, but none in the 3' flanking region. Gene conversions frequently occurred in the 5' flanking and transcribed regions of the duplicated genes of anthropoids but rarely in the 3' flanking sequences. However, an absence of conversions within the capuchin and spider monkeys promoter regions suggests that in platyrrhines selection acted against conversions as they could decrease (or inactivate) gamma 2 expression.

Amino Acid Sequence↗

An alphoid DNA sequence conserved in all human and great ape chromosomes: evidence for ancient centromeric sequences at human chromosomal regions 2q21 and 9q13.

Using vector-CENP-B box polymerase chain reaction (PCR) we isolated and cloned from a human chromosome 21-specific plasmid library, a 1 kb DNA sequence, named p alpha H21. In in situ hybridization experiments, p alpha H21 hybridized, under high stringency conditions, to the centromeric region of all the human, chimpanzee, gorilla and orangutan chromosomes. On human chromosomes p alpha H21 also identified non-centromeric sequences at 2q21 (locus D2F33S1) and 9q13 (locus D9F33S2). The possible derivation of these sequences from ancestral centromeres is discussed. Sequence analysis confirmed the alphoid nature of the whole p alpha H21 insert.

Animals↗

Nucleotide sequence of the matrix, fusion and putative SH protein genes of mumps virus and their deduced amino acid sequences.

cDNA clones representing the M, F and a putative SH gene of the SBL strain of mumps virus have been prepared and their nucleotide sequence determined. The M gene of mumps virus appears to contain 1253 nucleotides and codes for a protein of 375 amino acid residues (Mr 41,589). The protein is hydrophobic and the deduced amino acid sequence shows homologies with those of other paramyxoviruses. The F gene of the SBL strain was compared to that of the RW strain [Waxham et al. (1987) Virology 159, 381-388]. The F gene is 1727 nucleotides long and codes for a protein of 538 amino acids (Mr 58,789). There are substantial variations between the F gene sequences of various mumps virus strains. The F gene is followed by a small (315 nt) transcription unit which contains an open reading frame of 57 amino acids encoding a very hydrophobic protein (Mr 6712). This may be similar to the SH gene of SV5, although there appears to be no sequence homology between the SV5 SH protein and the putative SH protein of mumps virus. A physical and transcription map of mumps virus indicates the gene order to be 3'-N-P-M-F-SH-HN, similar to that of other paramyxoviruses and SV5 in particular.

Amino Acid Sequence↗

Sequence of the phosphoenolpyruvate carboxykinase-encoding cDNA from the rumen anaerobic fungus Neocallimastix frontalis: comparison of the amino acid sequence with animals and yeast.

The nucleotide sequence of the cDNA of the phosphoenolpyruvate carboxykinase-encoding gene from the fungus Neocallimastix frontalis, was determined. The deduced amino acid sequence (608 residues) and the predicted protein structure were compared to their counterparts in animals and yeast. Catalytic regions (substrate-binding site and nucleotide-binding domains) are highly conserved among fungal and animal organisms. The yeast sequence showed no similarity to the fungal sequence.

Amino Acid Sequence↗

The human mdr1 (multidrug-resistance) gene harbours a long homopyrimidine.homopurine sequence next to a cluster of Alu repeated sequences in intron 14.

In order to identify specific DNA sequences useful as 'genetic landmarks' in the construction of a complete map of the human mdr1 (multidrug-resistance) gene, we investigated the introns in the central region. In intron 14, we identified a long stretch of a homopyrimidine.homopurine sequence most probably adopting an unconventional DNA conformation, followed by a cluster of three Alu repeated sequences in an inverted orientation. Here, we describe the structure, formation and nucleotide sequence of these DNA elements.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Potato lectin: a three-domain glycoprotein with novel hydroxyproline-containing sequences and sequence similarities to wheat-germ agglutinin.

Potato (Solanum tuberosum) tuber lectin is a chitin-binding, hydroxyproline-rich glycoprotein, which may be involved in the defence mechanism of the plant. We had previously obtained evidence that it consists of at least two very dissimilar domains. The aim was to use a combination of accurate determinations of molecular weight and protein sequencing to gain more accurate information on the domains. Accurate determinations of the molecular weight of the lectin by a MALDI mass spectrometer have shown that the subunit molecular weight is 65,500 (+/- 1100) and that of a totally deglycosylated sample is 31,250 (+/- 30). This means that the lectin is 52.3 (+/- 1)% carbohydrate with a considerable number of glycoforms being present. Partial sequences and other analyses are consistent with the existence of three distinct domains. These are: (1) an N-terminal region which is rich in proline but poor in hydroxyproline; (2) a glycosylated region with a glycosylated molecular weight of 45,300 (+/- 1100) and a deglycosylated molecular weight of 11,050 (+/- 50) which is extremely rich in glycosylated hydroxyproline residues with a similar sequence to extensins; and (3) a cystine-rich domain which has the sugar binding site shows partial conservation of a repeated motif common to many chitin-binding proteins of the hevin family including wheat-germ agglutinin. The closest similarity seems to be to the sequence of potato basic chitinase.

Amino Acid Sequence↗