Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition↗

Sequenase sequence profiles used for HLA-DPB1 sequencing-based typing.

Sequencing-based HLA typing (SBT) is a PCR based high resolution HLA typing method in which polymorphic regions of the gene are sequenced and directly used for typing. Currently, for class II SBT, alleles are identified by comparison of the exon 2 sequence with their corresponding allele sequence library. Routine SBT requires reliable identification of heterozygosity, and automated assignment of the alleles. In sequencing strategies different enzymes can be used for primer extension. The most characteristic difference between sequences obtained by two protocols using Sequenaseregistered, or Taq-cycle sequencing, respectively, is a difference in incorporation of nucleotides in the primer extension leading to different sequence profiles. In Taq-cycling sequencing variable nucleotide incorporation results in irregular, but reproducible peak patterns, whereas Sequenase incorporates nucleotides in nearly equal amounts, resulting in more even peak patterns. In a previously published multi-center study we evaluated HLA-DPB1 SBT using Taq-cycle sequencing, and showed that typing can reliably be performed, considering the specific sequence profiles. In this study the applicability of Sequenase for HLA-DPB1 SBT was tested. A panel of samples were typed by SBT at five test sites which participate in the Sequencing Based Typing component of the 12th International Histocompatibility Workshop. The panel represents the existing polymorphism at all known polymorphic positions of exon 2, both in homozygous and heterozygous combinations. The assignment of homozygosity and heterozygosity was validated by Multi-Sequence Analysis, performing cluster analysis of chromatographic data of all sequences at each position. Sequence characteristics were examined and considered for appropriate assignment. Data reveals that Sequenase sequencing can also reliably be used for HLA-DPB1 typing.

Bacteriophage T7↗

Long range structural communication between sequences in supercoiled DNA. Sequence dependence of contextual influence on cruciform extrusion mechanism.

Sequence context may profoundly alter the character of structural transitions in supercoiled DNA (Sullivan, K. M., and Lilley, D. M. J. (1986) Cell 47, 817-827). The A + T-rich sequences of ColE1, which flank the inverted repeat, are responsible for cruciform extrusion following a mechanistic pathway which proceeds via a relatively large denatured region. This C-type mechanism results in kinetic properties which are very different from those of the S-type pathway, the normal mechanism of cruciform extrusion in the absence of the ColE1 flanking sequences. We have analyzed the sequence requirements for the induction of the C-type pathway. The 100-base pair left side sequence of ColE1 (colL) was subjected to systematic deletion using Bal31 exonucleolysis, showing that removal of 30 base pairs from its right end abolished extrusion by the C-type process. A cloned oligonucleotide of the same 30-base pair sequence was sufficient to confer C-type cruciform extrusion on an adjacent inverted repeat. An A + T-rich sequence from Drosophila was found to act like the ColE1 sequences. We have studied the effects of introducing sequences between the A + T-rich colL, and the inverted repeat on which it acts. A range of such fragments was found, from those which augment the effect of colL to those which block it completely. In general, it appears that the ability of a sequence to block the effect of colL depends on both the length and G + C content of the fragment. The sequences which are responsible for the extrusion by the C-type pathway are termed C-type inducing sequences, while sequences which are interposed between the inducing sequence and the inverted repeat, and which may either augment or attenuate the effect, but which cannot function as inducing sequences in isolation, are termed transmitting sequences. The results of these studies are most readily consistent with long range destabilization of DNA structure via telestability effects.

Animals↗

Comparison of HASTE and segmented-HASTE sequences with a T2-weighted fast spin-echo sequence in the screening evaluation of the brain.

OBJECTIVE: The purpose of this study was to evaluate the neuroradiologic application of half-Fourier acquisition single-shot turbo spin-echo (HASTE) and segmented-HASTE (s-HASTE) sequences in comparison with a T2-weighted fast spin-echo sequence. MATERIALS AND METHODS: First, HASTE, s-HASTE, and fast spin-echo sequences were evaluated for blurring artifacts with a stationary phantom and for motion artifacts with a moving phantom, which repeated constant or intermittent to-and-fro motions at variable intervals. Second, 30 consecutive patients with various intracranial diseases were prospectively examined with the three sequences. Lesions were classified into four groups according to size and signal intensity on fast spin-echo MR images as follows: large hyperintense, small hyperintense, small markedly hyperintense, and hypointense lesions. Signal intensities of the lesion, putamen, and gray matter were compared with the signal intensity of white matter, and contrast-to-noise ratios were calculated. Overall image quality, conspicuity of lesions, delineation of the junction between gray matter and white matter, conspicuity of the putamen, and certain types of artifacts were evaluated qualitatively. RESULTS: In the phantom study, the HASTE sequence was least affected by motion artifacts and the fast spin-echo sequence was most affected although the images of the HASTE sequence were most degraded by blurring artifacts. In the clinical study, we found no significant differences among the three sequences for contrast-to-noise ratios or conspicuity of large hyperintense and small markedly hyperintense lesions. However, the contrast-to-noise ratios of hypointense lesions and gray matter, and the conspicuity of hypointense lesions were significantly poorer for the HASTE sequence than for the fast spin-echo sequence. The contrast-to-noise ratios of small hyperintense lesions and the putamen, conspicuity of small hyperintense lesions and putamen, and delineation of the junction between gray matter and white matter were significantly poorer for HASTE and s-HASTE sequences than for the fast spin-echo sequence. Ghost artifacts, which were observed during the s-HASTE sequence, were sometimes superimposed on the image. CONCLUSION: The HASTE and s-HASTE sequences afford substantial time reduction and also decrease motion artifacts and thus have potential advantages for neuroradiologic application, especially in uncooperative or unsedated children. The s-HASTE sequence may be preferable to the HASTE sequence because of fewer blurring artifacts and higher T2 contrast. However, small hyperintense and hypointense lesions may be overlooked when HASTE and s-HASTE sequences are used.

Artifacts↗

[Comparison of magnetic resonance Spin-echo sequences and fat-suppressed sequences in bone diseases].

Thirty-two patients affected with skeletal conditions were examined with MRI using Short TI Inversion Recovery sequence and Spectral Presaturation with Inversion Recovery (SPIR) sequence as well as Spin-Echo (SE) T1-weighted sequence and Fast Spin-Echo (FSE) T2-weighted sequence to compare their value in the assessment of skeletal lesions. SPIR sequence was performed after intravenous injection of Gd-DTPA. The lesions included primary bone tumors (10 cases: 1 osteosarcoma, 1 periosteal sarcoma, 1 Ewing's sarcoma, 1 chondrosarcoma, 2 non-ossifying fibromas, 1 chondroma, 1 chondromyxoid fibroma, 1 desmoplastic fibroma and 1 bone cyst), metastases (7 cases: 3 prostate, 3 breast, 1 lung-squamous cell carcinoma), infections (12 cases: 9 osteomyelitis, 3 spondylodiscitis), sacroiliitis (1 case) and posttraumatic bone bruise (2 cases of bone marrow edema). The four sequences were compared by using both qualitative and quantitative evaluation. Qualitative evaluation showed that STIR sequence was better than SPIR sequence (performed with Gd-DTPA) for lesion conspicuity (p < .016) and for signal intensity uniformity (p < .03). Compared with SE T1 and FSE T2 sequences, fat-suppressed sequences were superior for conspicuity, margins, and extension of the lesions (range of p < .001-.017). Only SPIR with Gd-DTPA sequence, compared with SE T1 sequence for lesion conspicuity was not statistically significantly different. Quantitative evaluation showed statistically significant higher values of percent contrast (%C) and contrast-to-noise ratio (C/N) for STIR sequence compared with SPIR sequence (%C p < .004; C/N p < .040). This study suggests that STIR sequence and SE T1-weighted sequence provide high sensitivity in lesion detection and good anatomical definition. The use of a fat-suppressed sequence with Gd-DTPA can be useful for lesion characterization.

Adolescent↗

Classification of mouse VK groups based on the partial amino acid sequence to the first invariant tryptophan: impact of 14 new sequences from IgG myeloma proteins.

Fourteen new VK sequences derived from BALB/c IgG myeloma proteins were determined to the first invariant tryptophan (Trp 35). These partial sequences were compared with 65 other published VK sequences using a computer program. The 79 sequences were organized according to the length of the sequence from the amino terminus to the first invariant tryptophan (Trp 35), into seven groups (33, 34, 35, 36, 39, 40 and 41aa). A distance matrix of all 79 sequences was then computed, i.e. the number of amino acid substitutions necessary to convert one sequence to another was determined. From these data a dendrogram was constructed. Most of the VK sequences fell into clusters or closely related groups. The definition of a sequence group is arbitrary but facilitates the classification of VK proteins. We used 12 substitutions as the basis for defining a sequence group based on the known number of substitutions that are found in the VK21 proteins. By this criterion there were 18 groups in the Trp 35 dendrogram. Twelve of the 14 new sequences fell into one of these sequence groups; two formed new sequence groups. Collective amino acid sequencing is still encountering new VK structures indicating more sequences will be required to attain an accurate estimate of the total number of VK groups. Updated dendrograms can be quickly generated to include newly generated sequences.

Amino Acid Sequence↗

The Chinese hamster Alu-equivalent sequence: a conserved highly repetitious, interspersed deoxyribonucleic acid sequence in mammals has a structure suggestive of a transposable element.

A consensus sequence has been determined for a major interspersed deoxyribonucleic acid repeat in the genome of Chinese hamster ovary cells (CHO cells). This sequence is extensively homologous to (i) the human Alu sequence (P. L. Deininger et al., J. Mol. Biol., in press), (ii) the mouse B1 interspersed repetitious sequence (Krayev et al., Nucleic Acids Res. 8:1201-1215, 1980) (iii) an interspersed repetitious sequence from African green monkey deoxyribonucleic acid (Dhruva et al., Proc. Natl. Acad. Sci. U.S.A. 77:4514-4518, 1980) and (iv) the CHO and mouse 4.5S ribonucleic acid (this report; F. Harada and N. Kato, Nucleic Acids Res. 8:1273-1285, 1980). Because the CHO consensus sequence shows significant homology to the human Alu sequence it is termed the CHO Alu-equivalent sequence. A conserved structure surrounding CHO Alu-equivalent family members can be recognized. It is similar to that surrounding the human Alu and the mouse B1 sequences, and is represented as follows: direct repeat-CHO-Alu-A-rich sequence-direct repeat. A composite interspersed repetitious sequence has been identified. Its structure is represented as follows: direct repeat-residue 47 to 107 of CHO-Alu-non-Alu repetitious sequence-A-rich sequence-direct repeat. Because the Alu flanking sequences resemble those that flank known transposable elements, we think it likely that the Alu sequence dispersed throughout the mammalian genome by transposition.

Animals↗

Identification of peptides within a known protein sequence using COMSEQ analysis of data containing multiple sequences.

Modern methods of automated protein sequence analysis can provide high-quality data with which unambiguous amino-acid sequences can be determined, but analyses are more difficult when the sample is not pure. COMSEQ and auxillary programs were written to facilitate reconciliation of multiple amino-acid sequences potentially contained in noisy data with the known amino-acid sequence of the parent protein. The COMSEQ program prints a matrix in which the first vertical column represents the known amino-acid sequence of a selected protein. Each row of the matrix contains the sequencer yield corresponding to the amino acid in the first column, with each column corresponding to the sequencing reaction cycle. A diagonal which contains net increases of amino acids for each amino acid in the known sequence identifies a peptide potentially contained within the data. The number of matches for each diagonal over the entire known sequence are tabulated and presented as an aid to locating comparisons of greatest interest. The RNDSEQ program conducts multiple analyses using randomized versions of the known amino-acid sequence and tabulates the cumulative frequencies of potential sequence matches irrespective of the true known sequence. TRANSEQ is a utility program that translates edited sequence data from common databases into files that can be used by COMSEQ and RNDSEQ. The programs have been used successfully to identify two co-sequenced peptides from bovine serum albumin, an albumin peptide sequence in the presence of hemoglobin, and to identify two sequences of rat alpha-2u-globulin that differ in their amino termini.

Amino Acid Sequence↗

Signals for the selection of a splice site in pre-mRNA. Computer analysis of splice junction sequences and like sequences.

To evaluate the importance of the surrounding nucleotide sequence in the selection of a splice site for mRNA, we have carried out computer studies of eukaryotic protein genes whose entire nucleotide sequences were available. A splice site-like sequence that has a significant homology to the consensus splice junction sequences is frequently found within an intron and exon. It is found that the higher the homology of a candidate donor site sequence to the nine-nucleotide consensus sequence, the higher is its probability of being a donor site. For most of the donors, the stability of presumed base-pairing with U1-RNA is higher than that of donor-like sequences, if any, in the adjacent exon and intron. However, homology of a candidate acceptor sequence to the 15-nucleotide consensus is a poor criterion of an acceptor site. The presence of a sequence that could serve as a branch-point 18 to 37 nucleotides before an acceptor does not seem to be critical in distinguishing it from an acceptor-like sequence. For genes of human, rat, mouse and chicken, respectively, nucleotide frequencies around splice junctions of many genes have been calculated. They seem to be different at some positions around a donor site from species to species. The acceptors for these vertebrates have longer pyrimidine-rich regions than the previous consensus sequence. The newly derived nucleotide frequencies were used as the standard to calculate the weighted homology score of a candidate splice site sequence in a gene of the four species. This weighted homology score of the 40 to 60-nucleotide intron-exon sequence is a much better criterion of an acceptor. These results suggest that the most important signal in the selection of a splice resides in the surrounding nucleotide sequence. It is also suggested that the surrounding nucleotide sequence alone is not generally sufficient for the selection.

Animals↗

Nucleotide sequences of Caenorhabditis elegans core histone genes. Genes for different histone classes share common flanking sequence elements.

We have determined the nucleotide sequence of core histone genes and flanking regions from two of approximately 11 different genomic histone clusters of the nematode Caenorhabditis elegans. Four histone genes from one cluster (H3, H4, H2B, H2A) and two histone genes from another (H4 and H2A) were analyzed. The predicted amino acid sequences of the two H4 and H2A proteins from the two clusters are identical, whereas the nucleotide sequences of the genes have diverged 9% (H2A) and 12% (H4). Flanking sequences, which are mostly not similar, were compared to identify putative regulatory elements. A conserved sequence of 34 base-pairs is present 19 to 42 nucleotides 3' of the termination codon of all the genes. Within the conserved sequence is a 16-base dyad sequence homologous to the one typically found at the 3' end of histone genes from higher eukaryotes. The C. elegans core histone genes are organized as divergently transcribed pairs of H3-H4 and H2A-H2B and contain 5' conserved sequence elements in the shared spacer regions. One of the sequence elements, 5' CTCCNCCTNCCCACCNCANA 3', is located immediately upstream from the canonical TATA homology of each gene. Another sequence element, 5' CTGCGGGGACACATNT 3', is present in the spacer of each heterotypic pair. These two 5' conserved sequences are not present in the promoter region of histone genes from other organisms, where 5' conserved sequences are usually different for each histone class. They are also not found in non-histone genes of C. elegans. These putative regulatory sequences of C. elegans core histone genes are similar to the regulatory elements of both higher and lower eukaryotes. The coding regions of the genes and the 3' regulatory sequences are similar to those of higher eukaryotes, whereas the presence of common 5' sequence elements upstream from genes of different histone classes is similar to histone promoter elements in yeast.

Animals↗

Identification of novel transcribed sequences on human chromosome 22 by expressed sequence tag mapping.

To identify sequences on the human genome that are actually transcribed, we mapped expressed sequence tags (ESTs) of long cDNAs ranging from 4 kb to 7 kb along a 33.4-Mb sequence of human chromosome 22, the first human chromosome entirely sequenced. By the EST mapping of 30,683 long cDNAs in silico, 603 cDNA sequences were found to locate on chromosome 22 and classified into 169 clusters. Comparison of the genomic loci of these cDNA sequences with 679 genes already annotated on chromosome 22q revealed that 46 clusters represented newly identified transcribed sequences. To further characterize these sequences, we sequenced 12 cDNAs in their entirety out of 46 clusters. Of these 12 cDNAs, 6 were predicted to include a protein-coding region while the remaining 6 were unlikely to encode proteins. Interestingly, 3 out of the 12 cDNAs had the nucleotide sequences of the opposite strands of the genes previously annotated, which suggested that these genomic regions were transcribed bi-directionally. In addition to these newly identified 12 cDNAs, another 12 cDNAs were entirely sequenced since these cDNAs were likely to contain new information about the predicted protein-coding sequences previously annotated. In the cases of KIAA1670 and KIAA1672, these single cDNA sequences covered two separately annotated transcribed regions. For example, the sequence of a clone for KIAA1670 indicated that the CHKL and CPT1B genes were co-transcribed as a contiguous transcript without making both the protein-coding regions fused. In conclusion, the mapping of ESTs derived from long cDNAs followed by sequencing of the entire cDNAs provided indispensable information for the precise annotation of genes on the genome together with ESTs derived from short cDNAs.

Brain Chemistry↗

Effect of sequence length on the execution of familiar keying sequences: lasting segmentation and preparation?

The author assessed the mechanisms underlying skilled production of keying sequences in the discrete sequence-production task by examining the effect of sequence length on mean element execution rate (i.e., the rate effect). To that end, participants (N = 9) practiced fixed movement sequences consisting of 2, 4, and 6 key presses for a total of 588 trials per sequence. In the subsequent test phase, the sequences were executed with and without a verbal short-term memory task in both simple and choice reaction time (RT) paradigms. The rate effect was obtained in the discrete sequence-production task-including the typical quadratic increase in sequence execution time (SET, which excludes RT) with sequence length. The rate effect resulted primarily from 6-key sequences that included 1 or 2 relatively slow elements at individually different serial positions. Slowing of the depression of the 2nd response key (R2) in the 2-key sequence reduced the rate effect in the memory task condition, and faster execution of the 1st few elements in each sequence amplified the rate effect in simple RT. Last, the time to respond to random cues increased with position, suggesting that the mechanisms that underlie the rate effect in new sequences and in familiar sequences are different. The data were in line with the notion that coding of longer keying sequences involves motor chunks for the individual sequence segments and information on how those motor chunks are to be concatenated.

Adult↗

Coronavirus transcription mediated by sequences flanking the transcription consensus sequence.

In our studies of murine coronavirus transcription, we continue to use defective interfering (DI) RNAs of mouse hepatitis virus (MHV) in which we insert a transcription consensus sequence in order to mimic subgenomic RNA synthesis from the nondefective genome. Using our subgenomic DI system, we have studied the effects of sequences flanking the MHV transcription consensus sequence on subgenomic RNA transcription. We obtained the following results. (i) Insertion of a 12-nucleotide-long sequence including the UCUAAAC transcription consensus sequence at different locations of the DI RNA resulted in different efficiencies of subgenomic DI RNA synthesis. (ii) Differences in the amount of subgenomic DI RNA were defined by the sequences that flanked the 12-nucleotide-long sequence and were not affected by the location of the 12-nucleotide-long sequence on the DI RNA. (iii) Naturally occurring flanking sequences of intergenic sequences at gene 6-7, but not at genes 1-2 and 2-3, contained a transcription suppressive element(s). (iv) Each of three naturally occurring flanking sequences of an MHV genomic cryptic transcription consensus sequence from MHV gene 1 also contained a transcription suppressive element(s). These data showed that sequences flanking the transcription consensus sequence affected MHV transcription.

Animals↗

Average values of a dissimilarity measure not requiring sequence alignment are twice the averages of conventional mismatch counts requiring sequence alignment for a computer-generated model system.

Three measures of sequence dissimilarity have been compared on a computer-generated model system in which substitutions in random sequences were made at randomly selected sites and the replacement character was chosen at random from the set of characters different from the original occupant of the site. The three measures were the conventional mismatch count between aligned sequences (AMC = m) and two measures not requiring prior sequence alignment. The latter two measures were the squared Euclidean distance between vectors of counts of t-tuples (t = 1-6) of characters in the two sequences (multiplet distribution distances or MDD = d) and counts of characters not covered by word structures of statistically significant length common to the two sequences (common long words or CLW = SIB, SIS, or SAB). Average MDD distances were found to be two times average mismatch counts in the simulated sequences for all values of t from 1 to 6 and all degrees of substitution from one per sequence to so many as to produce, effectively, random sequences. This simple relation held independently of sequence length and of sequence composition. The relation was confirmed by exact results on small model systems and by formal asymptotic results in the limit of so few substitutions that no double hits occur and in the limit of two random sequences. The coefficient of variation for MDD distances was greater than that for mismatch counts for singlets but both measures approached the same low value for sextets. Needleman-Wunsch alignment produced incorrect mismatch counts at higher degrees of substitution. The model satisfied the conditions for the derivation of the Jukes-Cantor asymptotic adjustment, but its application produced increasingly bad results with increasing degrees of substitution in accord with earlier results on model and natural sequences. This fact was a consequence of the increase with increasing degrees of substitution of the sensitivity of the adjustment to error in the observations. Average CLW distances for a variety of common word structures were more or less parallel to MDD distances for appropriately long t-tuples. These results on model systems supported the validity of the two dissimilarity measures not requiring sequence alignment that was found in earlier work on natural sequences (Blaisdell 1989).

Base Sequence↗

An evaluation of the significance of amino acid sequence homologies in human histocompatibility antigens (HLA-A and HLA-B) with immunoglobulins and other proteins, using relatively short sequences.

A computer search was carried out for homologies between HLA-A and HLA-B antigen sequences and the sequences of constant and variable regions of immunoglobulins and of all other sequenced proteins. Searches were made both with relatively short peptide sequences from the HLA antigens and with those longer peptide sequences which were available in 1978. Significant homology of HLA antigen sequences to immunoglobulin constant region sequences was found in two cases: (1) a short decapeptide sequence which includes the fourth cysteine residue of HLA-B7 and (2) an 89-amino-acid residue (Ac-2) C-terminal fragment of the papainsolubilized HLA-B7 molecule. The difficulty of establishing statistically significant sequence homology with relatively short peptide sequences is emphasized by computer-based comparisons of the decapeptide sequence with randomly generated peptide sequences. It is concluded that statistically significant homology with short sequences can be assured only when extraordinarily high degrees of homology are present and additional constraints are included in the matches, for example, matches at relatively rare amino acid residues such as Cys, His and Trp. The homology of the 89-amino-acid residue sequence to constant region sequences of immunoglobulins is as great as or greater than that of beta 2-microglobulin. These findings and the unique domain structure involving a disulphide loop of comparable size strongly favour a common evolutionary origin for this region of HLA-A and -B, beta 2-microglobulin and immunoglobulin constant regions.

Amino Acid Sequence↗