Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Nucleotide sequence of the gene encoding the Newcastle disease virus fusion protein and comparisons of paramyxovirus fusion protein sequences.

The nucleotide sequence of cloned cDNA copies of the mRNA encoding the Newcastle disease virus fusion protein was determined. A single open reading frame in the sequence encodes a hydrophobic protein of 553 amino acids with a calculated molecular weight of 58 978. The previously determined protein sequence of the amino terminus of the F1 (Richardson, G.D. et al. (1980) Virology 105, 205-222) was located within the predicted protein sequence. The predicted protein sequence contains a hydrophobic stretch of 29 amino acids near the carboxy terminal end and likely represents the membrane spanning region of the protein. The F2 portion of the sequence contains one glycosylation site while F1 contains four which are potentially used. The predicted sequence contains 13 cysteine residues. Comparison of the NDV fusion protein sequence with three other paramyxovirus fusion protein sequences reveals little homology common to all four viruses except for the amino terminus of the F1 proteins. However, the positions of the cysteine residues within the sequence are conserved, particularly among the members of the paramyxovirus subgroup, suggesting the importance of disulfide bond formation in the conformation of paramyxovirus fusion proteins.

Amino Acid Sequence↗

The membrane-interactive tail of cytochrome b(5) can function as a stop-transfer sequence in concert with a signal sequence to give inversion of protein topology in the endoplasmic reticulum.

Sequence analyses of the C-terminal membrane intercalative region of the rat cytochrome b(5) indicated that this domain has, in addition to a signal sequence, a combined element of the classic stop-transfer sequence typically found in a variety of transmembrane proteins. Such bitopic protein arrangements arise by tandem but topogenically displaced activities of cleavable/noncleavable signal and stop-transfer sequences. A fusion precursor comprising an N-terminally linked prokaryotic signal sequence and the full-length of mammalian cytochrome b(5), including its C-terminal membrane insertion sequence, was engineered to investigate the outcome of this combination of signals on the targeting and topology of the cytochrome b(5) in the endoplasmic reticulum membrane. Precytochrome b(5) was cotranslationally translocated across the endoplasmic reticulum membrane. The signal-processed cytochrome b(5) was integrally anchored in the membrane with the globular domain facing the lumen. Thus, the topology of the signal sequence-directed cytochrome b(5) in the microsomal vesicle was reversed with respect to that of the native form. Posttranslational incubation of the precytochrome b(5) with microsomes resulted in a "loose" incorporation of the unprocessed form onto the surface of the vesicle. Our findings suggest that the membrane-insertion sequence of cytochrome b(5) has a functional stop-transfer sequence. We discuss the implications of these findings with respect to selective targeting of cytochrome b(5) to the endoplasmic reticulum membrane in the view that signal and stop-transfer sequences are often interchangeable or combined for topogenic functions.

Amino Acid Sequence↗

Sequence and analysis of bovine enteritic coronavirus (F15) genome. I. Sequence of the gene coding for the nucleocapsid protein; analysis of the predicted protein.

Sequences encoding the N protein of the bovine enteritic coronavirus-F15 strain (BECV-F15) have been cloned in PBR322 plasmid using cDNA produced by priming with oligo-dT on purified viral genomic RNA. Some 265 insert-containing clones were studied. Hybridization of these inserts with poly(A)+ RNA extracted from infected cells led to the conclusion that they were located at the 3'-end of the genome. After subcloning in M13 phage DNA, clones were sequenced by the Sanger technique. A 1,710-nucleotide sequence corresponding to the gene coding for the viral N-protein was established. It shows 2 overlapping open reading frames (ORF). The 3'-non-coding end of the gene has an 8-nucleotide sequence in common with the homologous genome areas of MHV, TGE and IBV viruses. This sequence may represent the polymerase RNA binding site. An upstream sequence surrounding the first AUG of the smaller ORF corresponds to a potentially functional initiation codon. The sequence of the primary translation product deduced from the DNA sequence predicts a polypeptide of 207 amino acids (22.9 Kd) with a high leucine (19.8%) content, possessing a hydrophobic N-terminal end. The larger ORF has a coding capacity of 448 amino acids (49.4 Kd), corresponding to the N-protein molecular weight. The deduced protein possesses 43 serine residues (9.6% of the total amino acid content) which may be phosphorylated and involved in N-protein/RNA binding. N-protein also has 5 regions with a high basic amino acid content. One of them is also serine-rich and has a strong homology site with MHV, TGE and IBV viruses. In the first part of the N-terminal, a 12-amino-acid sequence (PRWYFYYLGTGP) is highly conserved for BECV-F15, JHM, TGE and IBV viruses. BCV Mebus strain and BECV-F15 have only minor differences in their N-protein sequence.

Amino Acid Sequence↗

Manifold sequencing: efficient processing of large sets of sequencing reactions.

Automated instruments for DNA sequencing greatly simplify data collection in the Sanger sequencing procedure. By contrast, the so-called front-end problems of preparing sequencing templates, performing sequencing reactions, and loading these on the instruments remain major obstacles to extensive sequencing projects. We describe here the use of a manifold support to prepare and perform sequencing reactions on large sets of templates in parallel, as well as to load the reaction products on a sequencing instrument. In this manner, all reaction steps are performed without pipetting the samples. The strategy is applied to sequencing PCR-amplified clones of the human mitochondrial D-loop and for detection of heterozygous positions in the human major histocompatibility complex class II gene HLA-DQB, amplified from genomic DNA samples. This technique will promote sequencing in a clinical context and could form the basis of more efficient genomic sequencing strategies.

Base Sequence↗

DNA sequences tightly bound to proteins in mouse chromatin: identification of murine MER sequences.

The finding of stably (tightly) associated DNA-protein complexes in eukaryotic chromatin has provoked many hypotheses and speculations concerning their possible role. While the answer of this question is not envisaged yet, it is clear that elucidation of the nature of the individual components involved in such complexes is a necessary step in this direction. Here, the nature of several mouse DNA sequences in the vicinity of a putative stably attached protein is studied. Eight independently isolated clones containing such sequences were compared to known sequences in GenBank. Two clones were found to belong to different subfamilies of repetitive sequences, organized into a larger family--the L1md family. One clone harbors a sequence that is a member of the Alu-type family. Four of the cloned sequences are preset in low copy numbers, but the computer search found similar sequences in various genomic regions of different rodents. These facts, together with the finding that regions homologous to the above clones often flank other repetitive elements in the genome, suggest that the cloned sequences belong to new, not yet described families of repeats in the murine genome. It is possible that they correspond to the medium reiteration frequency sequences, MER-sequences, discovered recently in the human genome (Jurka, 1990; Kaplan and Duncan, 1990). Particularly intriguing is the homology found at the integration sites of polyoma virus in two transformed cell lines with two of these clones.

Animals↗

Fast comparison of a DNA sequence with a protein sequence database.

We describe a computer program, named DNA-Protein Search (DPS), for comparing a megabase DNA sequence with a protein sequence database. The DPS program addresses the problems of frameshifts and introns in the DNA sequence. The DPS program was used to compare each of the following sequences with the Swiss-Prot database: the 1.8-megabase sequence of the Haemophilus influenzae Rd genome, the 0.58-megabase sequence of the Mycoplasma genitalium genome, and the 0.56-megabase sequence of Saccharomyces cerevisiae chromosome VIII. The comparisons found new regions that are similar to protein sequences. The sensitivity of DPS was evaluated using as test data the known coding regions of the three DNA sequences. The results demonstrate that the DPS program is a useful tool for finding the coding regions of the DNA sequence. The DPS program uses an order of magnitude less computer memory and is several times faster than the BLASTX program.

Amino Acid Sequence↗

Membrane potential-driven protein import into mitochondria. The sorting sequence of cytochrome b(2) modulates the deltapsi-dependence of translocation of the matrix-targeting sequence.

The transport of preproteins into or across the mitochondrial inner membrane requires the membrane potential Deltapsi across this membrane. Two roles of Deltapsi in the import of cleavable preproteins have been described: an electrophoretic effect on the positively charged matrix-targeting sequences and the activation of the translocase subunit Tim23. We report the unexpected finding that deletion of a segment within the sorting sequence of cytochrome b(2), which is located behind the matrix-targeting sequence, strongly influenced the Deltapsi-dependence of import. The differential Deltapsi-dependence was independent of the submitochondrial destination of the preprotein and was not attributable to the requirement for mitochondrial Hsp70 or Tim23. With a series of preprotein constructs, the net charge of the sorting sequence was altered, but the Deltapsi-dependence of import was not affected. These results suggested that the sorting sequence contributed to the import driving mechanism in a manner distinct from the two known roles of Deltapsi. Indeed, a charge-neutral amino acid exchange in the hydrophobic segment of the sorting sequence generated a preprotein with an even better import, i.e. one with lower Deltapsi-dependence than the wild-type preprotein. The sorting sequence functioned early in the import pathway since it strongly influenced the efficiency of translocation of the matrix-targeting sequence across the inner membrane. These results suggest a model whereby an electrophoretic effect of Deltapsi on the matrix-targeting sequence is complemented by an import-stimulating activity of the sorting sequence.

Amino Acid Sequence↗

TBP flanking sequences: asymmetry of binding, long-range effects and consensus sequences.

We carried out in vitro selection experiments to systematically probe the effects of TATA-box flanking sequences on its interaction with the TATA-box binding protein (TBP). This study validates our previous hypothesis that the effect of the flanking sequences on TBP/TATA-box interactions is much more significant when the TATA box has a context-dependent DNA structure. Several interesting observations, with implications for protein-DNA interactions in general, came out of this study. (i) Selected sequences are selection-method specific and TATA-box dependent. (ii) The variability in binding stability as a function of the flanking sequences for (T-A)4 boxes is as large as the variability in binding stability as a function of the core TATA box itself. Thus, for (T-A)4 boxes the flanking sequences completely dominate and determine the binding interaction. (iii) Binding stabilities of all but one of the individual selected sequences of the (T-A)4 form is significantly higher than that of their mononucleotide-based consensus sequence. (iv) Even though the (T-A)4 sequence is symmetric the flanking sequence pattern is asymmetric. We propose that the plasticity of (T-A)n sequences increases the number of conformationally distinct TATA boxes without the need to extent the TBP contact region beyond the eight-base-pair long TATA box.

Base Pairing↗

Sequence heterogeneities among 16S ribosomal RNA sequences, and their effect on phylogenetic analyses at the species level.

We have analyzed what phylogenetic signal can be derived by small subunit rRNA comparison for bacteria of different but closely related genera (enterobacteria) and for different species or strains within a single genus (Escherichia or Salmonella), and finally how similar are the ribosomal operons within a single organism (Escherichia coli). These sequences have been analyzed by neighbor-joining, maximum likelihood, and parsimony. The robustness of each topology was assessed by bootstrap. Sequences were obtained for the seven rrn operons of E. coli strain PK3. These data demonstrated differences located in three highly variable domains. Their nature and localization suggest that since the divergence of E. coli and Salmonella typhimurium, most point mutations that occurred within each gene have been propagated among the gene family by conversions involving short domains, and that homogenization by conversions may not have affected the entire sequence of each gene. We show that the differences that exist between the different operons are ignored when sequences are obtained either after cloning of a single operon or directly from polymerase chain reaction (PCR) products. Direct sequencing of PCR products produces a mean sequence in which mutations present in the most variable domains become hidden. Cloning a single operon results in a sequence that differs from that of the other operons and of the mean sequence by several point mutations. For identification of unknown bacteria at the species level or below, a mean sequence or the sequence of a single nonidentified operon should therefore be avoided. Taking into account the seven operons and therefore mutations that accumulate in the most variable domains would perhaps increase tree resolution. However, if gene conversions that homogenize the rRNA multigene family are rare events, some nodes in phylogenetic trees will reflect these recombination events and these trees may therefore be gene trees rather than organismal trees.

Bacteria↗

The primary sequence of rhesus monkey rhadinovirus isolate 26-95: sequence similarities to Kaposi's sarcoma-associated herpesvirus and rhesus monkey rhadinovirus isolate 17577.

The primary sequence of the long unique region L-DNA (L for low GC) of rhesus monkey rhadinovirus (RRV) isolate 26-95 was determined. The L-DNA consists of 130,733 bp that contain 84 open reading frames (ORFs). The overall organization of the RRV26-95 genome was found to be very similar to that of human Kaposi sarcoma-associated herpesvirus (KSHV). BLAST search analysis revealed that in almost all cases RRV26-95 coding sequences have a greater degree of similarity to corresponding KSHV sequences than to other herpesviruses. All of the ORFs present in KSHV have at least one homologue in RRV26-95 except K3 and K5 (bovine herpesvirus-4 immediate-early protein homologues), K7 (nut-1), and K12 (Kaposin). RRV26-95 contains one MIP-1 and eight interferon regulatory factor (vIRF) homologues compared to three MIP-1 and four vIRF homologues in KSHV. All homologues are correspondingly located in KSHV and RRV with the exception of dihydrofolate reductase (DHFR). DHFR is correspondingly located near the left end of the genome in RRV26-95 and herpesvirus saimiri (HVS), but in KSHV the DHFR gene is displaced 16,069 nucleotides in a rightward direction in the genome. DHFR is also unusual in that the RRV26-95 DHFR more closely resembles HVS DHFR (74% similarity) than KSHV DHFR (55% similarity). Of the 84 ORFs in RRV26-95, 83 contain sequences similar to the recently determined sequences of the independent RRV isolate 17577. RRV26-95 and RRV17577 sequences differ in that ORF 67.5 sequences contained in RRV26-95 were not found in RRV17577. In addition, ORF 4 is significantly shorter in RRV26-95 than was reported for RRV17577 (395 versus 645 amino acids). Only four of the corresponding ORFs between RRV26-95 and RRV17577 exhibited less than 95% sequence identity: glycoproteins H and L, uracil DNA glucosidase, and a tegument protein (ORF 67). Both RRV26-95 and RRV17577 have unique ORFs between positions 21444 to 21752 and 110910 to 114899 in a rightward direction and from positions 116524 to 111082 in a leftward direction that are not found in KSHV. Our analysis indicates that RRV26-95 and RRV17577 are clearly independent isolates of the same virus species and that both are closely related in structural organization and overall sequence to KSHV. The availability of detailed sequence information, the ability to grow RRV lytically in cell culture, and the ability to infect monkeys experimentally with RRV will facilitate the construction of mutant strains of virus for evaluating the contribution of individual genes to biological properties.

Amino Acid Sequence↗

Automated cycle sequencing of PCR templates: relationships between fragment size, concentration and strand renaturation rates on sequencing efficiency.

With the Applied Biosystems 373A automated DNA sequencer, we have systematically investigated the amounts of double-stranded PCR fragments of varying size (200, 564, and 1126 bp) required to give sequence of defined lengths, up to the maximum possible. Sequencing was performed on purified double-stranded PCR products using the dye-terminator chemistry and a thermal cycling procedure. The minimal template concentrations allowing determination of short sequences (< or = 160 bases) were essentially identical for the fragments studied. Maximal possible sequence determination from the 200 bp fragment was achieved over a wide concentration range, despite the fact that within this range a significant fraction of the template renatured by the mid-point of the sequencing reaction time-course. We conclude that the cyclic sequencing process overcomes competitive strand reannealing of double-stranded PCR products. The sequencing concentration-response curves for the 564 bp and 1126 bp fragments were similar to each other, although the minimal template concentrations required to read > 300 bases were slightly increased for the 564 bp fragment. Excess template is undesirable for optimal sequence length determination, but this is unlikely to be solely due to strand reannealing as single-stranded M13 templates in super-optimal concentrations also showed marked reduction in sequencing efficiency.

Automation↗

Primary structure of the archaebacterial Methanococcus vannielii ribosomal protein L12. Amino acid sequence determination, oligonucleotide hybridization, and sequencing of the gene.

The primary structure of ribosomal protein L12 from Methanococcus vannielii has been determined by direct amino acid sequence analysis with automated liquid phase Edman degradation of the entire protein and manual 4-N,N'-dimethylaminoazobenzene-4'-isothiocyanate/phenylisothiocyanate sequencing of fragments obtained by enzymatic digestion and by partial acid hydrolysis. The knowledge of the amino acid sequences of these various fragments allowed the synthesis of two oligonucleotide probes complementary to the 5'- and the 3'-end of the gene, and they were used for hybridization with digested M. vannielii chromosomal DNA. Both oligonucleotide probes gave similar and clear hybridization signals. The plasmid pMvaX1 containing the entire gene of protein L12 was obtained. The nucleotide sequence complemented the partial amino acid sequence, and it is in full agreement with the protein sequence and the amino acid analysis. Comparison of secondary structural elements and hydrophobicity plots of the M. vannielii protein L12 with the known L12 sequences derived from other archaebacterial and eukaryotic sources show strong homologies among these sequences. They contain an exceptional highly conserved hydrophilic sequence area in the C-terminal part of the proteins. In comparison with eubacterial L12 proteins, the conservation is reduced to single amino acid residues. However, the eubacterial L12 proteins have hydrophilic regions similar to those of L12 from M. vannielii. These regions are predicted to be located at the surface of the proteins, as has been proven to be the case in crystallized Escherichia coli L12 protein. It is possible that the strongly conserved hydrophilic sequence regions form part of the factor-binding domain.

Amino Acid Sequence↗

New DNA sequence rules for high affinity binding to histone octamer and sequence-directed nucleosome positioning.

DNA sequences that position nucleosomes are of increasing interest because of their relationship to gene regulation in vivo and because of their utility in studies of nucleosome structure and function in vitro. However, at present our understanding of the rules for DNA sequence-directed nucleosome positioning is fragmentary, and existing positioning sequences have many limitations. We carried out a SELEX experiment starting with a large pool of chemically synthetic random. DNA molecules to identify those individuals having the highest affinity for histone octamer. A set of highest-affinity molecules were selected, cloned, and sequenced, their affinities (free energies) for histone octamer in nucleosome reconstitution measured, and their ability to position nucleosomes in vitro assessed by native gel electrophoresis. The selected sequences have higher affinity than previously known natural or non-natural sequences, and have a correspondingly strong nucleosome positioning ability. A variety of analyses including Fourier transform, real-space correlation, and direct counting computations were carried out to assess non-random features in the selected sequences. The results reveal sequence rules that were already identified in earlier studies of natural nucleosomal DNA, together with a large set of new rules having even stronger statistical significance. Possible physical origins of the selected molecules' high affinities are discussed. The sequences isolated in this study should prove valuable for studies of chromatin structure and function in vitro and, potentially, for studies in vivo.

Base Composition↗

The complete sequence of the rabbit erythroid cell-specific 15-lipoxygenase mRNA: comparison of the predicted amino acid sequence of the erythrocyte lipoxygenase with other lipoxygenases.

We report the complete sequence of the rabbit reticulocyte (RBC) 15-lipoxygenase (LOX) mRNA as deduced from (i) sequencing cDNA recombinants isolated by screening cDNA libraries or polymerase-chain-reactions, and (ii) the sequence originating from the transcription start point obtained by primer extension-sequencing reactions. Like the human leukocyte 5-LOX mRNA, the RBC 15-LOX mRNA contains a very short 5'-untranslated region with a long 3'-untranslated region. But, unlike the human leukocyte 5-LOX mRNA, the RBC 15-LOX mRNA contains an intriguing repeated sequence (ten copies with the consensus sequence C4PuC3TCTTC4AAG) just after the translational stop codon, which may be involved in its regulation during reticulocyte maturation. Comparison of the RBC 15-LOX mRNA sequence with those of the previously published human 5-LOX mRNA and the soybean 3-LOX gene shows only a few short regions of sequence similarity. However, the predicted amino acid sequences of the encoded LOX enzymes show certain conserved regions that are presumably involved in their catalytic activity, in particular a cluster of five conserved histidines that we predict chelate the iron moiety involved in the active site.

Amino Acid Sequence↗

Comparison of sequence masking algorithms and the detection of biased protein sequence regions.

MOTIVATION: Separation of protein sequence regions according to their local information complexity and subsequent masking of low complexity regions has greatly enhanced the reliability of function prediction by sequence similarity. Comparisons with alternative methods that focus on compositional sequence bias rather than information complexity measures have shown that removal of compositional bias yields at least as sensitive and much more specific results. Besides the application of sequence masking algorithms to sequence similarity searches, the study of the masked regions themselves is of great interest. Traditionally, however, these have been neglected despite evidence of their functional relevance. RESULTS: Here we demonstrate that compositional bias seems to be a more effective measure for the detection of biologically meaningful signals. Typical results on proteins are compared to results for sequences that have been randomized in various ways, conserving composition and local correlations for individual proteins or the entire set. It is remarkable that low-complexity regions have the same form of distribution in proteins as in randomized sequences, and that the signal from randomized sequences with conserved local correlations and amino acid composition almost matches the signal from proteins. This is not the case for sequence bias, which hence seems to be a genuinely biological phenomenon in contrast to patches of low complexity.

Aeropyrum↗

Novel repeated DNA sequences in safflower (Carthamus tinctorius L.) (Asteraceae): cloning, sequencing, and physical mapping by fluorescence in situ hybridization.

Two novel repetitive DNA sequences, pCtKpnI-1 and pCtKpnI-2, were isolated from Carthamus tinctorius (2n = 2x = 24) and cloned. Both represent tandemly repeated sequences. The pCtKpnI-1 and pCtKpnI-2 clones constitute repeat units of 343-345 bp and 367 bp, respectively, with 63% sequence heterogeneity between the two. Fluorescence in situ hybridization (FISH) was employed on metaphase chromosomes of C. tinctorius using, simultaneously, pCtKpnI-1 and pCtKpnI-2 repeated sequences. The pCtKpnI-1 sequence was found to be exclusively localized at subtelomeric regions on most of the chromosomes. On the other hand, sequence of the pCtKpnI-2 clone was distributed on two nucleolar and one nonnucleolar chromosome pairs. The satellite, and the intervening chromosome segment between the primary and secondary constrictions, in the two nucleolar chromosome pairs were wholly constituted by pCtKpnI-2 repeated sequence. The pCtKpnI-2 repeated sequence, showing partial homology to intergenic spacer (IGS) of 18S-25S ribosomal RNA genes of an Asteraceae taxon (Centaurea stoebe), and the 18S-25S rRNA gene clusters were located at independent, but juxtaposed sites in the nucleolar chromosomes. Variability in the number, size, and location of the two repeated sequences provided identification of most of the chromosomes in the otherwise not too distinctive homologues within the complement. This article reports the start of a molecular cytogenetics program targeting the genome of safflower, a major world oil crop about whose genetics very little is known.

Base Sequence↗

Genomic organization, sequence interrelationship, and physical localization using in situ hybridization of two tandemly repeated DNA sequences in the genus Olea.

Two tandemly repeated DNA sequences, the 81-bp family and pOS218, have been isolated from a Sau3AI Olea europaea ssp. sativa partial genomic library. Sequencing of the 81-bp element showed the monomer to be between 78 and 84 bases long and to contain 51-58% adenine and thymidine residues. Comparison between the monomers revealed heterogeneity of the sequence primary structure. The clone pOS218 is 218 bases long, and sequence comparison between the two elements revealed that an internal region of the pOS218 repeated DNA sequence had 79% homology to the 81 bp repeat sequence. A breakage-reunion mechanism, involving the CAAAA sequence, could be responsible for the derivation of pOS218 from the 81 bp family element. By using double target in situ hybridization, co-localization of the two sequences on Olea chromosomes was observed. The sequences were present at DAPI stained heterochromatic regions, as major or minor sites having a subtelomeric or interstitial location. Methylation studies using two sets of isoschizomers, Sau3AI-MboI and MspI-HpaII, demonstrated that most cytosine residues in the GATC sites and the internal cytosine in the CCGG sites of both elements were methylated in O. europaea ssp. sativa. No major difference in methylation was apparent between DNA extracted from young leaves or from callus of O. europaea ssp. sativa. Both elements are also present in Olea chrysophylla, Olea oleaster, and Olea africana, but are absent from other Oleaceae genera, including Phillyrea, Forsythia, Ligustrum, Parasyringa, and Jasminum.

Base Sequence↗

A fractal method to distinguish coding and non-coding sequences in a complete genome based on a number sequence representation.

A fractal method to distinguish coding and non-coding sequences in a complete genome is proposed, based on different statistical behaviors between these two kinds of sequences. We first propose a number sequence representation of DNA sequences. Multifractal analysis is then performed on the measure representation of the obtained number sequence. The three exponents C(-1), C1 and C2 are selected from the result of multifractal analysis. Each DNA may be represented by a point in the three-dimensional space generated by these three-component vectors. It is shown that points corresponding to coding and non-coding sequences in the complete genome of many prokaryotes are roughly distributed in different regions. Fisher's discriminant algorithm can be used to separate these two regions in the spanned space. If the point (C(-1),C1,C2) for a DNA sequence is situated in the region corresponding to coding sequences, the sequence is discriminated as a coding sequence; otherwise, the sequence is classified as a non-coding one. For all 51 prokaryotes we considered , the average discriminant accuracies pc,pnc,qc and qnc reach 72.28%, 84.65%, 72.53% and 84.18%, respectively.

Animals↗