Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “multiple clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Different expression strategy: multiple intronic gene clusters of box H/ACA snoRNA in Drosophila melanogaster.

The high degree of rRNA pseudouridylation in Drosophila melanogaster provides a good model for studying the genomic organization, structural and functional diversity of box H/ACA small nucleolar RNAs (snoRNAs). Accounting for both conserved sequence motifs and secondary structures, we have developed a computer-assisted method for box H/ACA snoRNA searching. Ten snoRNA clusters containing 42 box H/ACA snoRNAs were identified from D.melanogaster. Strikingly, they are located in the introns of eight protein-coding genes. In contrast to the mode of one snoRNA per intron so far observed in all animals, our results demonstrate for the first time a novel polycistronic organization that implies a different expression strategy for a box H/ACA snoRNA gene when compared to box C/D snoRNAs in D.melanogaster. Mutiple isoforms of the box H/ACA snoRNAs, from which most clusters are made up, were observed in D.melanogaster. The degree of sequence similarity between the isoforms varies from 99% to 70%, implying duplication events in different periods and a trend of enlarging the intronic snoRNA clusters. The variation in the functional elements of the isoforms could lead to partial alternation of snoRNA's function in loss or gain of rRNA complementary sequences and probably contributes to the great diversity of rRNA pseudouridylation in D.melanogaster.

Animals↗

Large cluster formation through multiple substitution with lanthanide cations (La, Ce, Nd, Sm, Eu, and Gd) of the polyoxoanion.

Several new large polyoxotungstates have been synthesized by reaction of lanthanide cations with the well-known "As(4)W(40)" anion, [(B-alpha-AsO(3)W(9)O(30))(4)(WO(2))(4)](28-) (1). The heteropolyanions [(H(2)O)(11)Ln(III)(Ln(III)(2)OH)(B-alpha-AsO(3)W(9)O(30))(4)(WO(2))(4)](20)(-) (Ln = Ce, Nd, Sm, Gd) (2-4) (Ln(3)As(4)W(40)) and [M(m)()(H(2)O)(10)(Ln(III)(2)OH)(2)(B-alpha-AsO(3)W(9)O(30))(4)(WO(2))(4)]((18-m)(-)) (Ln = La, Ce, Gd and M = Ba, K, none) (5-7) (Ln(4)As(4)W(40)) have been isolated as alkali metal and ammonium salts, respectively, and characterized by single-crystal X-ray analysis, elemental analysis, and IR and (183)W-NMR spectroscopy. The X-ray analyses revealed interanionic W-O-Ln bonds between adjacent Ln(x)()As(4)W(40) units forming a "dimer" for x = 3 and chains for x = 4. Upon dissolving in water these bonds hydrolyze and the monomeric species form. The straightforward syntheses which require the use of concentrated NaCl solutions (1-4 M) and the addition of stoichiometric amounts of Ba(2+) or K(+) reemphasize the importance of the presence of appropriate countercations for the assembly of large polyoxometalate structures.

Journal Article↗

Electron affinities of Al(n) clusters and multiple-fold aromaticity of the square Al4(2-) structure.

The concept of aromaticity was first invented to account for the unusual stability of planar organic molecules with 4n + 2 delocalized pi electrons. Recent photoelectron spectroscopy experiments on all-metal MAl(4)(-) systems with an approximate square planar Al(4)(2-) unit and an alkali metal led to the suggestion that Al(4)(2-) is aromatic. The square Al(4)(2-) structure was recognized as the prototype of a new family of aromatic molecules. High-level ab initio calculations based on extrapolating CCSD(T)/aug-cc-pVxZ (x = D, T, and Q) to the complete basis set limit were used to calculate the first electron affinities of Al(n)(), n = 0-4. The calculated electron affinities, 0.41 eV (n = 0), 1.51 eV (n = 1), 1.89 eV (n = 3), and 2.18 eV (n = 4), are all in excellent agreement with available experimental data. On the basis of the high-level ab initio quantum chemical calculations, we can estimate the resonance energy and show that it is quite large, large enough to stabilize Al(4)(2-) with respect to Al(4). Analysis of the calculated results shows that the aromaticity of Al(4)(2-) is unusual and different from that of C(6)H(6). Particularly, compared to the usual (1-fold) pi aromaticity in C(6)H(6), which may be represented by two Kekulé structures sharing a common sigma bond framework, the square Al(4)(2-) structure has an unusual "multiple-fold" aromaticity determined by three independent delocalized (pi and sigma) bonding systems, each of which satisfies the 4n + 2 electron counting rule, leading to a total of 4 x 4 x 4 = 64 potential resonating Kekulé-like structures without a common sigma frame. We also discuss the 2-fold aromaticity (pi plus sigma) of the Al(3)(-) anion, which can be represented by 3 x 3 = 9 potential resonating Kekulé-like structures, each with two localized chemical bonds. These results lead us to suggest a general approach (applicable to both organic and inorganic molecules) for examining delocalized chemical bonding. The possible electronic contribution to the aromaticity of a molecule should not be limited to only one particular delocalized bonding system satisfying a certain electron counting rule of aromaticity. More than one independent delocalized bonding system can simultaneously satisfy the electron counting rule of aromaticity, and therefore, a molecular structure could have multiple-fold aromaticity.

Journal Article↗

A 3-Mb sequence-ready contig map encompassing the multiple disease gene cluster on chromosome 11q13.1-q13.3.

Despite the presence of several human disease genes on chromosome 11q13, few of them have been molecularly cloned. Here, we report the construction of a contig map encompassing 11q13.1-q13.3 using bacteriophage P1 (P1), bacterial artificial chromosome (BAC), and P1-derived artificial chromosome (PAC). The contig map comprises 32 P1 clones, 27 BAC clones, 6 PAC clones, and 1 YAC clone and spans a 3-Mb region from D11S480 to D11S913. The map encompasses all the candidate loci of Bardet-Biedle syndrome type I (BBS1) and spinocerebellar ataxia type 5 (SCA5), one-third of the distal region for hereditary paraganglioma 2 (PGL2), and one-third of the central region for insulin-dependent diabetes mellitus 4 (IDDM4). In the process of map construction, 61 new sequence-tagged site (STS) markers were developed from the Not I linking clones and the termini of clone inserts. We have also mapped 30 ESTs on this map. This contig map will facilitate the isolation of polymorphic markers for a more refined analysis of the disease gene region and identification of candidate genes by direct cDNA selection, as well as prediction of gene function from sequence information of these bacterial clones.

Chromosome Mapping↗

Multiple risk factor clustering of hypertension in a screened cohort.

OBJECTIVE: A family history of hypertension, obesity, diabetes mellitus, hypercholesterolaemia and hypertriglyceridaemia have all been associated with the risk for hypertension. We evaluated whether the clustering of these risk factors increases the risk for hypertension or whether the accumulation of risk factors is associated with the blood pressure level in non-hypertensive subjects. METHODS AND SUBJECTS: We assessed the clinical data and family history of hypertension (in parents and siblings) for 9914 individuals (6163 men and 3751 women, 18-89 years old) who were screened in Okinawa, Japan, in 1997. RESULTS: In 9914 subjects (2465 hypertensive and 7449 non-hypertensive subjects), all the five factors were positively associated with hypertension. The odds ratios (95% confidence interval) for the number of risk factors were 1.88 (1.62-2.18) for one risk factor, 3.06 (2.62-3.57) for two, 5.25 (4.37-6.30) for three, 8.71 (6.48-11.72) for four and 24.48 (8.49-70.56) for five, after adjusting for age, sex, alcohol consumption, cigarette smoking and physical exercise habits. In non-hypertensive subjects, multivariate regression analyses showed that the number of risks was positively correlated with blood pressure; the regression coefficient was 1.96 (P < 0.0001) for systolic blood pressure, and 1.47 (P < 0.0001) for diastolic blood pressure after adjusting for age and sex. CONCLUSIONS: Clustering of risk factors was significantly associated with hypertension. The number of risk factors positively correlated with the blood pressure levels in nonhypertensive subjects. The accumulation of risk factors may play an important role in the pathogenesis of hypertension, and thus the aggregation of risk factors may need to be addressed in primary prevention efforts related to hypertension.

Adult↗

'Streptomyces nanchangensis', a producer of the insecticidal polyether antibiotic nanchangmycin and the antiparasitic macrolide meilingmycin, contains multiple polyketide gene clusters.

Several independent gene clusters containing varying lengths of type I polyketide synthase genes were isolated from 'Streptomyces nanchangensis' NS3226, a producer of nanchangmycin and meilingmycin. The former is a polyether compound similar to dianemycin and the latter is a macrolide compound similar to milbemycin, which shares the same macrolide ring as avermectin but has different side groups. Clusters A-H spanned about 133, 132, 104, 174, 122, 54, 37 and 59 kb, respectively. Two systems were developed for functional analysis of the gene clusters by gene disruption or replacement. (1) Streptomyces phage phiC31 and its derived vectors can infect and lysogenize this strain. (2) pSET152, an Escherichia coli plasmid with phiC31 attP site, and pHZ1358, a Streptomyces-Escherichia coli shuttle cosmid vector, both carrying oriT from RP4, can be mobilized from E. coli into NS3226 by conjugation. pHZ1358 was shown to be generally useful for generating mutant strains by gene disruption and replacement in NS3226 as well as in several other Streptomyces strains. A region in cluster A (approximately 133 kb) seemed to be involved in nanchangmycin production because replacement of several DNA fragments in this region by an apramycin resistance gene [aac3(IV)] gave rise to nanchangmycin non-producing mutants.

Anti-Bacterial Agents↗

Phylogeny of 54 representative strains of species in the family Pasteurellaceae as determined by comparison of 16S rRNA sequences.

Virtually complete 16S rRNA sequences were determined for 54 representative strains of species in the family Pasteurellaceae. Of these strains, 15 were Pasteurella, 16 were Actinobacillus, and 23 were Haemophilus. A phylogenetic tree was constructed based on sequence similarity, using the Neighbor-Joining method. Fifty-three of the strains fell within four large clusters. The first cluster included the type strains of Haemophilus influenzae, H. aegyptius, H. aphrophilus, H. haemolyticus, H. paraphrophilus, H. segnis, and Actinobacillus actinomycetemcomitans. This cluster also contained A. actinomycetemcomitans FDC Y4, ATCC 29522, ATCC 29523, and ATCC 29524 and H. aphrophilus NCTC 7901. The second cluster included the type strains of A. seminis and Pasteurella aerogenes and H. somnus OVCG 43826. The third cluster was composed of the type strains of Pasteurella multocida, P. anatis, P. avium, P. canis, P. dagmatis, P. gallinarum, P. langaa, P. stomatis, P. volantium, H. haemoglobinophilus, H. parasuis, H. paracuniculus, H. paragallinarum, and A. capsulatus. This cluster also contained Pasteurella species A CCUG 18782, Pasteurella species B CCUG 19974, Haemophilus taxon C CAPM 5111, H. parasuis type 5 Nagasaki, P. volantium (H. parainfluenzae) NCTC 4101, and P. trehalosi NCTC 10624. The fourth cluster included the type strains of Actinobacillus lignieresii, A. equuli, A. pleuropneumoniae, A. suis, A. ureae, H. parahaemolyticus, H. parainfluenzae, H. paraphrohaemolyticus, H. ducreyi, and P. haemolytica. This cluster also contained Actinobacillus species strain CCUG 19799 (Bisgaard taxon 11), A. suis ATCC 15557, H. ducreyi ATCC 27722 and HD 35000, Haemophilus minor group strain 202, and H. parainfluenzae ATCC 29242. The type strain of P. pneumotropica branched alone to form a fifth group. The branching of the Pasteurellaceae family tree was quite complex. The four major clusters contained multiple subclusters. The clusters contained both rapidly and slowly evolving strains (indicated by differing numbers of base changes incorporated into the 16S rRNA sequence relative to outgroup organisms). While the results presented a clear picture of the phylogenetic relationships, the complexity of the branching will make division of the family into genera a difficult and somewhat subjective task. We do not suggest any taxonomic changes at this time.

Base Sequence↗

A novel sodium bicarbonate cotransporter-like gene in an ancient duplicated region: SLC4A9 at 5q31.

BACKGROUND: Sodium bicarbonate cotransporter (NBC) genes encode proteins that execute coupled Na+ and HCO3- transport across epithelial cell membranes. We report the discovery, characterization, and genomic context of a novel human NBC-like gene, SLC4A9, on chromosome 5q31. RESULTS: SLC4A9 was initially discovered by genomic sequence annotation and further characterized by sequencing of long-insert cDNA library clones. The predicted protein of 990 amino acids has 12 transmembrane domains and high sequence similarity to other NBCs. The 23-exon gene has 14 known mRNA isoforms. In three regions, mRNA sequence variation is generated by the inclusion or exclusion of portions of an exon. Noncoding SLC4A9 cDNAs were recovered multiple times from different libraries. The 3' untranslated region is fragmented into six alternatively spliced exons and contains expressed Alu, LINE and MER repeats. SLC4A9 has two alternative stop codons and six polyadenylation sites. Its expression is largely restricted to the kidney. In silico approaches were used to characterize two additional novel SLC4A genes and to place SLC4A9 within the context of multiple paralogous gene clusters containing members of the epidermal growth factor (EGF), ankyrin (ANK) and fibroblast growth factor (FGF) families. Seven human EGF-SLC4A-ANK-FGF clusters were found. CONCLUSION: The novel sodium bicarbonate cotransporter-like gene SLC4A9 demonstrates abundant alternative mRNA processing. It belongs to a growing class of functionally diverse genes characterized by inefficient highly variable splicing. The evolutionary history of the EGF-SLC4A-ANK-FGF gene clusters involves multiple rounds of duplication, apparently followed by large insertions and deletions at paralogous loci and genome-wide gene shuffling.

Adult↗

RT-PCR: characterization of long multi-gene operons and multiple transcript gene clusters in bacteria.

Reverse transcription (RT)-PCR is a valuable tool widely used for analysis of gene expression. In bacteria, RT-PCR is helpful beyond standard protocols of northern blot RNA/DNA hybridization (to identify transcripts) and primer extension (to locate their start points), as these methods have been difficult with transcripts that are low in abundance or unstable, similar to long multi-gene operons. In this report, RT-PCR is adapted to analyze transcripts that form long multi-gene operons--where they start and where they stop. The transcripts can also be semiquantitated to follow the expression of genes under different growth conditions. Examples using RT-PCR are presented with two different multi-gene systems for metal cation resistance to silver and mercury ions. The silver resistance system [9 open reading frames (ORFs); 12.5 kb] is shown by RT-PCR to synthesize three nonoverlapping messenger RNAs that are transcribed divergently. In the mercury resistance system (8 ORFs; 6.3 kb), all the genes are transcribed in the same orientation, and two promoter sites produce overlapping transcripts. For RT-PCR, reverse transcriptase enzyme is used to synthesize first-strand cDNA that is used as a template for PCR amplification of single-gene products, from the beginning, middle or end of long multi-gene, multi-transcript gene clusters.

Bacillus cereus↗

Multiple risk factor clustering and risk of hypertension in Japanese male office workers.

Major risk factors associated with hypertension (a family history of hypertension, obesity, diabetes mellitus, hypercholesterolemia, hypertriglyceridemia, hyperuricemia, and increased white blood cell counts) were assessed in 5275 Japanese male office workers aged 23-59 years. After controlling for potential risk factors of hypertension, the odds ratio of hypertension compared with the absence of risk factors was 1.91, 2.65, 3.88, 6.54, and 8.18 for the presence of 1, 2, 3, 4, and > or = 5 risk factors, respectively (P for trend < 0.001). Systolic and diastolic blood pressure levels also increased in a dose-dependent manner as the number of risk factors increased. Among men not taking antihypertensive medication, the adjusted mean differences in systolic and diastolic blood pressures (mmHg) were 11.2 and 9.2 between men with the presence of > or = 5 risk factors and men without risk factors, respectively. These results indicate that the accumulation of risk factors is highly associated with the increased risk of hypertension in Japanese men.

Adult↗

Multiple snoRNA gene clusters from Arabidopsis.

Small nucleolar RNAs (snoRNAs) are involved in precursor ribosomal RNA (pre-rRNA) processing and rRNA base modification (2'-O-ribose methylation and pseudouridylation). In all eukaryotes, certain snoRNAs (e.g., U3) are transcribed from classical promoters. In vertebrates, the majority are encoded in introns of protein-coding genes, and are released by exonucleolytic cleavage of linearized intron lariats. In contrast, in maize and yeast, nonintronic snoRNA gene clusters are transcribed as polycistronic pre-snoRNA transcripts from which individual snoRNAs are processed. In this article, 43 clusters of snoRNA genes, an intronic snoRNA, and 10 single genes have been identified by cloning and by computer searches, giving a total of 136 snoRNA gene copies of 71 different snoRNA genes. Of these, 31 represent snoRNA genes novel to plants. A cluster of four U14 snoRNA genes and two clusters containing five different snoRNA genes (U31, snoR4, U33, U51, and snoR5) from Arabidopsis have been isolated and characterized. Of these genes, snoR4 is a novel box C/D snoRNA that has the potential to base pair with the 3' end of 5.8S rRNA and snoR5 is a box H/ACA snoRNA gene. In addition, 42 putative sites of 2'-O-ribose methylation in plant 5.8S, 18S, and 25S rRNAs have been mapped by primer extension analysis, including eight sites novel to plant rRNAs. The results clearly show that, in plants, the most common gene organization is polycistronic and that over a third of predicted and mapped methylation sites are novel to plant rRNAs. The variation in this organization among gene clusters highlights mechanisms of snoRNA evolution.

Arabidopsis↗

Clustering ensembles: models of consensus and weak partitions.

Clustering ensembles have emerged as a powerful method for improving both the robustness as well as the stability of unsupervised classification solutions. However, finding a consensus clustering from multiple partitions is a difficult problem that can be approached from graph-based, combinatorial, or statistical perspectives. This study extends previous research on clustering ensembles in several respects. First, we introduce a unified representation for multiple clusterings and formulate the corresponding categorical clustering problem. Second, we propose a probabilistic model of consensus using a finite mixture of multinomial distributions in a space of clusterings. A combined partition is found as a solution to the corresponding maximum-likelihood problem using the EM algorithm. Third, we define a new consensus function that is related to the classical intraclass variance criterion using the generalized mutual information definition. Finally, we demonstrate the efficacy of combining partitions generated by weak clustering algorithms that use data projections and random data splits. A simple explanatory model is offered for the behavior of combinations of such weak clustering components. Combination accuracy is analyzed as a function of several parameters that control the power and resolution of component partitions as well as the number of partitions. We also analyze clustering ensembles with incomplete information and the effect of missing cluster labels on the quality of overall consensus. Experimental results demonstrate the effectiveness of the proposed methods on several real-world data sets.

Algorithms↗

Prevalence and clustering patterns of human papillomavirus genotypes in multiple infections.

Prevalence of multiple human papillomavirus (HPV) infections, involvement of specific HPV phylogenetic clades in multiple infections, and clustering patterns of multiple infections at the clade level were assessed in 854 HIV (-) and 275 HIV (+) women cross-sectionally. Reverse line blot assay was used to detect 27 HPV genotypes. Involvement of specific clades in coinfections and clustering patterns were assessed using HPV clade/genotype as the unit of analyses. Expected frequencies assuming independence for all possible clade combinations in two-genotype infections were derived using a multinomial expansion and comparisons of observed and expected frequencies were done using a composite goodness-of-fit test. In all, 100 two-genotype infections were detected; 61 in HIV (-) and 39 in HIV (+) women. Clade A9 (HPV types 16, 31, 33, 35, 52, and 58) was significantly less likely to be involved in multiple infections compared with all other clades (55.2% versus 64.6%; adjusted odds ratios, 0.68; 95% confidence interval, 0.48-0.95). Observed patterns for all possible clade combinations (among HPV clades A3, A5, A6, A7, A9, and A10) in two-genotype infections did not significantly differ from those expected in the entire sample, across HIV, Pap smear, and age strata (all goodness-of-fit exact P > 0.20). These results indicate that clade A9 is less likely to be involved in multiple infections and that HPV genotypes predominantly establish multiple infections at random, with little positive/negative clustering for either phylogenetically related or unrelated types. The current method of analysis affords the opportunity to test clustering of a large number of HPV genotype/clade combinations at nominal alpha levels.

Adult↗

Clustering of residence of multiple sclerosis patients at age 13 to 20 years in Hordaland, Norway.

Geographic and temporal variation and migration studies point to an exogenous agent in the etiology of multiple sclerosis. If infectious etiology is involved, space-time clustering would also be expected. The authors analyzed 381 patients with a clinical onset of multiple sclerosis between 1953 and 1987 in the county of Hordaland, Norway. Patients within the same birth cohort had lived significantly closer to each other than would be expected during ages 13-20 years, with peak clustering at age 18 years (p = 0.002). Clustering was also shown between patients in pairs comprised of one individual with initial remittent disease and the other with chronic progressive course of disease, suggesting a similar etiology for both clinical patterns. Clustering between cases with widely divergent dates of clinical onset provides evidence of marked variation in latency. No similar clustering was observed in age-, sex-, and area-matched hospital controls without multiple sclerosis, and no clustering was found among the cases when using fixed number of years before onset. These results are compatible with a common infectious agent, such as the Epstein-Barr virus, acquired in adolescence in genetically vulnerable persons who are also not protected by an infection acquired before this age of susceptibility. Susceptibility could be related to the route of transmission or to other age-related covariates or it may be hormonally mediated.

Adolescent↗

Familial clustering of Hodgkin lymphoma and multiple sclerosis.

BACKGROUND: Epidemiologic similarities between Hodgkin lymphoma in young adults (i.e., between 15 and 44 years old) and multiple sclerosis have led to the suggestion that these diseases may have related etiologies. Previous investigations have not supported this hypothesis, but the negative results could have been caused by methodologic problems. We therefore assessed the risk of developing Hodgkin lymphoma for patients with multiple sclerosis and for their families and the risk of developing multiple sclerosis for patients with Hodgkin lymphoma and for their families. METHODS: We identified 11,790 patients with multiple sclerosis and 19,599 of their first-degree relatives in Danish population-based registers and followed them for the occurrence of Hodgkin lymphoma. Analogously, we identified 4381 patients with Hodgkin lymphoma and 7388 of their first-degree relatives and followed them for the occurrence of multiple sclerosis. The relative risks (RRs) of Hodgkin lymphoma and multiple sclerosis were expressed as standardized incidence ratios (i.e., the ratio between observed and expected numbers of outcomes based on age, sex, and period-specific incidence rates). All statistical tests were two-sided. RESULTS: Overall, six cases of Hodgkin lymphoma were identified in patients with multiple sclerosis (RR for Hodgkin lymphoma = 1.40, 95% confidence interval [CI] = 0.63 to 3.12), two of which occurred in young adults (RR = 1.59, 95% CI = 0.40 to 6.37). The risk of young-adult-onset Hodgkin lymphoma was statistically significantly increased in the first-degree relatives of patients with multiple sclerosis (RR = 1.93, 95% CI = 1.01 to 3.71; n = 9 such lymphomas). Two cases of multiple sclerosis were identified among young adult patients with Hodgkin lymphoma (RR for multiple sclerosis = 0.82, 95% CI = 0.20 to 3.27), and the risk for multiple sclerosis was statistically significantly increased in their first-degree relatives (RR = 2.76, 95% CI = 1.44 to 5.31; n = 9 such multiple sclerosis cases). CONCLUSION: The observed familial clustering of multiple sclerosis and young-adult-onset Hodgkin lymphoma is consistent with the hypothesis that the two conditions share environmental and/or constitutional etiologies.

Adolescent↗

Deriving non-homogeneous DNA Markov chain models by cluster analysis algorithm minimizing multiple alignment entropy.

Non-homogeneous Markov chain models can represent biologically important regions of DNA sequences. The statistical pattern that is described by these models is usually weak and was found primarily because of strong biological indications. The general method for extracting similar patterns is presented in the current paper. The algorithm incorporates cluster analysis, multiple alignment and entropy minimization. The method was first tested using the set of DNA sequences produced by Markov chain generators. It was shown that artificial gene sequences, which initially have been randomly set up along the multiple alignment panels, are aligned according to the hidden triplet phase. Then the method was applied to real protein-coding sequences and the resulting alignment clearly indicated the triplet phase and produced the parameters of the optimal 3-periodic non-homogeneous Markov chain model. These Markov models were already employed in the GeneMark gene prediction algorithm, which is used in genome sequencing projects. The algorithm can also handle the case in which the sequences to be aligned reveal different statistical patterns, such as Escherichia coli protein-coding sequences belonging to Class II and Class III. The algorithm accepts a random mix of sequences from different classes, and is able to separate them into two groups (clusters), align each cluster separately, and define a non-homogeneous Markov chain model for each sequence cluster.

Algorithms↗