Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article↗

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA↗

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals↗

EnteriX 2003: Visualization tools for genome alignments of Enterobacteriaceae.

We describe EnteriX, a suite of three web-based visualization tools for graphically portraying alignment information from comparisons among several fixed and user-supplied sequences from related enterobacterial species, anchored on a reference genome (http://bio.cse.psu.edu/). The first visualization, Enteric, displays stacked pairwise alignments between a reference genome and each of the related bacteria, represented schematically as PIPs (Percent Identity Plots). Encoded in the views are large-scale genomic rearrangement events and functional landmarks. The second visualization, Menteric, computes and displays 1 Kb views of nucleotide-level multiple alignments of the sequences, together with annotations of genes, regulatory sites and conserved regions. The third, a Java-based tool named Maj, displays alignment information in two formats, corresponding roughly to the Enteric and Menteric views, and adds zoom-in capabilities. The uses of such tools are diverse, from examining the multiple sequence alignment to infer conserved sites with potential regulatory roles, to scrutinizing the commonalities and differences between the genomes for pathogenicity or phylogenetic studies. The EnteriX suite currently includes >15 enterobacterial genomes, generates views centered on four different anchor genomes and provides support for including user sequences in the alignments.

Computer Graphics↗

Phydbac2: improved inference of gene function using interactive phylogenomic profiling and chromosomal location analysis.

Phydbac (phylogenomic display of bacterial genes) implemented a method of phylogenomic profiling using a distance measure based on normalized BLAST scores. This method was able to increase the predictive power of phylogenomic profiling by about 25% when compared to the classical approach based on Hamming distances. Here we present a major extension of Phydbac (named here Phydbac2), that extends both the concept and the functionality of the original web-service. While phylogenomic profiles remain the central focus of Phydbac2, it now integrates chromosomal proximity and gene fusion analyses as two additional non-similarity-based indicators for inferring pairwise gene functional relationships. Moreover, all presently available (January 2004) fully sequenced bacterial genomes and those of three lower eukaryotes are now included in the profiling process, thus increasing the initial number of reference genomes (71 in Phydbac) to 150 in Phydbac2. Using the KEGG metabolic pathway database as a benchmark, we show that the predictive power of Phydbac2 is improved by 27% over the previous version. This gain is accounted for on one hand, by the increased number of reference genomes (11%) and on the other hand, as a result of including chromosomal proximity into the distance measure (16%). The expanded functionality of Phydbac2 now allows the user to query more than 50 different genomes, including at least one member of each major bacterial group, most major pathogens and potential bio-terrorism agents. The search for co-evolving genes based on consensus profiles from multiple organisms, the display of Phydbac2 profiles side by side with COG information, the inclusion of KEGG metabolic pathway maps the production of chromosomal proximity maps, and the possibility of collecting and processing results from different Phydbac queries in a common shopping cart are the main new features of Phydbac2. The Phydbac2 web server is available at http://igs-server.cnrs-mrs.fr/phydbac/.

Artificial Gene Fusion↗

gaftools: a toolkit for analyzing and manipulating pangenome alignments.

MOTIVATION: Linear reference genomes are ubiquitously used in genomics research, despite known biases associated with their use. In recent years, there has been a shift towards graph-based reference genomes to address some of these biases, which has required development of new algorithms and file formats. This has created a necessity for new tools capable of utilizing these formats and performing operations similar to those carried out by traditional methods. RESULTS: In this paper we present "gaftools," a multi-purpose tool that introduces several utilities for processing graph alignments in GAF format. gaftools enables users to index and sort alignments, with graph ordering serving as a necessary step for the sorting process. Additionally, it allows users to view subsets of alignments and perform realignment using the wavefront alignment algorithm, among other features. Many of these functionalities are inspired by SAMtools, which provides similar operations for linear genomes, while gaftools adapts and extends them for pangenomes. AVAILABILITY: gaftools is available under MIT license at https://github.com/marschall-lab/gaftools.

Software↗

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals↗

Rawsamble: overlapping raw nanopore signals using a hash-based seeding mechanism.

MOTIVATION: Raw nanopore signal analysis is a common approach in genomics to provide fast and resource-efficient analysis without translating the signals to bases (i.e. without basecalling). However, existing solutions cannot interpret raw signals directly if a reference genome is unknown due to a lack of accurate mechanisms to handle increased noise in pairwise raw signal comparison. Our goal is to enable the direct analysis of raw signals without a reference genome. To this end, we propose Rawsamble, the first mechanism that can identify regions of similarity between all raw signal pairs, known as all-vs-all overlapping, using a hash-based search mechanism. RESULTS: We use these overlaps to construct de novo assembly graphs with an existing assembler, miniasm, off-the-shelf. To our knowledge, these are the first de novo assemblies ever constructed directly from raw signals without basecalling. Our extensive evaluations across multiple genomes of varying sizes show that Rawsamble provides a significant speedup (on average by 5.01× and up to 23.10×) and reduces peak memory usage (on average by 5.74× and up to by 22.00×) compared to a conventional genome assembly pipeline using the state-of-the-art tools for basecalling (Dorado's fastest mode) and overlapping (minimap2) on a CPU. We find that around one-third of Rawsamble's overlapping pairs are also found by minimap2. We find that when we use overlapping reads from Rawsamble, we can construct unitigs that are (i) as accurate as those built from minimap2's overlaps and (ii) up to half a chromosome in length (e.g. 2.3 million bases for E. coli). AVAILABILITY AND IMPLEMENTATION: Rawsamble is available at https://github.com/CMU-SAFARI/RawHash. We also provide the scripts to fully reproduce our results on our GitHub page.

Nanopores↗

Sequence context analysis of 8.2 million single nucleotide polymorphisms in the human genome.

We analyzed n-mers (n=3-8) in the local environment of 8,249,446 human SNPs and compared their distribution with that in the genome reference sequences. The results revealed that the short sequences, which contained at least one CpG dinucleotide, occurred more frequently in the local SNP sequences than in the genome sequences. To exclude the hypermutability effect of the methylated CpG dinucleotides on the sequence context of SNPs, we examined the distribution patterns for each of the six categories of substitution. We observed the similar pattern (i.e., CpG-containing n-mers vs. non-CpG-containing n-mers) in SNP categories A/G, C/T and C/G but the opposite pattern in category A/T. We next identified 34,928 putative CpG islands in the human genome and located 133,591 SNPs within these islands. In the CpG islands, CpG SNPs were 3.92-fold less prevalent relative to the presence of CpG dinucleotides. Conversely, in the human genome, the frequency of CpG dinucleotides at the polymorphic sites was 6.09 times that in the genome reference sequences. These results support the previous views of mutational suppression at the CpG sites in the CpG islands and hypermutability of the methylated CpG dinucleotides that are prevalent in the non-CpG island sequences in the human genome. Our study represents a comprehensive investigation of the sequence context of SNPs in the human genome and in human CpG islands.

Base Composition↗

Estimates of genetic diversity in the brown cattle population of Switzerland obtained from pedigree information.

The study investigates the genetic diversity present as well as its development in the Brown Cattle population of Switzerland from pedigree information. The population consisted of three subpopulations, the Braunvieh (BV), the original Braunvieh (OB) and the US-Brown Swiss (BS). The BV is a cross of OB with BS where crossing still continues. The OB is without any genetic influence of BS. The diversity measures effective population size, effective number of ancestors (explaining 99% of reference genome) and founder genome equivalents were calculated for 11 reference populations of animals born in a single year from 1992 onwards. The BS-subpopulation consisted of animals and their known ancestors which were used in the crossing scheme and was, therefore, quite small. The youngest animals were born in 2002, the oldest ones in the 1920s. Average inbreeding was by far the highest in BS, in spite of the lowest quality of pedigrees, and lowest in OB. Effective population size obtained from the difference between average inbreeding of offspring and their parents was, mostly due to the heavy use of few highly inbred BS-sires, strongly overestimated in some BV-reference populations. If this parameter was calculated from the yearly rate of inbreeding and a generation interval of 5 years, no bias was observed and ranking of populations from high to low was OB-BV-BS, i.e. equal to the other diversity parameters. The high genetic diversity found in OB was a consequence of the use of many natural service sires. Rate of decrease of effective number of ancestors was steeper in BV than OB was, however, equal for founder genome equivalents. Founder genome equivalents were more stable than effective population sizes calculated from the difference between average inbreeding of offspring and parents. The five most important ancestors contributed one-third of the 2002-reference genomes of BV and OB, in BV all were BS-sires. The relative amount of BS-genes in the BV-genome increased from 59.2% to 78.5% during the 11 years considered.

Animals↗

Phylogenetic tree information aids supervised learning for predicting protein-protein interaction based on distance matrices.

BACKGROUND: Protein-protein interactions are critical for cellular functions. Recently developed computational approaches for predicting protein-protein interactions utilize co-evolutionary information of the interacting partners, e.g., correlations between distance matrices, where each matrix stores the pairwise distances between a protein and its orthologs from a group of reference genomes. RESULTS: We proposed a novel, simple method to account for some of the intra-matrix correlations in improving the prediction accuracy. Specifically, the phylogenetic species tree of the reference genomes is used as a guide tree for hierarchical clustering of the orthologous proteins. The distances between these clusters, derived from the original pairwise distance matrix using the Neighbor Joining algorithm, form intermediate distance matrices, which are then transformed and concatenated into a super phylogenetic vector. A support vector machine is trained and tested on pairs of proteins, represented as super phylogenetic vectors, whose interactions are known. The performance, measured as ROC score in cross validation experiments, shows significant improvement of our method (ROC score 0.8446) over that of using Pearson correlations (0.6587). CONCLUSION: We have shown that the phylogenetic tree can be used as a guide to extract intra-matrix correlations in the distance matrices of orthologous proteins, where these correlations are represented as intermediate distance matrices of the ancestral orthologous proteins. Both the unsupervised and supervised learning paradigms benefit from the explicit inclusion of these intermediate distance matrices, and particularly so in the latter case, which offers a better balance between sensitivity and specificity in the prediction of protein-protein interactions.

Computational Biology↗

Development and evaluation of a one-pot RPA-Cas12a assay based on a primer-driven reverse screening strategy for preliminary screening of megalocytivirus-related viruses.

A primer-driven reverse-screening strategy was used to identify an RPA-Cas12a target suitable for the rapid preliminary screening of megalocytivirus-related viruses. The ISKNV reference genome NC_003494.1 was used as the initial template, and candidate amplification units were designed according to RPA primer-design requirements, primer physicochemical properties, and the availability of Cas12a protospacer-adjacent motif (PAM) sites and crRNA target sequences. Following preliminary amplification assessment, the retained candidate primers were aligned individually against 75 complete genome sequences of megalocytivirus-related viruses. Of these, 67 sequences met the predefined criteria for target-region integrity, primer-binding-site compatibility, and Cas12a recognition. Retrospective mapping to the reference genome located the candidate amplification region within ORF057L. Based on the resulting candidate detection unit, a one-pot RPA-Cas12a assay incorporating a commercially available lyophilized RPA amplification module was developed. Optimization showed that 400 nM reporter and 80 nM crRNA-1 provided relatively stable fluorescence output. A cut-off value of 1281.6 relative fluorescence units (RFU) was established as the mean plus three standard deviations of the endpoint fluorescence values obtained from 20 qPCR-negative samples. In analytical sensitivity testing, the assay generated fluorescence signals above the negative control at low plasmid copy numbers. However, because only a limited number of replicates were tested at these low template concentrations, these findings were not used to define a formal limit of detection. ISKNV, RSIV, and TRBIV samples tested positive, whereas the MRV sample produced an endpoint fluorescence value below the cut-off. Repeatability analysis of the same sample in six independent reactions yielded a coefficient of variation of 8.03%. Among the 39 samples examined, no discordant qualitative results were observed between the RPA-Cas12a assay and qPCR. These findings support the use of the ORF057L-targeted one-pot RPA-Cas12a assay as a rapid preliminary screening tool for megalocytivirus-related viruses. Nevertheless, its formal limit of detection, inter-batch stability, cross-reactivity with additional non-target pathogens, and clinical diagnostic performance require further evaluation.

Lyophilized RPA↗

Myco-Ed: Mycological curriculum for education and discovery.

Fungi are important and hyperdiverse organisms, yet chronically understudied. Most fungal clades have no reference genomes, impeding our understanding of their ecosystem functions and use as solutions in health and biotechnology. Also, opportunities for training in fungal biology and genomics are lacking, creating a bottleneck that hinders the recruitment and cultivation of a talented future mycological workforce. To address these issues, we developed Myco-Ed, an educational program offering training and scientific contributions through genome sequencing and analysis. Myco-Ed empowers students to pursue careers in fungal biology while improving fungal resources. Myco-Ed has been piloted at 12 institutions (15 classrooms) ranging from online e-Campuses to R1 universities, resulting in hundreds of fungal observations and many new high-quality reference genomes.

Curriculum↗

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny↗

Strains of Peru tomato virus infecting cocona (Solanum sessiliflorum), tomato and pepper in Peru with reference to genome evolution in genus Potyvirus.

Two isolates (SL1 and SL6) of Peru tomato virus (PTV, genus Potyvirus) were obtained from cocona plants (Solanum sessiliflorum) growing in Tingo María, the jungle of the Amazon basin in Peru. One PTV isolate (TM) was isolated from a tomato plant (Lycopersicon esculentum) growing in Huaral at the Peruvian coast. The three PTV isolates were readily transmissible by Myzus persicae. Isolate SL1, but not SL6, caused chlorotic lesions in inoculated leaves of Chenopodium amaranticolor and C. quinoa. Isolate TM differed from SL1 and SL6 in causing more severe mosaic symptoms in tomato, and vein necrosis in the leaves of cocona. Pepper cv. Avelar (Capsicum annuum) showed resistance to the PTV isolates SL1 and SL6 but not TM. The 5'- and 3'-proximal sequences of the three PTV isolates were cloned, sequenced and compared to the corresponding sequences of four PTV isolates from pepper, the only host from which PTV isolates have been previously characterised at the molecular level. Phylogenetic analyses on the P1 protein and coat protein amino acid sequences indicated, in accordance with the phenotypic data from indicator hosts, that the PTV isolates from cocona represented a distinguishable strain. In contrast, the PTV isolates from tomato and pepper were not grouped according to the host. Inclusion of the sequence data from the three PTV isolates of this study in a phylogenetic analysis with other PTV isolates and other potyviruses strengthen the membership of PTV in the so-called "PVY subgroup" of Potyvirus. This subgroup of closely related potyvirus species was also distinguishable from other potyviruses by their more uniform sizes of the protein-encoding regions within the polyprotein.

Animals↗

Evolutionary analysis of Arabidopsis, cyanobacterial, and chloroplast genomes reveals plastid phylogeny and thousands of cyanobacterial genes in the nucleus.

Chloroplasts were once free-living cyanobacteria that became endosymbionts, but the genomes of contemporary plastids encode only approximately 5-10% as many genes as those of their free-living cousins, indicating that many genes were either lost from plastids or transferred to the nucleus during the course of plant evolution. Previous estimates have suggested that between 800 and perhaps as many as 2,000 genes in the Arabidopsis genome might come from cyanobacteria, but genome-wide phylogenetic surveys that could provide direct estimates of this number are lacking. We compared 24,990 proteins encoded in the Arabidopsis genome to the proteins from three cyanobacterial genomes, 16 other prokaryotic reference genomes, and yeast. Of 9,368 Arabidopsis proteins sufficiently conserved for primary sequence comparison, 866 detected homologues only among cyanobacteria and 834 other branched with cyanobacterial homologues in phylogenetic trees. Extrapolating from these conserved proteins to the whole genome, the data suggest that approximately 4,500 of Arabidopsis protein-coding genes ( approximately 18% of the total) were acquired from the cyanobacterial ancestor of plastids. These proteins encompass all functional classes, and the majority of them are targeted to cell compartments other than the chloroplast. Analysis of 15 sequenced chloroplast genomes revealed 117 nuclear-encoded proteins that are also still present in at least one chloroplast genome. A phylogeny of chloroplast genomes inferred from 41 proteins and 8,303 amino acids sites indicates that at least two independent secondary endosymbiotic events have occurred involving red algae and that amino acid composition bias in chloroplast proteins strongly affects plastid genome phylogeny.

Arabidopsis↗

CGHScan: finding variable regions using high-density microarray comparative genomic hybridization data.

BACKGROUND: Comparative genomic hybridization can rapidly identify chromosomal regions that vary between organisms and tissues. This technique has been applied to detecting differences between normal and cancerous tissues in eukaryotes as well as genomic variability in microbial strains and species. The density of oligonucleotide probes available on current microarray platforms is particularly well-suited for comparisons of organisms with smaller genomes like bacteria and yeast where an entire genome can be assayed on a single microarray with high resolution. Available methods for analyzing these experiments typically confine analyses to data from pre-defined annotated genome features, such as entire genes. Many of these methods are ill suited for datasets with the number of measurements typical of high-density microarrays. RESULTS: We present an algorithm for analyzing microarray hybridization data to aid identification of regions that vary between an unsequenced genome and a sequenced reference genome. The program, CGHScan, uses an iterative random walk approach integrating multi-layered significance testing to detect these regions from comparative genomic hybridization data. The algorithm tolerates a high level of noise in measurements of individual probe intensities and is relatively insensitive to the choice of method for normalizing probe intensity values and identifying probes that differ between samples. When applied to comparative genomic hybridization data from a published experiment, CGHScan identified eight of nine known deletions in a Brucella ovis strain as compared to Brucella melitensis. The same result was obtained using two different normalization methods and two different scores to classify data for individual probes as representing conserved or variable genomic regions. The undetected region is a small (58 base pair) deletion that is below the resolution of CGHScan given the array design employed in the study. CONCLUSION: CGHScan is an effective tool for analyzing comparative genomic hybridization data from high-density microarrays. The algorithm is capable of accurately identifying known variable regions and is tolerant of high noise and varying methods of data preprocessing. Statistical analysis is used to define each variable region providing a robust and reliable method for rapid identification of genomic differences independent of annotated gene boundaries.

Algorithms↗