Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Robust prediction of consensus secondary structures using averaged base pairing probability matrices.

MOTIVATION: Recent transcriptomic studies have revealed the existence of a considerable number of non-protein-coding RNA transcripts in higher eukaryotic cells. To investigate the functional roles of these transcripts, it is of great interest to find conserved secondary structures from multiple alignments on a genomic scale. Since multiple alignments are often created using alignment programs that neglect the special conservation patterns of RNA secondary structures for computational efficiency, alignment failures can cause potential risks of overlooking conserved stem structures. RESULTS: We investigated the dependence of the accuracy of secondary structure prediction on the quality of alignments. We compared three algorithms that maximize the expected accuracy of secondary structures as well as other frequently used algorithms. We found that one of our algorithms, called McCaskill-MEA, was more robust against alignment failures than others. The McCaskill-MEA method first computes the base pairing probability matrices for all the sequences in the alignment and then obtains the base pairing probability matrix of the alignment by averaging over these matrices. The consensus secondary structure is predicted from this matrix such that the expected accuracy of the prediction is maximized. We show that the McCaskill-MEA method performs better than other methods, particularly when the alignment quality is low and when the alignment consists of many sequences. Our model has a parameter that controls the sensitivity and specificity of predictions. We discussed the uses of that parameter for multi-step screening procedures to search for conserved secondary structures and for assigning confidence values to the predicted base pairs. AVAILABILITY: The C++ source code that implements the McCaskill-MEA algorithm and the test dataset used in this paper are available at http://www.ncrna.org/papers/McCaskillMEA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Detection and analysis of alternative splicing in the silkworm by aligning expressed sequence tags with the genomic sequence.

We identified 277 alternative splice forms in silkworm genes based on aligning expressed sequence tags with genomic sequences, using a transcript assembly program. A large fraction (74%) of these alternative splices are located in protein-coding regions and alter protein products, whereas only 26% are in untranslated regions. From the alternative splices located in protein-coding regions, some (43%) affect protein domains that bind various biological molecules. The vast majority of the detected alternative forms in this study appear to be novel, and potentially affect biologically meaningful control of function in silkworm genes. Our results indicate that alternative splicing in silkworm largely produces protein diversity and functional diversity, and is a widely used mechanism for regulating gene expression.

Alternative Splicing↗

The complete nucleotide sequence, genome organization, and origin of human adenovirus type 11.

The complete DNA sequence and transcription map of human adenovirus type 11 are reported here. This is the first published sequence for a subgenera B human adenovirus and demonstrates a genome organization highly similar to those of other human adenoviruses. All of the genes from the early, intermediate, and late regions are present in the expected locations of the genome for a human adenovirus. The genome size is 34,794 bp in length and has a GC content of 48.9%. Sequence alignment with genomes of groups A (Ad12), C (Ad5), D (Ad17), E (Simian adenovirus 25), and F (Ad40) revealed homologies of 64, 54, 68, 75, and 52%, respectively. Detailed genomic analysis demonstrated that Ads 11 and 35 are highly conserved in all areas except the hexon hypervariable regions and fiber. Similarly, comparison of Ad11 with subgroup E SAV25 revealed poor homology between fibers but high homology in proteins encoded by all other areas of the genome. We propose an evolutionary model in which functional viruses can be reconstituted following fiber substitution from one serotype to another. According to this model either the Ad11 genome is a derivative of Ad35, from which the fiber was substituted with Ad7, or the Ad35 genome is the product of a fiber substitution from Ad21 into the Ad11 genome. This model also provides a possible explanation for the origin of group E Ads, which are evolutionarily derived from a group C fiber substitution into a group B genome.

Adenoviruses, Human↗

Alignment of a restriction map with the genetic map of bacteriophage T4.

A restriction map of the bacteriophage T4 genome was aligned with the T4 genetic map. Included were the cleavage sites for BamHI, BglII, KpnI, PvuI, SalI, and XbaI. The alignment utilized the fact that the T4 genetic map had been oriented previously with respect to a T2/T4 heteroduplex map. DNA fragments from a BglII digestion of cytosine-containing DNA from a T4 dCTPase- denA denB(rIIH23B) alc mutant were hybridized with full-length chromosomal strands of bacteriophage T2, and the heteroduplexes were examined by electron microscopy. From their lengths and patterns of substitution and deletion loops, the heteroduplexes formed with 6 of the 13 BglII fragments could be unambiguously identified and positioned on the T2/T4 heteroduplex map. The ends of the T4 DNA strands in the heteroduplexes directly identified the location of 10 BglII cleavage sites. The remaining three BglII cleavage sites could be assigned to the T2/T4 heteroduplex map based on their relative locations on the restriction map. It was also possible to identify the source of the DNA strands (i.e., T2 or T4) in four previously unassigned deletion loops on the T2/T4 heteroduplex. Among the BglII fragments identified in heteroduplexes was the fragment containing the rIIH23B deletion; this deletion was used as the primary point of reference for alignment of the T4 restriction map with the T2/T4 heteroduplex map and, hence, with the T4 genetic map.

Chromosome Mapping↗

AVID: A global alignment program.

In this paper we describe a new global alignment method called AVID. The method is designed to be fast, memory efficient, and practical for sequence alignments of large genomic regions up to megabases long. We present numerous applications of the method, ranging from the comparison of assemblies to alignment of large syntenic genomic regions and whole genome human/mouse alignments. We have also performed a quantitative comparison of AVID with other popular alignment tools. To this end, we have established a format for the representation of alignments and methods for their comparison. These formats and methods should be useful for future studies. The tools we have developed for the alignment comparisons, as well as the AVID program, are publicly available. See Web Site References section for AVID Web address and Web addresses for other programs discussed in this paper.

Algorithms↗

PipTools: a computational toolkit to annotate and analyze pairwise comparisons of genomic sequences.

Sequence conservation between species is useful both for locating coding regions of genes and for identifying functional noncoding segments. Hence interspecies alignment of genomic sequences is an important computational technique. However, its utility is limited without extensive annotation. We describe a suite of software tools, PipTools, and related programs that facilitate the annotation of genes and putative regulatory elements in pairwise alignments. The alignment server PipMaker uses the output of these tools to display detailed information needed to interpret alignments. These programs are provided in a portable format for use on common desktop computers and both the toolkit and the PipMaker server can be found at our Web site (http://bio.cse.psu.edu/). We illustrate the utility of the toolkit using annotation of a pairwise comparison of the mouse MHC class II and class III regions with orthologous human sequences and subsequently identify conserved, noncoding sequences that are DNase I hypersensitive sites in chromatin of mouse cells.

Animals↗

theBIGbam: compression and interactive exploration of large-scale sequencing alignments with circular mapping support.

SUMMARY: theBIGbam (github.com/bhagavadgitadu22/theBIGbam) is a genome browser and alignment viewer designed for massive metagenomic and metatranscriptomic datasets. The tool takes BAM files containing read alignments, together with genome assemblies in FASTA format or annotated genome sequences in GenBank format. Alternatively, it can start from raw FASTQ reads and generate alignments using a modified mapper that supports circular genomes, enabling seamless read mapping across genome ends. theBIGbam can compress hundreds of gigabytes of input files 10- to 100-fold into dedicated databases while retaining key per-position information, including coverage depth and recurrent mismatches, insertions, and deletions between reads and the reference. These databases can be served to a local web browser, enabling interactive exploration of any contig in any sample using DNAFeaturesViewer for genome maps and Bokeh for mapping-derived features. Contig-sample pairs available for visualization can be filtered using a range of summary metrics calculated per contig, per sample, and per contig-sample pair to guide users toward the most relevant signals. Through its interactive visualization, theBIGbam facilitates the exploration of complex datasets, while its integrated database-combining assembly features, annotated features, and mapping-derived features-provides the information needed to investigate biological hypotheses systematically. Designed to complement existing browsing tools like IGV and Anvi'o, theBIGbam is particularly suited for examining misassemblies, subpopulations, microdiversity, and contig topology in large-scale datasets. AVAILABILITY AND IMPLEMENTATION: theBIGbam is an open-source Rust/Python package that can be installed from Bioconda or PyPI. The source code and documentation are available on GitHub (github.com/bhagavadgitadu22/theBIGbam).

Software↗

Alignment of Escherichia coli K12 DNA sequences to a genomic restriction map.

We use the extensive published information describing the genome of Escherichia coli and new restriction map alignment software to align DNA sequence, genetic, and physical maps. Restriction map alignment software is used which considers restriction maps as strings analogous to DNA or protein sequences except that two values, enzyme name and DNA base address, are associated with each position on the string. The resulting alignments reveal a nearly linear relationship between the physical and genetic maps of the E. coli chromosome. Physical map comparisons with the 1976, 1980, and 1983 genetic maps demonstrate a better fit with the more recent maps. The results of these alignments are genomic kilobase coordinates, orientation and rank of the alignment that best fits the genetic data. A statistical measure based on extreme value distribution is applied to the alignments. Additional computer analyses allow us to estimate the accuracy of the published E. coli genomic restriction map, simulate rearrangements of the bacterial chromosome, and search for repetitive DNA. The procedures we used are general enough to be applicable to other genome mapping projects.

Amino Acid Sequence↗

Refined annotation of the Arabidopsis genome by complete expressed sequence tag mapping.

Expressed sequence tags (ESTs) currently encompass more entries in the public databases than any other form of sequence data. Thus, EST data sets provide a vast resource for gene identification and expression profiling. We have mapped the complete set of 176,915 publicly available Arabidopsis EST sequences onto the Arabidopsis genome using GeneSeqer, a spliced alignment program incorporating sequence similarity and splice site scoring. About 96% of the available ESTs could be properly aligned with a genomic locus, with the remaining ESTs deriving from organelle genomes and non-Arabidopsis sources or displaying insufficient sequence quality for alignment. The mapping provides verified sets of EST clusters for evaluation of EST clustering programs. Analysis of the spliced alignments suggests corrections to current gene structure annotation and provides examples of alternative and non-canonical pre-mRNA splicing. All results of this study were parsed into a database and are accessible via a flexible Web interface at http://www.plantgdb.org/AtGDB/.

Alternative Splicing↗

Computational methods for alternative splicing prediction.

The fact that a large majority of mammalian genes are subject to alternative splicing indicates that this phenomenon represents a major mechanism for increasing proteome complexity. Here, we provide an overview of current methods for the computational prediction of alternative splicing based on the alignment of genome and transcript sequences. Specific features and limitations of different approaches and software are discussed, particularly those affecting prediction accuracy and assembly of alternative transcripts.

Algorithms↗

Computational analysis of 3'-ends of ESTs shows four classes of alternative polyadenylation in human, mouse, and rat.

Alternative initiation, splicing, and polyadenylation are key mechanisms used by many organisms to generate diversity among mature mRNA transcripts originating from the same transcription unit. While previous computational analyses of alternative polyadenylation have focused on polyadenylation activities within or downstream of the normal 3'-terminal exons, we present the results of the first genome-wide analysis of patterns of alternative polyadenylation in the human, mouse, and rat genomes occurring over the entire transcribed regions of mRNAs using 3'-ESTs with poly(A) tails aligned to genomic sequences. Four distinct classes of patterns of alternative polyadenylation result from this analysis: tandem poly(A) sites, composite exons, hidden exons, and truncated exons. We estimate that at least 49% (human), 31% (mouse), and 28% (rat) of polyadenylated transcription units have alternative polyadenylation. A portion of these alternative polyadenylation events result in new protein isoforms.

Alternative Splicing↗

Identification and classification of conserved RNA secondary structures in the human genome.

The discoveries of microRNAs and riboswitches, among others, have shown functional RNAs to be biologically more important and genomically more prevalent than previously anticipated. We have developed a general comparative genomics method based on phylogenetic stochastic context-free grammars for identifying functional RNAs encoded in the human genome and used it to survey an eight-way genome-wide alignment of the human, chimpanzee, mouse, rat, dog, chicken, zebra-fish, and puffer-fish genomes for deeply conserved functional RNAs. At a loose threshold for acceptance, this search resulted in a set of 48,479 candidate RNA structures. This screen finds a large number of known functional RNAs, including 195 miRNAs, 62 histone 3'UTR stem loops, and various types of known genetic recoding elements. Among the highest-scoring new predictions are 169 new miRNA candidates, as well as new candidate selenocysteine insertion sites, RNA editing hairpins, RNAs involved in transcript auto regulation, and many folds that form singletons or small functional RNA families of completely unknown function. While the rate of false positives in the overall set is difficult to estimate and is likely to be substantial, the results nevertheless provide evidence for many new human functional RNAs and present specific predictions to facilitate their further characterization.

3' Untranslated Regions↗

Space-efficient whole genome comparisons with Burrows-Wheeler transforms.

The starting point for any alignment of mammalian genomes is the computation of exact matches satisfying various criteria. Time-efficient, O(n), data structures for this computation, such as the suffix tree, require O(n log(n)) space, several times the space of the genomes themselves. Thus, any reasonable whole-genome comparative project finds itself requiring tens of Gigabytes of RAM to maintain time-efficiency. This is beyond most modern workstations. With a new data structure, the compressed suffix array (CSA) implemented via the Burrows-Wheeler transform, we can trade time-efficiency for space-efficiency, taking O(n log(n)) time, but running in O(n) space, typically in total space less than or equal to that of the genomes themselves. If space is more expensive than time, this is an appropriate approach to consider. The most space-efficient implementation of this data structure requires 5 bits per nucleotide character to build on-line, in the worst case, and 2.5 bits per character to store once built. We present a description of this data structure and how it is used to obtain matches. An implementation (called bbbwt) is demonstrated by aligning two mammalian genomes on a modest workstation equipped with under 2 GB of free RAM in time superior to that of the implementations of other data structures.

Animals↗

Integration of animal linkage and BAC contig maps using overgo hybridization.

The alignment of genome linkage maps, defined primarily by segregation of sequence-tagged site (STS) markers, with BAC contig physical maps and full genome sequences requires high throughput mechanisms to identify BAC clones that contain specific STS. A powerful technique for this purpose is multi-dimensional hybridization of "overgo" probes. The probes are chosen from available STS sequence data by selecting unique probe sequences that have a common melting temperature. We have hybridized sets of 216 overgo probes in subset pools of 36 overgos at a time to filter-spotted chicken BAC clone arrays. A four-dimensional pooling strategy, including one degree of redundancy, has been employed. This requires 24 hybridizations to completely assign BACs for all 216 probes. Results to date are consistent with about a 10% failure rate in overgo probe design and a 15-20% false negative detection rate within a group of 216 markers. Three complete rounds of overgo hybridization, each to sets of about 39,000 BACs (either BAMHI or ECORI partial digest inserts) generated a total of 1853 BAC alignments for 517 mapped chicken genome STS markers. These data are publicly available, and they have been used in the assembly of a first generation BAC contig map of the chicken genome.

Animals↗

Zone equalisation normalisation for improved alignment of epigenetic signal.

MOTIVATION: High-throughput genomic technologies have transformed our understanding of biological systems, yet direct comparison and visualisation of these complex datasets remains challenging. Existing normalisation methods often fail to align genomic signal across samples due to sensitivity to sequencing depth differences and localised high-signal artefacts, leading to inconsistent replicate behaviour and increased downstream variability. RESULTS: We introduce Zone Equalisation Normalisation (ZEN), a novel approach designed to improve cross-sample signal alignment of genomic data. ZEN rescales genomic signal based on variance estimated within biologically enriched regions, reducing the influence of extreme outliers while preserving underlying biological structure. Using a diverse collection of data and our new genome-wide benchmarking approach, we reveal that ZEN improves biological and technical replicate alignment across the majority of tested conditions and experimental platforms. We further show that this improved signal comparability is associated with fewer differential accessibility calls between technical replicates and a more conservative set of biological differences. Together, these results demonstrate that ZEN provides a complementary framework to improve the accuracy and reliability of genomic data analysis and that normalisation choice can affect downstream analyses and biological interpretation. AVAILABILITY AND IMPLEMENTATION: ZEN is available as an open-source Python package via conda and PyPI. Source code, documentation, tutorials, and code to reproduce the analyses are available at https://github.com/Genome-Function-Initiative-Oxford/Zone-Equalisation-Normalisation and Zenodo (https://doi.org/10.5281/zenodo.21067751).

Epigenesis, Genetic↗

Heterogeneity in regional GC content and differential usage of codons and amino acids in GC-poor and GC-rich regions of the genome of Apis mellifera.

The honeybee (Apis mellifera) has a genome with a wide variation in GC content showing 2 clear modal GC values, in some ways reminiscent of an isochore-like structure. To gain insight into causes and consequences of this pattern, we used a comparative approach to study the genome-wide alignment of primarily coding sequence of A. mellifera with Drosophila melanogaster and Anopheles gambiae. The latter 2 species show a higher average GC content than A. mellifera and no indications of bimodality, suggesting that the GC-poor mode is a derived condition in honeybee. In A. mellifera, synonymous sites of genes generally adopt the GC content of the region in which they reside. A large proportion of genes in GC-poor regions have not been assigned to the honeybee assembly because of the low sequence complexity of their genome neighborhood. The synonymous substitution rate between A. mellifera and the other species is very close to saturation, but analyses of nonsynonymous substitutions as well as amino acid substitutions indicate that the GC-poor regions are not evolving faster than the GC-rich regions. We describe the codon usage and amino acid usage and show that they are remarkably heterogeneous within the honeybee genome between the 2 different GC regions. Specifically, the genes located in GC-poor regions show a much larger deviation in both codon usage bias and amino acid usage from the Dipterans than the genes located in the GC-rich regions.

Amino Acids↗

SVbyEye: a visual tool to characterize structural variation among whole-genome assemblies.

MOTIVATION: We are now in the era of being able to routinely generate highly contiguous (near telomere-to-telomere) genome assemblies of human and nonhuman species. Complex structural variation and regions of rapid evolutionary turnover are being discovered for the first time. Thus, efficient and informative visualization tools are needed to evaluate and directly observe structural differences between two or more genomes. RESULTS: We developed SVbyEye, an open-source R package to visualize and annotate sequence-to-sequence alignments along with various functionalities to process these alignments. The tool facilitates the characterization of complex structural variants in the context of sequence homology helping resolve the mechanisms underlying their formation. AVAILABILITY AND IMPLEMENTATION: SVbyEye is available on GitHub (https://github.com/daewoooo/SVbyEye) and via Zenodo (https://doi.org/10.5281/zenodo.15303553).

Software↗

Genetic structure and evolution of RAC-GTPases in Arabidopsis thaliana.

Rho GTPases regulate a number of important cellular functions in eukaryotes, such as organization of the cytoskeleton, stress-induced signal transduction, cell death, cell growth, and differentiation. We have conducted an extensive screening, characterization, and analysis of genes belonging to the Ras superfamily of GTPases in land plants (embryophyta) and found that the Rho family is composed mainly of proteins with homology to RAC-like proteins in terrestrial plants. Here we present the genomic and cDNA sequences of the RAC gene family from the plant Arabidopsis thaliana. On the basis of amino acid alignments and genomic structure comparison of the corresponding genes, the 11 encoded AtRAC proteins can be divided into two distinct groups of which one group apparently has evolved only in vascular plants. Our phylogenetic analysis suggests that the plant RAC genes underwent a rapid evolution and diversification prior to the emergence of the embryophyta, creating a group that is distinct from rac/cdc42 genes in other eukaryotes. In embryophyta, RAC genes have later undergone an expansion through numerous large gene duplications. Five of these RAC duplications in Arabidopsis thaliana are reported here. We also present an hypothesis suggesting that the characteristic RAC proteins in higher plants have evolved to compensate the loss of RAS proteins.

Amino Acid Sequence↗