Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Genomic BLAST: custom-defined virtual databases for complete and unfinished genomes.

BLAST (Basic Local Alignment Search Tool) searches against DNA and protein sequence databases have become an indispensable tool for biomedical research. The proliferation of the genome sequencing projects is steadily increasing the fraction of genome-derived sequences in the public databases and their importance as a public resource. We report here the availability of Genomic BLAST, a novel graphical tool for simplifying BLAST searches against complete and unfinished genome sequences. This tool allows the user to compare the query sequence against a virtual database of DNA and/or protein sequences from a selected group of organisms with finished or unfinished genomes. The organisms for such a database can be selected using either a graphic taxonomy-based tree or an alphabetical list of organism-specific sequences. The first option is designed to help explore the evolutionary relationships among organisms within a certain taxonomy group when performing BLAST searches. The use of an alphabetical list allows the user to perform a more elaborate set of selections, assembling any given number of organism-specific databases from unfinished or complete genomes. This tool, available at the NCBI web site http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/genom_table_cgi, currently provides access to over 170 bacterial and archaeal genomes and over 40 eukaryotic genomes.

Amino Acid Sequence↗

Genome instability in BVDV: an examination of the sequence and structural influences on RNA recombination.

The cytopathogenic biotype of the pestivirus, bovine viral diarrhea virus, is frequently a product of nonhomologous recombination in the region of the genome encoding the viral NS2-NS3 proteins. The possibility that sequences or structures in this region contributed to a hotspot for RNA recombination was examined. A PCR-based strategy was used to examine viral genomic RNA isolated from tissue samples of cattle persistently infected with the noncytopathogenic biotype of the virus. Analysis of two different regions of the viral genome revealed that recombination was not restricted to particular sequences. Alignment of the genomic sequences undergoing recombination and examination of the predicted secondary structures of the participating RNAs revealed that the dissociation of partial, newly synthesized negative strand RNAs from the positive strand template occurred at many different sites on the molecule. Similarly, it appeared that the reassociation of the RNA polymerase complex with a second positive strand template was frequently influenced by short regions of homology between the nascent RNA strand and open secondary structures in the template molecule.

Animals↗

Organization of the equine immunoglobulin heavy chain constant region genes; III. Alignment of c mu, c gamma, c epsilon and c alpha genes.

Previous restriction analysis of cloned equine DNA and genomic DNA of equine peripheral blood mononuclear cells had indicated the existence of one c epsilon, one c alpha and up to six c gamma genes in the haploid equine genome. The c epsilon and c alpha genes have been aligned on a 30 kb DNA fragment in the order 5' c epsilon-c alpha 3'. Here we describe the alignment of the equine c mu and c gamma genes by deletion analysis of one IgM, four IgG and two equine light chain expressing heterohybridomas. This analysis establishes the existence of six c gamma genes per haploid genome. The genomic alignment of the cH-genes is 5' c mu/(/) c gamma 1/(/) c gamma 2/(/) c gamma 3/(/) c gamma 4/(/) c gamma 5/(/) c gamma 6/(/) c epsilon-c alpha 3', naming the c gamma genes according to their position relative to c mu. For three of the c gamma genes the corresponding IgG isotypes could be identified as IgGa for c gamma 1, IgG(T) for c gamma 3 and IgGb for c gamma 4.

Animals↗

VISTA family of computational tools for comparative analysis of DNA sequences and whole genomes.

Comparative analysis of DNA sequences is becoming one of the major methods for discovery of functionally important genomic intervals. Presented here the VISTA family of computational tools was built to help researchers in this undertaking. These tools allow the researcher to align DNA sequences, quickly visualize conservation levels between them, identify highly conserved regions, and analyze sequences of interest through one of the following approaches: . Browse precomputed whole-genome alignments of vertebrates and other groups of organisms. . Submit sequences to Genome VISTA to align them to whole genomes. . Submit two or more sequences to mVISTA to align them with each other (a variety of alignment programs with several distinct capabilities are made available).. Submit sequences to Regulatory VISTA (rVISTA) to perform transcription factor binding site predictions based on conservation within sequence alignments.Use stand-alone alignment and visualization programs to run comparative sequence analysis locally All VISTA tools use standard algorithms for visualization and conservation analysis to make comparison of results from different programs more straightforward. The web page http://genome.lbl.gov/vista/ serves as a portal for access to all VISTA tools. Our support group can be reached by email at vista@lbl.gov.

Algorithms↗

The UCSC Genome Browser Database.

The University of California Santa Cruz (UCSC) Genome Browser Database is an up to date source for genome sequence data integrated with a large collection of related annotations. The database is optimized to support fast interactive performance with the web-based UCSC Genome Browser, a tool built on top of the database for rapid visualization and querying of the data at many levels. The annotations for a given genome are displayed in the browser as a series of tracks aligned with the genomic sequence. Sequence data and annotations may also be viewed in a text-based tabular format or downloaded as tab-delimited flat files. The Genome Browser Database, browsing tools and downloadable data files can all be found on the UCSC Genome Bioinformatics website (http://genome.ucsc.edu), which also contains links to documentation and related technical information.

Animals↗

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software↗

An applications-focused review of comparative genomics tools: capabilities, limitations and future challenges.

A team at the Lawrence Livermore National Laboratory (LLNL) was given the task of using computational tools to speed up the development of DNA diagnostics for pathogen detection. This work will be described in another paper in this issue (see pages 133-149). To achieve this goal it was necessary to understand the merits and limitations of the various available comparative genomics tools. A review of some recent tools for multisequence/genome alignment and substring comparison is presented, within the general framework of applicability to a large-scale application. We note that genome alignments are important for many things, only one of which is pathogen detection. Understanding gene function, gene regulation, gene networks, phylogenetic studies and other aspects of evolution all depend on accurate nucleic acid and protein sequence alignment. Selecting appropriate tools can make a large difference in the quality of results obtained and the effort required.

Algorithms↗

Identification of regions in multiple sequence alignments thermodynamically suitable for targeting by consensus oligonucleotides: application to HIV genome.

BACKGROUND: Computer programs for the generation of multiple sequence alignments such as "Clustal W" allow detection of regions that are most conserved among many sequence variants. However, even for regions that are equally conserved, their potential utility as hybridization targets varies. Mismatches in sequence variants are more disruptive in some duplexes than in others. Additionally, the propensity for self-interactions amongst oligonucleotides targeting conserved regions differs and the structure of target regions themselves can also influence hybridization efficiency. There is a need to develop software that will employ thermodynamic selection criteria for finding optimal hybridization targets in related sequences. RESULTS: A new scheme and new software for optimal detection of oligonucleotide hybridization targets common to families of aligned sequences is suggested and applied to aligned sequence variants of the complete HIV-1 genome. The scheme employs sequential filtering procedures with experimentally determined thermodynamic cut off points: 1) creation of a consensus sequence of RNA or DNA from aligned sequence variants with specification of the lengths of fragments to be used as oligonucleotide targets in the analyses; 2) selection of DNA oligonucleotides that have pairing potential, greater than a defined threshold, with all variants of aligned RNA sequences; 3) elimination of DNA oligonucleotides that have self-pairing potentials for intra- and inter-molecular interactions greater than defined thresholds. This scheme has been applied to the HIV-1 genome with experimentally determined thermodynamic cut off points. Theoretically optimal RNA target regions for consensus oligonucleotides were found. They can be further used for improvement of oligo-probe based HIV detection techniques. CONCLUSIONS: A selection scheme with thermodynamic thresholds and software is presented in this study. The package can be used for any purpose where there is a need to design optimal consensus oligonucleotides capable of interacting efficiently with hybridization targets common to families of aligned RNA or DNA sequences. Our thermodynamic approach can be helpful in designing consensus oligonucleotides with consistently high affinity to target variants in evolutionary related genes or genomes.

Base Sequence↗

Cytogenetic alignment of the bovine chromosome 13 genome map by fluorescence in-situ hybridization of human chromosome 10 and 20 comparative markers.

A bovine bacterial artificial chromosome (BAC) library was screened for the presence of six genes (IL2RA, VIM, THBD, PLC-II, CSNK2A1 and TOP1) previously assigned to human chromosomes 10 or 20 (HSA10 or HSA20). Four of the genes were found represented in the bovine BAC library by at least one clone. The identified BAC clones were used as probes in single-color fluorescence in-situ hybridization (FISH) to determine the chromosomal band location of each gene. As predicted by the human/bovine comparative map and comparative chromosome painting analysis, the four genes mapped to bovine chromosome 13 (BTA13). Dual-color FISH was then used to integrate these four type I markers into the existing BTA13 genome map. These FISH results anchor the BTA13 genome map from bands 14-23, and confirm the presence of a conserved HSA10 homologous synteny group on BTA13 centromeric to a HSA20 homologous segment.

Animals↗

Identification of a gadd45beta 3' enhancer that mediates SMAD3- and SMAD4-dependent transcriptional induction by transforming growth factor beta.

GADD45beta regulates cell growth, differentiation, and cell death following cellular exposure to diverse stimuli, including DNA damage and transforming growth factor-beta (TGFbeta). We examined how cells transduce the TGFbeta signal from the cell surface to the gadd45beta genomic locus and describe how GADD45beta contributes to TGFbeta biology. Following an alignment of gadd45beta genomic sequences from multiple organisms, we discovered a novel TGFbeta-responsive enhancer encompassing the third intron of the gadd45beta gene. Using three different experimental approaches, we found that SMAD3 and SMAD4, but not SMAD2, mediate transcription from this enhancer. Three lines of evidence support our conclusions. First, overexpression of SMAD3 and SMAD4 activated the transcriptional activity from this enhancer. Second, silencing of SMAD protein levels using short interfering RNAs revealed that TGFbeta-induced activation of the endogenous gadd45beta gene required SMAD3 and SMAD4 but not SMAD2. In contrast, we found that the regulation of plasminogen activator inhibitor type I depended upon all three SMAD proteins. Last, SMAD3 and SMAD4 reconstitution in SMAD-deficient cancer cells restored TGFbeta induction of gadd45beta. Finally, we assessed the function of GADD45beta within the TGFbeta response and found that GADD45beta-deficient cells arrested in G2 following TGFbeta treatment. These data support a role for SMAD3 and SMAD4 in activating gadd45beta through its third intron to facilitate G2 progression following TGFbeta treatment.

Animals↗

The bioinformatics resource for oral pathogens.

Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.

Bacteria↗

MAVID multiple alignment server.

MAVID is a multiple alignment program suitable for many large genomic regions. The MAVID web server allows biomedical researchers to quickly obtain multiple alignments for genomic sequences and to subsequently analyse the alignments for conserved regions. MAVID has been successfully used for the alignment of closely related species such as primates and also for the alignment of more distant organisms such as human and fugu. The server is fast, capable of aligning hundreds of kilobases in less than a minute. The multiple alignment is used to build a phylogenetic tree for the sequences, which is subsequently used as a basis for identifying conserved regions in the alignment. The server can be accessed at http://baboon.math.berkeley.edu/mavid/.

Animals↗

Sequence and organization of the Spodoptera exigua multicapsid nucleopolyhedrovirus genome.

The nucleotide sequence of the DNA genome of Spodoptera exigua multicapsid nucleopolyhedrovirus (SeMNPV), a group II NPV, was determined and analysed. The genome contains 135611 bp and has a G+C content of 44 mol%. Computer-assisted analysis revealed 139 ORFs of 150 nucleotides or larger; 103 have homologues in Autographa californica MNPV (AcMNPV) and a further 16 have homologues in other baculoviruses. Twenty ORFs are unique to SeMNPV. Major differences in SeMNPV gene content and arrangement were found compared with the group I NPVs AcMNPV, Bombyx mori (Bm) NPV and Orgyia pseudotsugata (Op) MNPV and the group II NPV Lymantria dispar (Ld) MNPV. Eighty-five ORFs were conserved among all five baculoviruses and are considered as candidate core baculovirus genes. Two putative p26 and odv-e66 homologues were identified in SeMNPV, each of which appeared to have been acquired independently and not by gene duplication. The SeMNPV genome lacks homologues of the major budded virus glycoprotein gene gp64, the immediate-early transactivator ie-2 and bro (baculovirus repeat ORF) genes that are found in AcMNPV, BmNPV, OpMNPV and LdMNPV. Gene parity analysis of baculovirus genomes suggests that SeMNPV and LdMNPV have a recent common ancestor and that they are more distantly related to the group I baculoviruses AcMNPV, BmNPV and OpMNPV. The orientation of the SeMNPV genome is reversed compared with the genomes of AcMNPV, BmNPV, OpMNPV and LdMNPV. However, the gene order in the 'central' part of baculovirus genomes is highly conserved and appears to be a key feature in the alignment of baculovirus genomes.

Animals↗

Genome of an avian papillomavirus.

A papillomavirus which we designate FPV was isolated from chaffinches (Fringilla coelebs). A physical map of the FPV genome was constructed, and selected regions of this genome were studied by nucleotide sequence analysis. The results make it possible to align the FPV genome with the genome of bovine papillomavirus type 1 and to show, moreover, that avian and mammalian papillomaviruses have a similar genome organization.

Amino Acid Sequence↗

Tumour-infiltrating lymphocytes differ across MammaPrint® classifications in breast cancer.

MammaPrint&#xae; refines risk stratification in early oestrogen receptor-positive, HER2-negative breast cancer, evolving from a binary to a four-tier classification (UltraLow-Risk, Low-Risk, High-Risk 1, and High-Risk 2). The relationship between routine histopathological features, immune infiltration, and genomic risk within this framework remains incompletely characterized in luminal disease. We retrospectively analysed 492 luminal breast carcinomas with available MammaPrint&#xae; results. Clinicopathological variables (including histological subtype, grade, Ki-67, hormone receptor expression, lymphovascular invasion (LVI), and HER2-low status) were recorded. Stromal tumour-infiltrating lymphocytes (TILs) were quantified according to the criteria of Salgado et al. and spatially categorized as immune-deserted, stromal-restricted, immune-excluded, or inflamed patterns. CD4 and CD8 infiltration was assessed by immunohistochemistry on tissue microarrays. Associations with binary and four-tier MammaPrint&#xae; categories were examined using multivariable models. High-Risk tumours (41%) were enriched for increased grade, Ki-67, and LVI and lower PR expression. In High-Risk tumours, inflamed spatial patterns were more frequent and median TIL levels were significantly higher compared with Low-Risk tumours (15% vs 5%, P < 0.001). CD4 and CD8 infiltration increased with genomic risk, and CD4 retained a modest but statistically significant association after adjustment for conventional pathological variables. In the four-tier model, UltraLow-Risk/Low-Risk tumours showed minimal TILs and were enriched for invasive lobular carcinoma, whereas High-Risk 1/High-Risk 2 tumours displayed progressively higher proliferative and immune features. No association was observed between HER2-low status and genomic risk. Immune infiltration parallels proliferative and genomic risk gradients in luminal breast cancer. These findings indicate that immune descriptors align with the genomic risk continuum, although their independent prognostic contribution beyond established genomic assays requires further evaluation.

Humans↗

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software↗

Experimental analysis of the annotation of promoters in the public database.

The ability to identify and examine promoter elements is important to researchers who wish to understand how gene expression is regulated in normal and pathological states. Unfortunately, the number of human promoters that have been directly experimentally defined is small. In order to determine if promoter sequences can be identified by simply aligning mRNA and genomic sequences, we have used a reporter gene assay to assess the promoter activity of the immediate 5' region flanking 38 mRNAs mapping to chromosome 21. For comparison, we have measured the activities of 19 sequences not thought to be promoters and 39 sequences taken from the Eukaryotic Promoter Database. Our results suggest that alignment of reference mRNAs to genomic sequence allows promoters to be identified for at least 75% of genes. These data provide the first empirical evidence that the current state of annotation of the genome is sufficient to allow molecular geneticists to correctly identify promoter sequences for most genes for which reference mRNA and genomic sequences are available.

Cell Line↗

Beyond blacklists: a critical assessment of exclusion set generation strategies and alternative approaches.

MOTIVATION: Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Alternatively, "sponge" or decoy sequences have been proposed to reduce alignment artifacts. RESULTS: We examined the widely used Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to sensitivity to input data, aligner choice, and read length. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. We additionally investigated the effect of the T2T-CHM13 genome assembly on improving biological signals. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets, and recommend the use of the T2T-CHM13 assembly or, for the hg38 genome assembly, "sponge" sequences as an alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Software↗