Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

High-resolution alignment of a 1-megabase-long genome region of three strains of Rhodobacter capsulatus.

A detailed restriction map of the genome of Rhodobacter capsulatus SB1003 was constructed recently by using an ordered set of overlapping cosmids. Pulsed-field gel electrophoresis-generated restriction patterns of the chromosomes of 14 other R. capsulatus strains were compared. Two of them, St. Louis and 2.3.1, were chosen for high-resolution alignment of their genomes with that of SB1003. A 1-Mb segment of the R. capsulatus SB1003 cosmid set was used as a source of ordered probes to group cosmids from the other strains. Selected cosmids were linked into one 800-kb contig and two smaller contigs of 100 kb each. EcoRV and BamHI restriction maps of the newly ordered cosmids were constructed by using lambda terminase. Long-range gene order in the new strains was mainly conserved for the regions studied. However, one large genome rearrangement inverted a 470-kb DNA fragment of the St. Louis strain between the rrnA and rrnB operons. A 50-kb deletion covering three SB1003 probes was found in strain 2.3.1 near rrnB. Conservation of about 50% of the positions of restriction sites in all these strains and nearly 80% for the pair 2.3.1- St. Louis made it possible to produce high-resolution alignment of the contiguous 800-kb genome segment. Ten deletions of 2 to 27 kb, one 30-kb inversion, and three translocations were found in this region. Strong clustering of the positions of polymorphic restriction sites was observed. For a 50-kb size interval, two patterns of the distribution of restriction sites were found, one with about 90% and the other with 5 to 30% conservation of sites. This structure may be explained by independent acquisition of these divergent regions from other Rhodobacter strains.

Chromosome Inversion↗

[Gene cloning and sequencing of chicken anemia virus(CAV) isolated from Harbin].

A Chicken anemia virus has been isolated from a chicken flock in Harbin of China. The genome of the ivrus was cloned through polymerase chain reaction(PCR) and sequence of the genome was analyzed. The cycle genome is made of 2298 base pairs including three overlapping open reading frames(vp1, vp2, vp3) and a regulative region. Comparing sequence of the genome through BLAST in GenBank, this sequence exhibits 96.9% identity with other genome of CA Vs and least. Multiple alignment of this genome of this virus, 26p4, strain isolated in Germany, strain isolated in Malaysia and Cux-1 found that this sequence exhibits 98.2% (42/2298), 98.2% (42/2298), 96.9% (72/2298) and 97.5% (60/2319) identify with them, respectively. A new CAV strain was isolated and it has better identify with CAV isolated in Europe countries than is Asia country Malaysia. Multiple alignment of VP1, VP2, VP3 of 26p4, strain isolated in Germany, strain isolated in Malaysia, Cux-1 and strain isolated in Harbin of China found the VP2 the most conservative.

Amino Acid Sequence↗

In vitro transformation and molecular characterization of Colobus monkey venereal papillomavirus DNA.

The DNA of a monkey papillomavirus (CgPV-1), originally isolated from a penile lesion on a Colobus monkey was cloned into the EcoRI site of the pUC18 vector and characterized. Using a variety of restriction enzymes a physical map of the DNA was constructed. Cross-hybridization with a variety of animal and human papillomaviruses under high (Tm-22 degrees C) and low (Tm-40 degrees C) stringency conditions indicated various degrees of homology. CgPV-1 showed higher homology with HPVs than it did with any other animal papillomaviruses tested. DNA similarities with the human papillomaviruses HPV-16 and HPV-18 that are frequently associated with cervical cancer, were manifested by extensive cross-hybridization under stringent conditions. Functional alignment of the genomic map of CgPV-1 with that of HPV-16 was carried out by determination of homology between specific restriction fragments of the two viral genomes in cross-hybridization analyses. This alignment was refined by sequencing two regions of approximately 200 bp of the CgPV-1 DNA, and aligning them by computer with their homologous HPV-16 counterparts. CgPV-1 DNA in its pUC18 vector, transformed NIH 3T3 cells with roughly the same efficiency as BPV-1, as determined by the number of transformed foci generated per ug of DNA. The data presented indicate that the state of the CgPV-1 viral DNA in these transformed cells is integrated and partially deleted, not unlike the genomes of HPV-16 and HPV-18 characterized in cell lines derived from cervical cancers.

Animals↗

Dcode.org anthology of comparative genomic tools.

Comparative genomics provides the means to demarcate functional regions in anonymous DNA sequences. The successful application of this method to identifying novel genes is currently shifting to deciphering the non-coding encryption of gene regulation across genomes. To facilitate the practical application of comparative sequence analysis to genetics and genomics, we have developed several analytical and visualization tools for the analysis of arbitrary sequences and whole genomes. These tools include two alignment tools, zPicture and Mulan; a phylogenetic shadowing tool, eShadow for identifying lineage- and species-specific functional elements; two evolutionary conserved transcription factor analysis tools, rVista and multiTF; a tool for extracting cis-regulatory modules governing the expression of co-regulated genes, Creme 2.0; and a dynamic portal to multiple vertebrate and invertebrate genome alignments, the ECR Browser. Here, we briefly describe each one of these tools and provide specific examples on their practical applications. All the tools are publicly available at the http://www.dcode.org/ website.

Base Sequence↗

Efficient discovery of single-nucleotide polymorphisms in coding regions of human genes.

Single nucleotide polymorphisms in protein coding regions (cSNPs) are of great interest for their effects on phenotype and potential for mapping disease genes. We have identified 5,400 novel exonic SNPs from alignments of public EST data to the draft human genome sequence, and approximately 12,000 more novel exonic SNPs from EST cluster alignments. We found 82% of the genomic-aligned SNPs and 63% of the EST-only SNPs to be detectably polymorphic in 20 Finnish DNA samples. 37% of the SNPs mapped to known protein coding regions, yielding 6,500 distinct, novel cSNPs from the two datasets. These data reveal selection against mutations that alter protein structure, and distinct classes of genes under strongly positive vs. negative pressure from natural selection for amino acid replacement (detected by K(A)/K(S)ratio). We have searched these cSNPs for compatibility with the amino acid profile at each site and structural impact on protein core stability.

Chromosome Mapping↗

Origins of primate chromosomes - as delineated by Zoo-FISH and alignments of human and mouse draft genome sequences.

This review examines recent advances in comparative eutherian cytogenetics, including Zoo-FISH data from 30 non-primate species. These data provide insights into the nature of karyotype evolution and enable the confident reconstruction of ancestral primate and boreo-eutherian karyotypes with diploid chromosome numbers of 48 and 46 chromosomes, respectively. Nine human autosomes (1, 5, 6, 9, 11, 13, 17, 18, and 20) represent the syntenies of ancestral boreo-eutherian chromosomes and have been conserved for about 95 million years. The average rate of chromosomal exchanges in eutherian evolution is estimated to about 1.9 rearrangements per 10 million years (involving 3.4 chromosome breaks). The integrated analysis of Zoo-FISH data and alignments of human and mouse draft genome sequences allow the identification of breakpoints involved in primate evolution. Thus, the boundaries of ancestral eutherian conserved segments can be delineated precisely. The mapping of rearrangements onto the phylogenetic tree visualizes landmark chromosome rearrangements, which might have been involved in cladogenesis in eutherian evolution.

Animals↗

Identification of probable genomic packaging signal sequence from SARS-CoV genome by bioinformatics analysis.

AIM: To predict the probable genomic packaging signal of SARS-CoV by bioinformatics analysis. The derived packaging signal may be used to design antisense RNA and RNA interfere (RNAi) drugs treating SARS. METHODS: Based on the studies about the genomic packaging signals of MHV and BCoV, especially the information about primary and secondary structures, the putative genomic packaging signal of SARS-CoV were analyzed by using bioinformatic tools. Multi-alignment for the genomic sequences was performed among SARS-CoV, MHV, BCoV, PEDV and HCoV 229E. Secondary structures of RNA sequences were also predicted for the identification of the possible genomic packaging signals. Meanwhile, the N and M proteins of all five viruses were analyzed to study the evolutionary relationship with genomic packaging signals. RESULTS: The putative genomic packaging signal of SARS-CoV locates at the 3' end of ORF1b near that of MHV and BCoV, where is the most variable region of this gene. The RNA secondary structure of SARS-CoV genomic packaging signal is very similar to that of MHV and BCoV. The same result was also obtained in studying the genomic packaging signals of PEDV and HCoV 229E. Further more, the genomic sequence multi-alignment indicated that the locations of packaging signals of SARS-CoV, PEDV, and HCoV overlaped each other. It seems that the mutation rate of packaging signal sequences is much higher than the N protein, while only subtle variations for the M protein. CONCLUSIONS: The probable genomic packaging signal of SARS-CoV is analogous to that of MHV and BCoV, with the corresponding secondary RNA structure locating at the similar region of ORF1b. The positions where genomic packaging signals exist have suffered rounds of mutations, which may influence the primary structures of the N and M proteins consequently.

Amino Acid Sequence↗

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA↗

Analysis of primate genomic variation reveals a repeat-driven expansion of the human genome.

We performed a detailed analysis of both single-nucleotide and large insertion/deletion events based on large-scale comparison of 10.6 Mb of genomic sequence from lemur, baboon, and chimpanzee to human. Using a human genomic reference, optimal global alignments were constructed from large (>50-kb) genomic sequence clones. These alignments were examined for the pattern, frequency, and nature of mutational events. Whereas rates of single-nucleotide substitution remain relatively constant (1-2 x 10(-9) substitutions/site/year), rates of retrotransposition vary radically among different primate lineages. These differences have lead to a 15%-20% expansion of human genome size over the last 50 million years of primate evolution, 90% of it due to new retroposon insertions. Orthologous comparisons with the chimpanzee suggest that the human genome continues to significantly expand due to shifts in retrotransposition activity. Assuming that the primate genome sequence we have sampled is representative, we estimate that human euchromatin has expanded 30 Mb and 550 Mb compared to the primate genomes of chimpanzee and lemur, respectively.

Animals↗

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗

AnimalQTLdb: a livestock QTL database tool set for positional QTL information mining and beyond.

The Animal Quantitative Trait Loci (QTL) database (AnimalQTLdb) is designed to house all publicly available QTL data on livestock animal species from which researchers can easily locate and compare QTL within species. The database tools are also added to link the QTL data to other types of genomic information, such as radiation hybrid (RH) maps, finger printed contig (FPC) physical maps, linkage maps and comparative maps to the human genome, etc. Currently, this database contains data on 1287 pig, 630 cattle and 657 chicken QTL, which are dynamically linked to respective RH, FPC and human comparative maps. We plan to apply the tool to other animal species, and add more structural genome information for alignment, in an attempt to aid comparative structural genome studies (http://www.animalgenome.org/QTLdb/).

Animals↗

Gene verification and discovery by Walking Tree Method.

The Walking Tree Method [3, 4, 5, 18] is an approximate string alignment method that can handle insertions, deletions, substitutions, translocations, and more than one level of inversions all together. Moreover, it tends to highlight gene locations, and helps discover unknown genes. Its recent improvements in runtime and space use extends its capability in exploring large strings. We will briefly describe the Walking Tree Method with its recent improvements [18], and demonstrate its speed and ability to align real complete genomes such as Borrelia burgdorferi (910724 base pairs of its single chromosome) and Chlamydia trachomatis (1042519 base pairs) in reasonable time, and to locate and verify genes.

Animals↗

DIALIGN P: fast pair-wise and multiple sequence alignment using parallel processors.

BACKGROUND: Parallel computing is frequently used to speed up computationally expensive tasks in Bioinformatics. RESULTS: Herein, a parallel version of the multi-alignment program DIALIGN is introduced. We propose two ways of dividing the program into independent sub-routines that can be run on different processors: (a) pair-wise sequence alignments that are used as a first step to multiple alignment account for most of the CPU time in DIALIGN. Since alignments of different sequence pairs are completely independent of each other, they can be distributed to multiple processors without any effect on the resulting output alignments. (b) For alignments of large genomic sequences, we use a heuristics by splitting up sequences into sub-sequences based on a previously introduced anchored alignment procedure. For our test sequences, this combined approach reduces the program running time of DIALIGN by up to 97%. CONCLUSIONS: By distributing sub-routines to multiple processors, the running time of DIALIGN can be crucially improved. With these improvements, it is possible to apply the program in large-scale genomics and proteomics projects that were previously beyond its scope.

Computational Biology↗

Gene identification in novel eukaryotic genomes by self-training algorithm.

Finding new protein-coding genes is one of the most important goals of eukaryotic genome sequencing projects. However, genomic organization of novel eukaryotic genomes is diverse and ab initio gene finding tools tuned up for previously studied species are rarely suitable for efficacious gene hunting in DNA sequences of a new genome. Gene identification methods based on cDNA and expressed sequence tag (EST) mapping to genomic DNA or those using alignments to closely related genomes rely either on existence of abundant cDNA and EST data and/or availability on reference genomes. Conventional statistical ab initio methods require large training sets of validated genes for estimating gene model parameters. In practice, neither one of these types of data may be available in sufficient amount until rather late stages of the novel genome sequencing. Nevertheless, we have shown that gene finding in eukaryotic genomes could be carried out in parallel with statistical models estimation directly from yet anonymous genomic DNA. The suggested method of parallelization of gene prediction with the model parameters estimation follows the path of the iterative Viterbi training. Rounds of genomic sequence labeling into coding and non-coding regions are followed by the rounds of model parameters estimation. Several dynamically changing restrictions on the possible range of model parameters are added to filter out fluctuations in the initial steps of the algorithm that could redirect the iteration process away from the biologically relevant point in parameter space. Tests on well-studied eukaryotic genomes have shown that the new method performs comparably or better than conventional methods where the supervised model training precedes the gene prediction step. Several novel genomes have been analyzed and biologically interesting findings are discussed. Thus, a self-training algorithm that had been assumed feasible only for prokaryotic genomes has now been developed for ab initio eukaryotic gene identification.

Algorithms↗

gaftools: a toolkit for analyzing and manipulating pangenome alignments.

MOTIVATION: Linear reference genomes are ubiquitously used in genomics research, despite known biases associated with their use. In recent years, there has been a shift towards graph-based reference genomes to address some of these biases, which has required development of new algorithms and file formats. This has created a necessity for new tools capable of utilizing these formats and performing operations similar to those carried out by traditional methods. RESULTS: In this paper we present "gaftools," a multi-purpose tool that introduces several utilities for processing graph alignments in GAF format. gaftools enables users to index and sort alignments, with graph ordering serving as a necessary step for the sorting process. Additionally, it allows users to view subsets of alignments and perform realignment using the wavefront alignment algorithm, among other features. Many of these functionalities are inspired by SAMtools, which provides similar operations for linear genomes, while gaftools adapts and extends them for pangenomes. AVAILABILITY: gaftools is available under MIT license at https://github.com/marschall-lab/gaftools.

Software↗

Gene sequences useful for predicting relatedness of whole genomes in bacteria.

Thirty-two protein-encoding genes that are distributed widely among bacterial genomes were tested for the potential usefulness of their DNA sequences in assigning bacterial strains to species. From publicly available data, it was possible to make 49 pairwise comparisons of whole bacterial genomes that were related at the genus or subgenus level. DNA sequence identity scores for eight of the genes correlated strongly with overall sequence identity scores for the genome pairs. Even single-gene alignments could predict overall genome relatedness with a high degree of precision and accuracy. Predictions could be refined further by including two or three genes in the analysis. The proposal that sequence analysis of a small set of protein-encoding genes could reliably assign novel strains or isolates to bacterial species is strongly supported.

Bacteria↗

Sequence changes in six variants of rice tungro bacilliform virus and their phylogenetic relationships.

The DNA of three biological variants, G1, Ic and G2, which originated from the same greenhouse isolate of rice tungro bacilliform virus (RTBV) at the International Rice Research Institute (IRRI), was cloned and sequenced. Comparison of the sequences revealed small differences in genome sizes. The variants were between 95 and 99% identical at the nucleotide and amino acid levels. Alignment of the three genome sequences with those of three published RTBV sequences (Phi-1, Phi-2 and Phi-3) revealed numerous nucleotide substitutions and some insertions and deletions. The published RTBV sequences originated from the same greenhouse isolate at IRRI 20, 11 and 9 years ago. All open reading frames (ORFs) and known functional domains were conserved across the six variants. The cysteine-rich region of ORF3 showed the greatest variation. When the six DNA sequences from IRRI were compared with that of an isolate from Malaysia (Serdang), similar changes were observed in the cysteine-rich region in addition to other nucleotide substitutions and deletions across the genome. The aligned nucleotide sequences of the IRRI variants and Serdang were used to analyse phylogenetic relationships by the bootstrapped parsimony, distance and maximum-likelihood methods. The isolates clustered in three groups: Serdang alone; Ic and G1; and Phi-1, Phi-2, Phi-3 and G2. The distribution of phylogenetically informative residues in the IRRI sequences shared with the Serdang sequence and the differing tree topologies for segments of the genome suggested that recombination, as well as substitutions and insertions or deletions, has played a role in the evolution of RTBV variants. The significance and implications of these evolutionary forces are discussed in comparison with badnaviruses and caulimoviruses.

Amino Acid Sequence↗