Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

A comparative genomics strategy for targeted discovery of single-nucleotide polymorphisms and conserved-noncoding sequences in orphan crops.

Completed genome sequences provide templates for the design of genome analysis tools in orphan species lacking sequence information. To demonstrate this principle, we designed 384 PCR primer pairs to conserved exonic regions flanking introns, using Sorghum/Pennisetum expressed sequence tag alignments to the Oryza genome. Conserved-intron scanning primers (CISPs) amplified single-copy loci at 37% to 80% success rates in taxa that sample much of the approximately 50-million years of Poaceae divergence. While the conserved nature of exons fostered cross-taxon amplification, the lesser evolutionary constraints on introns enhanced single-nucleotide polymorphism detection. For example, in eight rice (Oryza sativa) genotypes, polymorphism averaged 12.1 per kb in introns but only 3.6 per kb in exons. Curiously, among 124 CISPs evaluated across Oryza, Sorghum, Pennisetum, Cynodon, Eragrostis, Zea, Triticum, and Hordeum, 23 (18.5%) seemed to be subject to rigid intron size constraints that were independent of per-nucleotide DNA sequence variation. Furthermore, we identified 487 conserved-noncoding sequence motifs in 129 CISP loci. A large CISP set (6,062 primer pairs, amplifying introns from 1,676 genes) designed using an automated pipeline showed generally higher abundance in recombinogenic than in nonrecombinogenic regions of the rice genome, thus providing relatively even distribution along genetic maps. CISPs are an effective means to explore poorly characterized genomes for both DNA polymorphism and noncoding sequence conservation on a genome-wide or candidate gene basis, and also provide anchor points for comparative genomics across a diverse range of species.

Base Sequence↗

GenoMiner: a tool for genome-wide search of coding and non-coding conserved sequence tags.

GenoMiner is a software tool that searches for regions of similarity between user-submitted genome or transcript sequences and user-specified whole genome assemblies. The program then identifies conserved sequence tags (CSTs) in these homologous regions and provides a prediction of their coding or non-coding nature. The analysis is carried out through three steps: (1) definition of sequence regions homologous to the query sequence in the selected target genomes by a fast BLAT alignment; (2) identification of CSTs by a more sensitive BLAST-like alignment between the query and the homologous regions in the target genomes and (3) assessment of the coding or non-coding nature of detected CSTs through the computation of a suitable coding potential score. GenoMiner allows the user to search the query sequence against a number of vertebrate genome assemblies in a single run providing a user-friendly graphical output.

Algorithms↗

Occurrence of Helicobacter pylori DNA in the coastal environment of southern Italy (Straits of Messina).

AIMS: The occurrence of Helicobacter pylori in the coastal zone of the Straits of Messina (Italy) as free-living and associated with plankton was studied. METHODS AND RESULTS: Monthly sampling of seawater and plankton was carried out from April 2002 to March, 2003. All environmental samples analysed by cultural method, did not show the presence of H. pylori. The DNA extracted from all environmental samples was tested by PCR by using primers for H. pylori 16S rRNA, ureA and cagA. 16S rRNA PCR yielded amplified products of 522-bp in 15 of 36 (41.7%) of the environmental samples. By using the ureA primers to amplify the urea signal sequences, the predicted PCR products of 491-bp were obtained from eight (22.2%) of 36 environmental samples. PCR with cagA primers yielded amplified products of 349-bp in DNA extracted of seven of 36 (19.4%) of the environmental samples. When 16S rRNA, ureA and cagA amplified gene sequences were aligned with H. pylori 26695 and J99 genome sequences, we obtained a percentage of alignment over 90%. CONCLUSIONS: The detection of H. pylori genes in marine samples allows us to consider the marine environment a possible reservoir for this pathogenic bacterium. SIGNIFICANCE AND IMPACT OF THE STUDY: The direct detection of H. pylori genes may be relevant in order to consider the marine environment as significant reservoir for this bacterium.

Antigens, Bacterial↗

An information-based sequence distance and its application to whole mitochondrial genome phylogeny.

MOTIVATION: Traditional sequence distances require an alignment and therefore are not directly applicable to the problem of whole genome phylogeny where events such as rearrangements make full length alignments impossible. We present a sequence distance that works on unaligned sequences using the information theoretical concept of Kolmogorov complexity and a program to estimate this distance. RESULTS: We establish the mathematical foundations of our distance and illustrate its use by constructing a phylogeny of the Eutherian orders using complete unaligned mitochondrial genomes. This phylogeny is consistent with the commonly accepted one for the Eutherians. A second, larger mammalian dataset is also analyzed, yielding a phylogeny generally consistent with the commonly accepted one for the mammals. AVAILABILITY: The program to estimate our sequence distance, is available at http://www.cs.cityu.edu.hk/~cssamk/gencomp/GenCompress1.htm. The distance matrices used to generate our phylogenies are available at http://www.math.uwaterloo.ca/~mli/distance.html.

Animals↗

Molecular cloning and genomic organization of a second probable allatostatin receptor from Drosophila melanogaster.

We (C. Lenz et al. (2000) Biochem. Biophys. Res. Commun. 269, 91-96) and others (N. Birgül et al. (1999) EMBO J. 18, 5892-5900) have recently cloned a Drosophila receptor that was structurally related to the mammalian galanin receptors, but turned out to be a receptor for a Drosophila peptide belonging to the insect allatostatin neuropeptide family. In the present paper, we screened the Berkeley "Drosophila Genome Project" database with "electronic probes" corresponding to the conserved regions of the four rat (delta, kappa, mu, nociceptin/orphanin FQ) opioid receptors. This yielded alignment with a Drosophila genomic database clone that contained a DNA sequence coding for a protein having, again, structural similarities with the rat galanin receptors. Using PCR with primers coding for the presumed exons of this second Drosophila receptor gene, 5'- and 3'-RACE, and Drosophila cDNA as template, we subsequently cloned the cDNA of this receptor. The receptor cDNA codes for a protein that is strongly related to the first Drosophila receptor (60% amino acid sequence identity in the transmembrane region; 47% identity in the overall sequence) and that is, therefore, most likely to be a second Drosophila allatostatin receptor (named DAR-2). The DAR-2 gene has three introns and four exons. Two of these introns coincide with two introns in the first Drosophila receptor (DAR-1) gene, and have the same intron phasing, showing that the two receptor genes are clearly evolutionarily related. The DAR-2 gene is located at the right arm of the third chromosome, position 98 D-E. This is the first report on the existence of two different allatostatin receptors in an animal.

Amino Acid Sequence↗

Diagnosing missed cases of spinal muscular atrophy in genome, exome, and panel sequencing data sets.

PURPOSE: We set out to develop a publicly available tool that could accurately diagnose spinal muscular atrophy (SMA) in exome, genome, or panel sequencing data sets aligned to a GRCh37, GRCh38, or T2T reference genome. METHODS: The SMA Finder algorithm detects the most common genetic causes of SMA by evaluating reads that overlap the c.840 position of the SMN1 and SMN2 paralogs. It uses these reads to determine whether an individual most likely has 0 functional copies of SMN1. RESULTS: We developed SMA Finder and evaluated it on 16,626 exomes and 3911 genomes from the Broad Institute Center for Mendelian Genomics, 1157 exomes and 8762 panel samples from Tartu University Hospital, and 198,868 exomes and 198,868 genomes from the UK Biobank. SMA Finder's false-positive rate was below 1 in 200,000 samples, its positive predictive value was greater than 96%, and its true-positive rate was 29 out of 29. Most of these SMA diagnoses had initially been clinically misdiagnosed as limb-girdle muscular dystrophy. CONCLUSION: Our extensive evaluation of SMA Finder on exome, genome, and panel sequencing samples found it to have nearly 100% accuracy and demonstrated its ability to reduce diagnostic delays, particularly in individuals with milder subtypes of SMA. Given this accuracy, the common misdiagnoses identified here, the widespread availability of clinical confirmatory testing for SMA, and the existence of treatment options, we propose that it is time to add SMN1 to the American College of Medical Genetics list of genes with reportable secondary findings after genome and exome sequencing.

Humans↗

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic↗

Genomic mapping of a calicivirus VPg.

We identified a primate calicivirus (Pan-1) VPg in Pan-1-infected cells. The Pan-1 VPg was associated with both genomic and subgenomic RNAs. RNase digestion of Pan-1 RNA yielded a residual protein of 16 kDa. The N-terminal sequence of Pan-1 VPg was determined by direct amino acid sequencing and mapped to a region of the genome equivalent to picornavirus VPgs. Alignment of this protein sequence with similar regions of other calicivirus genomes allowed identification of conserved amino acid motifs and potential boundaries of the calicivirus VPg genes. Proteinase K treatment abolished the infectivity of Pan-1 RNA, suggesting that Pan-1 VPg is required for RNA infectivity.

Amino Acid Sequence↗

Comparative sequence analysis of the eastern equine encephalitis virus pathogenic strains FL91-4679 and GA97 to other North American strains.

Eastern equine encephalitis (EEE) virus is a significant public health concern due to the high mortality rates observed in infected humans, equines and game birds. The EEE genomic sequences available prior to this report are based on laboratory strains with unknown passage histories that may contain an array of cell culture adaptations. Here we report the complete genomic sequences of two recently isolated EEE pathogenic strains with low passage histories. FL91-4697 was isolated in Florida from Aedes albopictus mosquitoes and GA97 was derived from brain tissue of a human fatality that occurred in 1997. Sequence alignment of these new genomes with the documented EEE's permitted us to generate a North American consensus sequence and identify regions of significant diversity. Sequence analysis of the FL91-4679 genome was essential to the production of an EEE infectious clone that is being used to create live attenuated vaccine candidates.

Aedes↗

transAlign: using amino acids to facilitate the multiple alignment of protein-coding DNA sequences.

BACKGROUND: Alignments of homologous DNA sequences are crucial for comparative genomics and phylogenetic analysis. However, multiple alignment represents a computationally difficult problem. For protein-coding DNA sequences, it is more advantageous in terms of both speed and accuracy to align the amino-acid sequences specified by the DNA sequences rather than the DNA sequences themselves. Many implementations making use of this concept of "translated alignments" are incomplete in the sense that they require the user to manually translate the DNA sequences and to perform the amino-acid alignment. As such, they are not well suited to large-scale automated alignments of large and/or numerous DNA data sets. RESULTS: transAlign is an open-source Perl script that aligns protein-coding DNA sequences via their amino-acid translations to take advantage of the superior multiple-alignment capabilities and speed of an amino-acid alignment. It operates by translating each DNA sequence into its corresponding amino-acid sequence, passing the entire matrix to ClustalW for alignment, and then back-translating the resulting amino-acid alignment to derive the aligned DNA sequences. In the translation step, transAlign determines the optimal orientation and reading frame for each DNA sequence according to the desired genetic code. It also checks for apparent frame shifts in the DNA sequences and can handle frame-shifted sequences in one of three ways (delete, align as amino acids regardless, or profile align as DNA). As a set of comparative benchmarks derived from six protein-coding genes for mammals shows, the strategy implemented in transAlign always improves the speed and usually the apparent accuracy of the alignment of protein-coding DNA sequences. CONCLUSION: transAlign represents one of few full and cross-platform implementations of the concept of translated alignments. Both the advantages accruing from performing a translated alignment and the suite of user-definable options available in the program mean that transAlign is ideally suited for large-scale automated alignments of very large and/or very numerous protein-coding DNA data sets. However, the good performance offered by the program also translates to the alignment of any set of protein-coding sequences. transAlign, including the source code, is freely available at http://www.tierzucht.tum.de/Bininda-Emonds/ (under "Programs").

Algorithms↗

PHIRE, a deterministic approach to reveal regulatory elements in bacteriophage genomes.

MOTIVATION: In silico genome analysis of bacteriophage genomes focuses mainly on gene discovery and functional assignment. The search for regulatory elements contained within these genome sequences is often based on prior knowledge of other genomic elements or on learning algorithms of experimentally determined data, potentially leading to a biased prediction output. The PHage In silico Regulatory Elements (PHIRE) program is a standalone program in Visual Basic. It performs an algorithmic string-based search on bacteriophage genome sequences to uncover and extract subsequence alignments hinting at regulatory elements contained within these genomes, in a deterministic manner without any prior experimental or predictive knowledge. RESULTS: The PHIRE program was tested on known phage genomes with experimentally verified regulatory elements. PHIRE was able to extract phage regulatory sequences correctly for bacteriophages T7, T3, YeO3-12 and lambda, based solely on the genome sequence. For 11 bacteriophages, new predictions of conserved phage-specific putative regulatory elements were made, further corroborating this approach. AVAILABILITY: http://www.agr.kuleuven.ac.be/logt/PHIRE.htm. Freely available for academic use. Commercial users should contact the corresponding author.

Bacteriophages↗

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article↗

The mouse Vcs2 gene is a composite structure which evolved by gene fusion and encodes five distinct salivary mRNA species.

Genes of the VCS (variable coding sequence) family are characterized by an extensive evolutionary divergence in the protein-coding sequence. The VCS family has been characterized by cDNA cloning from submandibular glands in the rat, mouse and humans. At the genomic level, the sequences of two members of this family are known in the rat Rattus norvegicus: the VCSA1 gene, encoding the prohormone-like polypeptide SMR1, and the VCSB1 gene, encoding a salivary Pro-rich polypeptide. No genomic data were available for the VCS genes of other species. To understand the evolution of the VCS gene family better, we have now sequenced 23 kilobases (kb) of the mouse Vcs2 gene. The Vcs2 sequence reveals numerous genomic reorganizations such as an inversion, insertions of short elements and an unusually high number of long interspersed repeated elements (LINEs), which make up 42% of this region. Interestingly, Vcs2 is composed of three different VCS-like regions. The first of these regions contains all the exons necessary to encode the previously described mouse submandibular gland polypeptide MSG2alpha. This region aligns with the entire genomic sequences of rat VCSA1 and VCSB1 genes. The two other regions align with fragments of these rat sequences. The three regions are arrayed in tandem and flanked by LINEs. In particular, the third region also contains exons that were found in mRNA species from the submandibular gland. In total, we have characterized five mRNAs from mouse submandibular glands which have in common their first exon, and are produced by alternative splicing. Vcs2 is thus a single gene that arose by the fusion of three genes (or pseudogenes) of the VCS multigene family.

Alternative Splicing↗

Software for optimization of SNP and PCR-RFLP genotyping to discriminate many genomes with the fewest assays.

BACKGROUND: Microbial forensics is important in tracking the source of a pathogen, whether the disease is a naturally occurring outbreak or part of a criminal investigation. RESULTS: A method and SPR Opt (SNP and PCR-RFLP Optimization) software to perform a comprehensive, whole-genome analysis to forensically discriminate multiple sequences is presented. Tools for the optimization of forensic typing using Single Nucleotide Polymorphism (SNP) and PCR-Restriction Fragment Length Polymorphism (PCR-RFLP) analyses across multiple isolate sequences of a species are described. The PCR-RFLP analysis includes prediction and selection of optimal primers and restriction enzymes to enable maximum isolate discrimination based on sequence information. SPR Opt calculates all SNP or PCR-RFLP variations present in the sequences, groups them into haplotypes according to their co-segregation across those sequences, and performs combinatoric analyses to determine which sets of haplotypes provide maximal discrimination among all the input sequences. Those set combinations requiring that membership in the fewest haplotypes be queried (i.e. the fewest assays be performed) are found. These analyses highlight variable regions based on existing sequence data. These markers may be heterogeneous among unsequenced isolates as well, and thus may be useful for characterizing the relationships among unsequenced as well as sequenced isolates. The predictions are multi-locus. Analyses of mumps and SARS viruses are summarized. Phylogenetic trees created based on SNPs, PCR-RFLPs, and full genomes are compared for SARS virus, illustrating that purported phylogenies based only on SNP or PCR-RFLP variations do not match those based on multiple sequence alignment of the full genomes. CONCLUSION: This is the first software to optimize the selection of forensic markers to maximize information gained from the fewest assays, accepting whole or partial genome sequence data as input. As more sequence data becomes available for multiple strains and isolates of a species, automated, computational approaches such as those described here will be essential to make sense of large amounts of information, and to guide and optimize efforts in the laboratory. The software and source code for SPR Opt is publicly available and free for non-profit use at http://www.llnl.gov/IPandC/technology/software/softwaretitles/spropt.php.

Cluster Analysis↗

Pair stochastic tree adjoining grammars for aligning and predicting pseudoknot RNA structures.

MOTIVATION: Since the whole genome sequences for many species are currently available, computational predictions of RNA secondary structures and computational identifications of those non-coding RNA regions by comparative genomics become important, and require more advanced alignment methods. Recently, an approach of structural alignments for RNA sequences has been introduced to solve these problems. By structural alignments, we mean a pairwise alignment to align an unfolded RNA sequence into a folded RNA sequence of known secondary structure. Pair HMMs on tree structures (PHMMTSs) proposed by Sakakibara are efficient automata-theoretic models for structural alignments of RNA secondary structures, but are incapable of handling pseudoknots. On the other hand, tree adjoining grammars (TAGs) is a subclass of context-sensitive grammar, which is suitable for modeling pseudoknots. Our goal is to extend PHMMTSs by incorporating TAGs to be able to handle pseudoknots. RESULTS: We propose the pair stochastic tree adjoining grammars (PSTAGs) for modeling RNA secondary structures including pseudoknots and show the strong experimental evidences that modeling pseudoknot structures significantly improves the prediction accuracies of RNA secondary structures. First, we extend the notion of PHMMTSs defined on alignments of 'trees' to PSTAGs defined on alignments of "TAG (derivation) trees", which represent a top-down parsing process of TAGs and are functionally equivalent to derived trees of TAGs. Second, we modify PSTAGs so that it takes as input a pair of a linear sequence and a TAG tree representing a pseudoknot structure of RNA to produce a structural alignment. Then, we develop a polynomial-time algorithm for obtaining an optimal structural alignment by PSTAGs, based on dynamic programming parser. We have done several computational experiments for predicting pseudoknots by PSTAGs, and our computational experiments suggests that prediction of RNA pseudoknot structures by our method are more efficient and biologically plausible than by other conventional methods. The binary code for PSTAG method is freely available from our website at http://www.dna.bio.keio.ac.jp/pstag/.

Algorithms↗

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

Computational discovery of internal micro-exons.

Very short exons, also known as micro-exons, occur in large numbers in some eukaryotic genomes. Existing annotation tools have a limited ability to recognize these short sequences, which range in length up to 25 bp. Here, we describe a computational method for the identification of micro-exons using near-perfect alignments between cDNA and genomic DNA sequences. Using this method, we detected 319 micro-exons in 4 complete genomes, of which 224 were previously unknown, human (170), the nematode Caenorhabditis elegans (4), the fruit fly Drosophila melanogaster (14), and the mustard plant Arabidopsis thaliana (36). Comparison of our computational method with popular cDNA alignment programs shows that the new algorithm is both efficient and accurate. The algorithm also aids in the discovery of micro-exon-skipping events and cross-species micro-exon conservation.

Alternative Splicing↗

Identification of two HIV type 1 circulating recombinant forms in Brazil.

Recombination is an important way to generate genetic diversity. Accumulation of HIV-1 full-length genomes in databases demonstrated that recombination is pervasive in viral strains collected globally. Recombinant forms achieving epidemiological relevance are termed circulating recombinant forms (CRFs). CRF12_BF was up to now the only CRF described in South America. The objective was to identify the first CRF in Brazil conducting full genome analysis of samples sharing the same partial genome recombinant structure. Ten samples obtained from individuals residing in Santos, Brazil, sharing the same recombination pattern based on partial genome sequence data, were selected from a larger group to undergo full length genome analysis. Near full length genomes were assembled from overlapping fragments. Mosaic genomes were evaluated by Bootscan, alignment inspection, and phylogenetic analysis using neighbor joining and maximum likelihood. Full genomes were also analyzed by split decomposition. We were able to identify five mosaic genomes. Two of these structures were represented by at least three samples derived from epidemiologically unlinked individuals. These structures were named CRF28_BF and CRF29_BF and are the second and third CRFs composed exclusively by subtypes B and F as well as the second and third CRFs encountered in South America. Other recombinant forms studied here resembled CRF28_BF and CRF29_BF. Our results suggest that a diverse population of related recombinants, including CRFs may play an important part in the Brazilian and South American epidemic.

Adult↗