Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

AGenDA: homology-based gene prediction.

We present a www server for homology-based gene prediction. The user enters a pair of evolutionary related genomic sequences, for example from human and mouse. Our software system uses CHAOS and DIALIGN to calculate an alignment of the input sequences and then searches for conserved splicing signals and start/stop codons around regions of local sequence similarity. This way, candidate exons are identified that are used, in turn, to calculate optimal gene models. The server returns the constructed gene model by email, together with a graphical representation of the underlying genomic alignment.

Algorithms↗

The bovine fatty acid binding protein 4 gene is significantly associated with marbling and subcutaneous fat depth in Wagyu x Limousin F2 crosses.

Fatty acid binding protein 4 (FABP4), which is expressed in adipose tissue, interacts with peroxisome proliferator-activated receptors and binds to hormone-sensitive lipase and therefore, plays an important role in lipid metabolism and homeostasis in adipocytes. The objective of this study was to investigate associations of the bovine FABP4 gene with fat deposition. Both cDNA and genomic DNA sequences of the bovine gene were retrieved from the public databases and aligned to determine its genomic organization. Primers targeting two regions of the FABP4 gene were designed: from nucleotides 5433-6106 and from nucleotides 7417-7868 (AAFC01136716). Direct sequencing of polymerase chain reaction (PCR) products on two DNA pools from high- and low-marbling animals revealed two single nucleotide polymorphisms (SNPs): AAFC01136716.1:g.7516G>C and g.7713G>C. The former SNP, detected by PCR-restriction fragment length polymorphism using restriction enzyme MspA1I, was genotyped on 246 F2 animals in a Waygu x Limousin F2 reference population. Statistical analysis showed that the FABP4 genotype significantly affected marbling score (P = 0.0398) and subcutaneous fat depth (P = 0.0246). The FABP4 gene falls into a suggestive/significant quantitative trait loci interval for beef marbling that was previously reported on bovine chromosome 14 in three other populations.

Animals↗

Utility and distribution of conserved noncoding sequences in the grasses.

Control of gene expression requires cis-acting regulatory DNA sequences. Historically these sequences have been difficult to identify. Conserved noncoding sequences (CNSs) have recently been identified in mammalian genes through cross-species genomic DNA comparisons, and some have been shown to be regulatory sequences. Using sequence alignment algorithms, we compared genomic noncoding DNA sequences of the liguleless1 (lg1) genes in two grasses, maize and rice, and found several CNSs in lg1. These CNSs are present in multiple grass species that represent phylogenetically disparate lineages. Six other maize/rice genes were compared and five contained CNSs. Based on nucleotide substitution rates, these CNSs exist because they have biological functions. Our analysis suggests that grass CNSs are smaller and far less frequent than those identified in mammalian genes and that mammalian gene regulation may be more complex than that of grasses. CNSs make excellent pan-grass PCR-based genetic mapping tools. They should be useful as characters in phylogenetic studies and as monitors of gene regulatory complexity.

Animals↗

An Eulerian path approach to global multiple alignment for DNA sequences.

With the rapid increase in the dataset of genome sequences, the multiple sequence alignment problem is increasingly important and frequently involves the alignment of a large number of sequences. Many heuristic algorithms have been proposed to improve the speed of computation and the quality of alignment. We introduce a novel approach that is fundamentally different from all currently available methods. Our motivation comes from the Eulerian method for fragment assembly in DNA sequencing that transforms all DNA fragments into a de Bruijn graph and then reduces sequence assembly to a Eulerian path problem. The paper focuses on global multiple alignment of DNA sequences, where entire sequences are aligned into one configuration. Our main result is an algorithm with almost linear computational speed with respect to the total size (number of letters) of sequences to be aligned. Five hundred simulated sequences (averaging 500 bases per sequence and as low as 70% pairwise identity) have been aligned within three minutes on a personal computer, and the quality of alignment is satisfactory. As a result, accurate and simultaneous alignment of thousands of long sequences within a reasonable amount of time becomes possible. Data from an Arabidopsis sequencing project is used to demonstrate the performance.

Algorithms↗

Characterization of a novel genital human papillomavirus by overlapping PCR: candHPV86 identified in cervicovaginal cells of a woman with cervical neoplasia.

A novel human papillomavirus (HPV), candHPV86, was cloned and characterized from cervicovaginal cells obtained from a 37-year-old Hispanic woman with cervical intraepithelial neoplasia grade 1 (CIN1) using an overlapping PCR technique. Primers were designed by phylogenetic alignment of closely related HPV genomes using the L1 fragment sequence amplified by GP5+/6+. The 7983 bp complete nucleotide sequence of the HPV genome was determined by sequence walking. A basic local alignment sequence tool (BLAST) homology search using the L1 open reading frame demonstrated that this HPV was most closely related to HPVHAN2294 (GenBank, AJ400628; 86% homology) and HPV84 (84% homology). candHPV86 was placed in the HPV genome homology group A3 by phylogenetic analyses. The overlapping PCR technique is applicable for characterizing the complete spectrum and variation of HPVs in a population.

Adult↗

Fold and function predictions for Mycoplasma genitalium proteins.

BACKGROUND: Uncharacterized proteins from newly sequenced genomes provide perfect targets for fold and function prediction. RESULTS: For 38% of the entire genome of Mycoplasma genitalium, sequence similarity to a protein with a known structure can be recognized using a new sequence alignment algorithm. When comparing genomes of M. genitalium and Escherichia coli, > 80% of M. genitalium proteins have a significant sequence similarity to a protein in E. coli and there are > 40 examples that have not been recognized before. For all cases of proteins with significant profile similarities, there are strong analogies in their functions, if the functions of both proteins are known. The results presented here and other recent results strongly support the argument that such proteins are actually homologous. Assuming this homology allows one to make tentative functional assignments for > 50 previously uncharacterized proteins, including such intriguing cases as the putative beta-lactam antibiotic resistance protein in M. gentalium. CONCLUSIONS: Using a new profile-to-profile alignment algorithm, the three-dimensional fold can be predicted for almost 40% of proteins from a genome of the small bacterium M. genitalium, and tentative function can be assigned to almost 80% of the entire genome. Some predictions lead to new insights about known functions or point to hitherto unexpected features of M. genitalium.

Bacterial Proteins↗

Dictyostelium transposable element DIRS-1 preferentially inserts into DIRS-1 sequences.

Sequence analysis of genomic clones containing the intact Dictyostelium transposable element DIRS-1 reveals that in five of six cases DIRS-1 has inserted into other DIRS-1 sequences. The nucleotide sequences just beyond the endpoints of the terminal repeats of five different genomic clones can be aligned with different regions of the internal nucleotide sequence of DIRS-1. In the three genomic clones which contain flanking sequences on both sides of the element, both flanking sequences are homologous with DIRS-1. In one of these clones, both extended flanking sequences represent the full 4.1-kilobase EcoRI fragment of DIRS-1, which has been interrupted by the insertion of an intact DIRS-1 element. There is no duplication or deletion (except possibly 1 base) of the DIRS-1 sequence upon insertion of a second DIRS-1 transposon. DIRS-1-into-DIRS-1 insertions can occur in either a colinear or inverted orientation with respect to the target sequence; the target sequence need not be an intact DIRS-1 element. We also describe a cDNA clone which could be derived by transcription of a sequence that resulted from a DIRS-1-into-DIRS-1 insertion and discuss its significance concerning the function of the heat-shock promoters found in the terminal repeats of DIRS-1 and in other DIRS-1-related sequences.

Base Sequence↗

Repression of the herpes simplex virus 1 alpha 4 gene by its gene product (ICP4) within the context of the viral genome is conditioned by the distance and stereoaxial alignment of the ICP4 DNA binding site relative to the TATA box.

Infected cell protein no. 4 (ICP4), the major regulatory protein encoded by the alpha 4 gene of herpes simplex virus 1, binds to a site (alpha 4-2) at the transcription initiation site of the alpha 4 gene. An earlier report described the construction of recombinant viruses that contained chimeric genes (alpha 4-tk) that consisted of the 5' untranscribed and transcribed noncoding domains of the alpha 4 gene fused to the coding sequences of the thymidine kinase gene and showed that disruption of the alpha 4-2 binding site by mutagenesis derepressed transcription of this gene (N. Michael and B. Roizman, Proc. Natl. Acad. Sci. USA 90:2286-2290, 1993). This experimental design was used to determine the effect of displacement of the alpha 4-2 binding site on the repression of alpha 4 gene transcription by ICP4. We report the following findings. (i) In the absence of the alpha 4-2 binding site, at 4 h after infection, alpha 4-tk RNA levels increased 10-fold relative to the corresponding RNA levels of a gene that contained the alpha 4-2 site at its natural location. Displacement of the alpha 4-2 binding site by approximately one, two, and three turns of the DNA helix, i.e., by 10, 21, and 30 nucleotides downstream of the original site, increased the concentration of alpha 4-tk RNA 2.4-, 3.5-, and 5.8-fold, respectively. (ii) Displacement of 16 nucleotides, i.e., approximately 1.5 helical turns, increased the accumulation of alpha 4-tk by 5.3-fold, i.e., more than predicted by displacement alone. (iii) At 8 h after infection in the absence of the binding site, the accumulation of alpha 4-tk RNA increased 13.6-fold. However, in cells infected with recombinants that carried displaced alpha 4-2 binding sites, RNA accumulation decreased relative to the levels seen at 4 h after infection. The insertion of DNA sequences in order to displace the alpha 4-2 binding site had no effect on accumulation of RNA in the presence of cycloheximide, i.e., in the absence of ICP4, or on maximum accumulation of alpha 4-tk RNA in the absence of the alpha 4-2 binding site.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Multiple genome rearrangement and breakpoint phylogeny.

Multiple alignment of macromolecular sequences generalizes from N = 2 to N > or = 3 the comparison of N sequences which have diverged through the local processes of insertion, deletion and substitution. Gene-order sequences diverge through non-local genome rearrangement processes such as inversion (or reversal) and transposition. In this paper we show which formulations of multiple alignment have counterparts in multiple rearrangement. Based on difficulties inherent in rearrangement edit-distance calculation and interpretation, we argue for the simpler "breakpoint analysis." Consensus-based multiple rearrangement of N > or = 3 orders can be solved exactly through reduction to instances of the Travelling Salesman Problem (TSP). We propose a branch-and-bound solution to TSP particularly suited to these instances. Simulations show how non-uniqueness of the solution is attenuated with increasing numbers of data genomes. Tree-based multiple alignment can be achieved to a great degree of accuracy by decomposing the tree into a number of overlapping 3-stars centered on the non-terminal nodes, and solving the consensus-based problem iteratively for these nodes until convergence. Accuracy improves with very careful initializations at the non-terminal nodes. The degree of non-uniqueness of solutions depends on the position of the node in the tree in terms of path length to the terminal vertices.

Algorithms↗

rVISTA 2.0: evolutionary analysis of transcription factor binding sites.

Identifying and characterizing the transcription factor binding site (TFBS) patterns of cis-regulatory elements represents a challenge, but holds promise to reveal the regulatory language the genome uses to dictate transcriptional dynamics. Several studies have demonstrated that regulatory modules are under positive selection and, therefore, are often conserved between related species. Using this evolutionary principle, we have created a comparative tool, rVISTA, for analyzing the regulatory potential of noncoding sequences. Our ability to experimentally identify functional noncoding sequences is extremely limited, therefore, rVISTA attempts to fill this great gap in genomic analysis by offering a powerful approach for eliminating TFBSs least likely to be biologically relevant. The rVISTA tool combines TFBS predictions, sequence comparisons and cluster analysis to identify noncoding DNA regions that are evolutionarily conserved and present in a specific configuration within genomic sequences. Here, we present the newly developed version 2.0 of the rVISTA tool, which can process alignments generated by both the zPicture and blastz alignment programs or use pre-computed pairwise alignments of several vertebrate genomes available from the ECR Browser and GALA database. The rVISTA web server is closely interconnected with the TRANSFAC database, allowing users to either search for matrices present in the TRANSFAC library collection or search for user-defined consensus sequences. The rVISTA tool is publicly available at http://rvista.dcode.org/.

Algorithms↗

Gene transfer to the nucleus and the evolution of chloroplasts.

Photosynthetic eukaryotes, particularly unicellular forms, possess a fossil record that is either wrought with gaps or difficult to interpret, or both. Attempts to reconstruct their evolution have focused on plastid phylogeny, but were limited by the amount and type of phylogenetic information contained within single genes. Among the 210 different protein-coding genes contained in the completely sequenced chloroplast genomes from a glaucocystophyte, a rhodophyte, a diatom, a euglenophyte and five land plants, we have now identified the set of 45 common to each and to a cyanobacterial outgroup genome. Phylogenetic inference with an alignment of 11,039 amino-acid positions per genome indicates that this information is sufficient--but just rarely so--to identify the rooted nine-taxon topology. We mapped the process of gene loss from chloroplast genomes across the inferred tree and found that, surprisingly, independent parallel gene losses in multiple lineages outnumber phylogenetically unique losses by more that 4:1. We identified homologues of 44 different plastid-encoded proteins as functional nuclear genes of chloroplast origin, providing evidence for endosymbiotic gene transfer to the nucleus in plants.

Cell Nucleus↗

Genomic organization of zebrafish cone-rod homeobox gene and exclusion as a candidate gene for retinal degeneration in niezerka and mikre oko.

PURPOSE: To determine the genomic organization of the zebrafish crx gene and to evaluate if mutations in crx are responsible for the retinal degeneration phenotype in the zebrafish (Danio rerio) mutants niezerka (nie(m743)) and mikre oko (mok(m632)). METHODS: Overlapping fragments were PCR amplified from genomic DNA isolated from homozygous mutant embryos and wild-type siblings (sibs). Amplicons were sequenced and sequence data assembled into contigs. Genomic organization was determined by alignment of contigs with published cDNA sequences and zebrafish genomic sequence from Sanger and Ensembl databases. Linkage analysis used DNA from mapping panels of single homozygous mutant animals with mixed genetic backgrounds. RESULTS: The analysis indicated that the zebrafish crx gene consisted of three exons and 2 introns, and spans 3.8 kb of genomic DNA. The splice junctions were all located within the coding region. Highly repetitive sequences present in non-coding regions of crx and extended tetra-nucleotide repeats in intronic regions were associated with sequence variation between different strains. Homozygous mok(m632) or nie(m743) mutants and their respective wild-type sibs, showed identical patterns of heterozygosity and sequence variations within each line. No mutation in crx were identified in homozygous mok(m632) or nie(m743). Consistent with the absence of identified mutations, linkage analysis excluded linkage of the mutant phenotypes to crx. CONCLUSIONS: Despite the presence of sequence variations in their respective genetic backgrounds, within each line the sequence of crx was identical. Consistent with the absence of mutations, further analysis excluded linkage of the mutant phenotypes to crx. Analysis is in progress to map these loci and identify the genes responsible for the retinal degeneration phenotype in these mutant lines.

Animals↗

An adaptive and iterative algorithm for refining multiple sequence alignment.

Multiple sequence alignment is a basic tool in computational genomics. The art of multiple sequence alignment is about placing gaps. This paper presents a heuristic algorithm that improves multiple protein sequences alignment iteratively. A consistency-based objective function is used to evaluate the candidate moves. During the iterative optimization, well-aligned regions can be detected and kept intact. Columns of gaps will be inserted to assist the algorithm to escape from local optimal alignments. The algorithm has been evaluated using the BAliBASE benchmark alignment database. Results show that the performance of the algorithm does not depend on initial or seed alignments much. Given a perfect consistency library, the algorithm is able to produce alignments that are close to the global optimum. We demonstrate that the algorithm is able to refine alignments produced by other software, including ClustalW, SAGA and T-COFFEE. The program is available upon request.

Algorithms↗

A physical map of the chicken genome.

Strategies for assembling large, complex genomes have evolved to include a combination of whole-genome shotgun sequencing and hierarchal map-assisted sequencing. Whole-genome maps of all types can aid genome assemblies, generally starting with low-resolution cytogenetic maps and ending with the highest resolution of sequence. Fingerprint clone maps are based upon complete restriction enzyme digests of clones representative of the target genome, and ultimately comprise a near-contiguous path of clones across the genome. Such clone-based maps are used to validate sequence assembly order, supply long-range linking information for assembled sequences, anchor sequences to the genetic map and provide templates for closing gaps. Fingerprint maps are also a critical resource for subsequent functional genomic studies, because they provide a redundant and ordered sampling of the genome with clones. In an accompanying paper we describe the draft genome sequence of the chicken, Gallus gallus, the first species sequenced that is both a model organism and a global food source. Here we present a clone-based physical map of the chicken genome at 20-fold coverage, containing 260 contigs of overlapping clones. This map represents approximately 91% of the chicken genome and enables identification of chicken clones aligned to positions in other sequenced genomes.

Animals↗

Pair stochastic tree adjoining grammars for aligning and predicting pseudoknot RNA structures.

MOTIVATION: Since the whole genome sequences of many species have been determined, computational prediction of RNA secondary structures and computational identification of those non-coding RNA regions by comparative genomics become important. Therefore, more advanced alignment methods are required. Recently, an approach of structural alignment for RNA sequences has been introduced to solve these problems. Pair hidden Markov models on tree structures (PHMMTSs) proposed by Sakakibara are efficient automata-theoretic models for structural alignment of RNA secondary structures, although PHMMTSs are incapable of handling pseudoknots. On the other hand, tree adjoining grammars (TAGs), a subclass of context-sensitive grammars, are suitable for modeling pseudoknots. Our goal is to extend PHMMTSs by incorporating TAGs to be able to handle pseudoknots. RESULTS: We propose pair stochastic TAGs (PSTAGs) for aligning and predicting RNA secondary structures including a simple type of pseudoknot which can represent most known pseudoknot structures. First, we extend PHMMTSs defined on alignment of 'trees' to PSTAGs defined on alignment of 'TAG trees' which represent derivation processes of TAGs and are functionally equivalent to derived trees of TAGs. Then, we develop an efficient dynamic programming algorithm of PSTAGs for obtaining an optimal structural alignment including pseudoknots. We implement the PSTAG algorithm and demonstrate the properties of the algorithm by using it to align and predict several small pseudoknot structures. We believe that our implemented program based on PSTAGs is the first grammar-based and practically executable software for comparative analyses of RNA pseudoknot structures, and, further, non-coding RNAs.

Algorithms↗

AGenDA: gene prediction by comparative sequence analysis.

UNLABELLED: Comparative sequence analysis is a powerful approach to identify functional elements in genomic sequences. Herein, we describe AGenDA (Alignment-based GENe Detection Algorithm), a novel method for gene prediction that is based on long-range alignment of syntenic regions in eukaryotic genome sequences. Local sequence homologies identified by the DIALIGN program are searched for conserved splice signals to define potential protein-coding exons; these candidate exons are then used to assemble complete gene structures. The performance of our method was tested on a set of 105 human-mouse sequence pairs. These test runs showed that sensitivity and specificity of AGenDA are comparable with the best gene- prediction program that is currently available. However, since our method is based on a completely different type of input information, it can detect genes that are not detectable by standard methods and vice versa. Thus, our approach seems to be a useful addition to existing gene-prediction programs. AVAILABILITY: DIALIGN is available through the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak.uni-bielefeld.de/dialign/ The gene-prediction program AGenDA described in this paper will be available through the BiBiServ or MIPS web server at http://mips.gsf.de.

Algorithms↗

Alu-associated enhancement of single nucleotide polymorphisms in the human genome.

Identifying features shaping the architecture of sequence variations is important for understanding genome evolution and mapping disease loci. In this study, high-resolution scanning of Alu-centered alignments of the human genome sequences has revealed a striking elevation of the frequency of single nucleotide polymorphisms (SNP) in the body and tail of Alu sequences compared to flanking regions. This enhancement in SNP density is evident for all twenty-four chromosomes, and in both the Alu-body and Alu-tail, which together may be referred to as the Alu-SNPs. Reduced levels of Alu-SNPs in the sex chromosomes, especially in the non-recombining NRY region of the Y chromosome, are consistent with recombination events playing an important role in the enhancement. The Alu elements are unstable recombination-mutation hotspots in the human genome, and it is suggested that the Alu-SNPs represent a key manifestation of this instability. Variations in Alu-SNPs among the HapMap populations of northern and western European ancestry (CEU), Han Chinese from Beijing (CHB), Japanese from Tokyo (JPT), and Yoruba from Ibadan, Nigeria (YRI) indicate that the Alu-SNPs provide useful sequence markers, in addition to the Alu-insertion polymorphisms themselves, for the delineation of human genome evolution. That Alu-SNP levels are highest in the youngest Alu-Y, intermediate in the Alu-S of intermediate age, and lowest in the oldest Alu-J is consistent with the occurrence of not only genetic drift but also natural selection on the Alu-SNPs. Such evolutionary selection in turn suggests that Alu-SNPs might include potential sites of disease association, and therefore deserve detailed investigation.

Alu Elements↗