Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Mitogenomic evolution and interrelationships of the Cypriniformes (Actinopterygii: Ostariophysi): the first evidence toward resolution of higher-level relationships of the world's largest freshwater fish clade based on 59 whole mitogenome sequences.

Fishes of the order Cypriniformes are almost completely restricted to freshwater bodies and number > 3400 species placed in 5 families, each with poorly defined subfamilies and/or tribes. The present study represents the first attempt toward resolution of the higher-level relationships of the world's largest freshwater-fish clade based on whole mitochondrial (mt) genome sequences from 53 cypriniforms (including 46 newly determined sequences) plus 6 outgroups. Unambiguously aligned, concatenated mt genome sequences (14,563 bp) were divided into 5 partitions (first, second, and third codon positions of the protein-coding genes, rRNA genes, and tRNA genes), and partitioned Bayesian analyses were conducted, with protein-coding genes being treated in 3 different manners (all positions included; third codon positions converted into purine [R] and pyrimidine [Y] [RY-coding]; third codon positions excluded). The resultant phylogenies strongly supported monophyly of the Cypriniformes as well as that of the families Cyprinidae, Catostomidae, and a clade comprising Balitoridae + Cobitidae, with the 2 latter loach families being reciprocally paraphyletic. Although all of the data sets yielded nearly identical tree topologies with regard to the shallower relationships, deeper relationships among the 4 major clades (the above 3 major clades plus Gyrinocheilidae, represented by a single species Gyrinocheilus aymonieri in this study), were incongruent depending on the data sets. Treatment of the rapidly saturated third codon-position transitions appeared to be a source of such incongruities, and we advocate that RY-coding, which takes only transversions into account, effectively removes this likely "noise" from the data set and avoids the apparent lack of signal by retaining all available positions in the data set.

Animals↗

Recent advances in gene structure prediction.

De novo gene predictors are programs that predict the exon-intron structures of genes using the sequences of one or more genomes as their only input. In the past two years, dual-genome de novo predictors, which exploit local rates and patterns of mutation inferred from alignments between two genomes, have led to significant improvements in accuracy. Systems that exploit more than two genomes simultaneously have only recently begun to appear and are not yet competitive on practical tasks, but offer the greatest hope for near-term improvements. Dual-genome de novo prediction for compact eukaryotic genomes such as those of Arabidopsis thaliana and Caenorhabditis elegans is already quite accurate. Although mammalian gene prediction lags behind in accuracy, it is yielding ever more useful results. Coupled with significant improvements in pseudogene detection methods, which have eliminated many false positives, we have reached the point where de novo gene predictions are being used as hypotheses to drive experimental annotation via systematic RT-PCR and sequencing.

Animals↗

Evolutionary turnover of mammalian transcription start sites.

Alignments of homologous genomic sequences are widely used to identify functional genetic elements and study their evolution. Most studies tacitly equate homology of functional elements with sequence homology. This assumption is violated by the phenomenon of turnover, in which functionally equivalent elements reside at locations that are nonorthologous at the sequence level. Turnover has been demonstrated previously for transcription-factor-binding sites. Here, we show that transcription start sites of equivalent genes do not always reside at equivalent locations in the human and mouse genomes. We also identify two types of partial turnover, illustrating evolutionary pathways that could lead to complete turnover. These findings suggest that the signals encoding transcription start sites are highly flexible and evolvable, and have cautionary implications for the use of sequence-level conservation to detect gene regulatory elements.

Animals↗

A tale of two templates: automatically resolving double traces has many applications, including efficient PCR-based elucidation of alternative splices.

Trace Recalling is a novel method for deconvoluting double traces that result from simultaneously sequencing two DNA templates. Trace Recalling identifies up to two bases at each position of such a trace. The resulting ambiguity sequence is aligned to the genome, identifying one template sequence. A second template sequence is then inferred from this alignment. This technique makes possible many exciting biological applications. Here we present two such applications, alternate splice finding and elucidation of multiple insertion sites in a random insertional mutagenesis library. Our results demonstrate that RT-PCR followed by Trace Recalling is a more efficient and cost effective way to find alternate splices than traditional methods. We also present a method for mapping double-insertion events in a random insertional-mutagenesis library.

Algorithms↗

The intronerator: exploring introns and alternative splicing in Caenorhabditis elegans.

The Intronerator (http://www.cse.ucsc.edu/ approximately kent/intronerator/ ) is a set of web-based tools for exploring RNA splicing and gene structure in Caenorhabditis elegans. It includes a display of cDNA alignments with the genomic sequence, a catalog of alternatively spliced genes and a database of introns. The cDNA alignments include >100 000 ESTs and almost 1000 full-length cDNAs. ESTs from embryos and mixed stage animals as well as full-length cDNAs can be compared in the alignment display with each other and with predicted genes. The alt-splicing catalog includes 844 open reading frames for which there is evidence of alternative splicing of pre-mRNA. The intron database includes 28 478 introns, and can be searched for patterns near the splice junctions.

Alternative Splicing↗

Theseus: fast and optimal affine-gap sequence-to-graph alignment.

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed. RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: MSA and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs. AVAILABILITY AND IMPLEMENTATION: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

Algorithms↗

Association between divergence and interspersed repeats in mammalian noncoding genomic DNA.

The amount of noncoding genomic DNA sequence that aligns between human and mouse varies substantially in different regions of their genomes, and the amount of repetitive DNA also varies. In this report, we show that divergence in noncoding nonrepetitive DNA is strongly correlated with the amount of repetitive DNA in a region. We investigated aligned DNA in four large genomic regions with finished human sequence and almost or completely finished mouse sequence. These regions, totaling 5.89 Mb of DNA, are on different chromosomes and vary in their base composition. An analysis based on sliding windows of 10 kb shows that the fraction of aligned noncoding nonrepetitive DNA and the fraction of repetitive DNA are negatively correlated, both at the level of an entire region and locally within it. This conclusion is strongly supported by a randomization study, in which repetitive elements are removed and randomly relocated along the sequences. Thus, regions of noncoding genomic DNA that accumulated fewer point mutations since the primate-rodent divergence also suffered fewer retrotransposition events. These results indicate that some regions of the genome are more "flexible" over the time scale of mammalian evolution, being able to accommodate many point mutations and insertions, whereas other regions are more "rigid" and accumulate fewer changes. Stronger conservation is generally interpreted as indicating more extensive or more important function. The evidence presented here of correlated variation in the rates of different evolutionary processes across noncoding DNA must be considered in assessing such conservation for evidence of selection.

Animals↗

The Drosophila gene collection: identification of putative full-length cDNAs for 70% of D. melanogaster genes.

Collections of full-length nonredundant cDNA clones are critical reagents for functional genomics. The first step toward these resources is the generation and single-pass sequencing of cDNA libraries that contain a high proportion of full-length clones. The first release of the Drosophila Gene Collection Release 1 (DGCr1) was produced from six libraries representing various tissues, developmental stages, and the cultured S2 cell line. Nearly 80,000 random 5' expressed sequence tags (5' expressed sequence tags [ESTs]from these libraries were collapsed into a nonredundant set of 5849 cDNAs, corresponding to ~40% of the 13,474 predicted genes in Drosophila. To obtain cDNA clones representing the remaining genes, we have generated an additional 157,835 5' ESTs from two previously existing and three new libraries. One new library is derived from adult testis, a tissue we previously did not exploit for gene discovery; two new cap-trapped normalized libraries are derived from 0-22-h embryos and adult heads. Taking advantage of the annotated D. melanogaster genome sequence, we clustered the ESTs by aligning them to the genome. Clusters that overlap genes not already represented by cDNA clones in the DGCr1 were analyzed further, and putative full-length clones were selected for inclusion in the new DGC. This second release of the DGC (DGCr2) contains 5061 additional clones, extending the collection to 10,910 cDNAs representing >70% of the predicted genes in Drosophila.

Animals↗

Tandem Duplication-Driven Neofunctionalization of UDP-Glycosyltransferases Shapes the Diversification of Triterpenoid Saponins in the Cucurbitaceae.

Tandem duplication of tailoring enzymes allows evolutionary innovation that diversifies plant specialized metabolism. Here, we present an interesting example of how tandem duplicated UDP-glycosyltransferases undergo neofunctionalization and shape the chemical diversity of triterpenoid saponins in the Cucurbitaceae family. A chromosome-level genome of Siraitia grosvenorii was assembled and aligned with multiple cucurbit genomes, revealing a specific UGT73AM tandem duplication responsible for regio-selective glycosylation (e.g. the rare 1,4-linked disaccharide) of diverse saponins such as mogrosides, ginsenosides, and momordicines. Comparative genomics depicted the evolutionary trajectory of a universal saponin-biosynthesizing UGT73 tandem arrays syntenously preserved across core eudicots, where lineage-specific UGT copies contribute to distinct metabolic phenotypes. A crystal structure of SgUGT73AM30 (mogrol 25-O-glycosyltransferase) in complex with UDP and mogrol was obtained to elucidate the molecular basis of the regio-specific decoration on vicinal diol of the substrates. Altogether, these findings provide insights into tandem duplication-driven diversification of glycosyltransferases and lay the foundation for engineered glycosylation of valuable triterpenoid saponins.

Saponins↗

Coronaviruses use discontinuous extension for synthesis of subgenome-length negative strands.

We have developed a new model for coronavirus transcription, which we call discontinuous extension, to explain how subgenome-length negatives stands are derived directly from the genome. The current model called leader-primed transcription, which states that subgenomic mRNA is transcribed directly from genome-length negative-strands, cannot explain many of the recent experimental findings. For instance, subgenomic mRNAs are transcribed directly via transcription intermediates that contain subgenome-length negative-strand templates; however subgenomic mRNA does not appear to be copied directly into negative strands. In our model the subgenome-length negative strands would be derived using the genome as a template. After the polymerase had copied the 3'-end of the genome, it would detach at any one of the several intergenic sequences and reattach to the sequence immediately downstream of the leader sequence at the 5'-end of genome RNA. Base pairing between the 3'-end of the nascent subgenome-length negative strands, which would be complementary to the intergenic sequence at the end of the leader sequence at the 5'-end of genome, would serve to align the nascent negative strand to the genome and permit the completion of synthesis, i.e., discontinuous extension of the 3'-end of the negative strand. Thus, subgenome-length negative strands would arise by discontinuous synthesis, but of negative strands, not of positive strands as proposed originally by the leader-primed transcription model.

Animals↗

Covariation in frequencies of substitution, deletion, transposition, and recombination during eutherian evolution.

Six measures of evolutionary change in the human genome were studied, three derived from the aligned human and mouse genomes in conjunction with the Mouse Genome Sequencing Consortium, consisting of (1) nucleotide substitution per fourfold degenerate site in coding regions, (2) nucleotide substitution per site in relics of transposable elements active only before the human-mouse speciation, and (3) the nonaligning fraction of human DNA that is nonrepetitive or in ancestral repeats; and three derived from human genome data alone, consisting of (4) SNP density, (5) frequency of insertion of transposable elements, and (6) rate of recombination. Features 1 and 2 are measures of nucleotide substitutions at two classes of "neutral" sites, whereas 4 is a measure of recent mutations. Feature 3 is a measure dominated by deletions in mouse, whereas 5 represents insertions in human. It was found that all six vary significantly in megabase-sized regions genome-wide, and many vary together. This indicates that some regions of a genome change slowly by all processes that alter DNA, and others change faster. Regional variation in all processes is correlated with, but not completely accounted for, by GC content in human and the difference between GC content in human and mouse.

Animals↗

Human papillomavirus type 48.

The cloning and partial characterization of the genome of human papillomavirus (HPV) type 48 is presented. Hybridization and short DNA sequence analyses permitted the alignment of the genome to the HPV genetic map.

Carcinoma, Squamous Cell↗

Evidence of recombination among enteroviruses.

Human enteroviruses consist of more than 60 serotypes, reflecting a wide range of evolutionary divergence. They have been genetically classified into four clusters on the basis of sequence homology in the coding region of the single-stranded RNA genome. To explore further the genetic relationships between human enteroviruses and to characterize the evolutionary mechanisms responsible for variation, previously sequenced genomes were subjected to detailed comparison. Bootstrap and genetic similarity analyses were used to systematically scan the alignments of complete genomic sequences. Bootstrap analysis provided evidence from an early recombination event at the junction of the 5' noncoding and coding regions of the progenitors of the current clusters. Analysis within the genetic clusters indicated that enterovirus prototype strains include intraspecies recombinants. Recombination breakpoints were detected in all genomic regions except the capsid protein coding region. Our results suggest that recombination is a significant and relatively frequent mechanism in the evolution of enterovirus genomes.

Enterovirus↗

A search tool for identification and analysis of conserved sequence patterns in Saccharomyces spp. orthologous promoter.

We describe a web-based resource to identify, search and analyze sequence patterns conserved in the multiple sequence alignments of orthologous promoters from closely related / distant Saccharomyces spp. The webtool interfaces with a database where conserved sequence patterns (greater than 4 bp) have been previously extracted from genome-wide promoter alignments, allowing one to carry out user-defined genome-wide searches for conserved sequences to assist in the discovery of novel promoter elements based on comparative genomics. The web-based server can be accessed at http://www2.imtech.res.in/ anand/sacch_prom_pat.html.

Base Sequence↗

An SNP map of the human genome generated by reduced representation shotgun sequencing.

Most genomic variation is attributable to single nucleotide polymorphisms (SNPs), which therefore offer the highest resolution for tracking disease genes and population history. It has been proposed that a dense map of 30,000-500,000 SNPs can be used to scan the human genome for haplotypes associated with common diseases. Here we describe a simple but powerful method, called reduced representation shotgun (RRS) sequencing, for creating SNP maps. RRS re-samples specific subsets of the genome from several individuals, and compares the resulting sequences using a highly accurate SNP detection algorithm. The method can be extended by alignment to available genome sequence, increasing the yield of SNPs and providing map positions. These methods are being used by The SNP Consortium, an international collaboration of academic centres, pharmaceutical companies and a private foundation, to discover and release at least 300,000 human SNPs. We have discovered 47,172 human SNPs by RRS, and in total the Consortium has identified 148,459 SNPs. More broadly, RRS facilitates the rapid, inexpensive construction of SNP maps in biomedically and agriculturally important species. SNPs discovered by RRS also offer unique advantages for large-scale genotyping.

Algorithms↗

PromAn: an integrated knowledge-based web server dedicated to promoter analysis.

PromAn is a modular web-based tool dedicated to promoter analysis that integrates distinct complementary databases, methods and programs. PromAn provides automatic analysis of a genomic region with minimal prior knowledge of the genomic sequence. Prediction programs and experimental databases are combined to locate the transcription start site (TSS) and the promoter region within a large genomic input sequence. Transcription factor binding sites (TFBSs) can be predicted using several public databases and user-defined motifs. Also, a phylogenetic footprinting strategy, combining multiple alignment of large genomic sequences and assignment of various scores reflecting the evolutionary selection pressure, allows for evaluation and ranking of TFBS predictions. PromAn results can be displayed in an interactive graphical user interface, PromAnGUI. It integrates all of this information to highlight active promoter regions, to identify among the huge number of TFBS predictions those which are the most likely to be potentially functional and to facilitate user refined analysis. Such an integrative approach is essential in the face of a growing number of tools dedicated to promoter analysis in order to propose hypotheses to direct further experimental validations. PromAn is publicly available at http://bips.u-strasbg.fr/PromAn.

Binding Sites↗

Combination of multiple alignment analysis and surface mapping paves a way for a detailed pathway reconstruction--the case of VHL (von Hippel-Lindau) protein and angiogenesis regulatory pathway.

Using the tumor suppressor VHL protein as an example, we show that detailed analysis of conservation versus variation pattern in the multiple alignment can be coupled with the genomic pathway/complex conservation analysis to provide a more complete picture of the entire interaction/regulatory network. Results from the present study have allowed us to hypothesize that two additional proteins are involved in the VHL-mediated regulation of angiogenesis. Detailed modeling also has led to a prediction of the possible interaction mode between the known and the proposed parts of the VHL complex. To aid in an analysis of the VHL protein regulation of HIF-1 alpha degradation, an important and only partially understood process that directly influences angiogenesis, we performed a comprehensive search for the orthologs of the VHL as well as for VHL-interacting proteins in all the available eukaryotic genomes. Analysis of a multiple alignment of thus identified VHL orthologs reveals an unusually high degree of conservation of the surface amino acid residues that almost exactly correspond to positions mutated in the VHL disease-associated tumors. In addition, these positions form well-defined clusters in three-dimensional space, and presence or absence of individual clusters correlates with the presence or absence of pathway elements in different genomes. We have also shown that relation trees derived from the multiple sequence alignment, functional surface-mapping, and HIF-1 alpha degradation pathway structure are in complete agreement, linking the functional and structural evolution of the VHL protein and VHL-dependent HIF-1 alpha degradation complex.

ATP-Dependent Proteases↗

PAK paradox: Paramecium appears to have more K(+)-channel genes than humans.

K(+)-selective ion channels (K(+) channels) have been found in bacteria, archaea, eucarya, and viruses. In Paramecium and other ciliates, K(+) currents play an essential role in cilia-based motility. We have retrieved and sequenced seven closely related Paramecium K(+)-channel gene (PAK) sequences by using previously reported fragments. An additional eight unique K(+)-channel sequences were retrieved from an indexed library recently used in a pilot genome sequencing project. Alignments of these protein translations indicate that while these 15 genes have diverged at different times, they all maintain many characteristics associated with just one subclass of metazoan K(+) channels (CNG/ERG type). Our results indicate that most of the genes are expressed, because all predicted frameshifts and several gaps in the homolog alignments contain Paramecium intron sequences deleted from reverse transcription-PCR products. Some of the variations in the 15 genomic nucleotide sequences involve an absence of introns, even between very closely related sequences, suggesting a potential occurrence of reverse transcription in the past. Extrapolation from the available genome sequence indicates that Paramecium harbors as many as several hundred of this one type of K(+)-channel gene. This quantity is far more numerous than those of K(+)-channel genes of all types known in any metazoan (e.g., approximately 80 in humans, approximately 30 in flies, and approximately 15 in Arabidopsis). In an effort to understand this plurality, we discuss several possible reasons for their maintenance, including variations in expression levels in response to changes in the freshwater environment, like that seen with other major plasma membrane proteins in Paramecium.

Alternative Splicing↗