Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Symmetry observations in long nucleotide sequences: a commentary on the Discovery Note of Qi and Cuticchia.

The relative quantities of bases in DNA were determined chemically many years before sequencing technologies permitted direct counting of bases. Apparently unaware of the rich literature on the topic, bioinformaticists are today rediscovering the 'wheels' of Chargaff, Wyatt and other biochemists. It follows from Chargaff's second parity rule (%A = %T, %G = %C for single stranded DNA) that the symmetries observed for the two pairs of complementary mononucleotide bases, should also apply to the eight pairs of complementary dinucleotide bases, the thirty-two pairs of complementary trinucleotide bases, etc. This was made explicit by Prabhu in 1993 in a study of complete genomes and long genome segments from a wide range of taxa, and was rediscovered by Qi and Cuticchia in 2001 in a study of complete genomes. It follows from Chargaff's GC-rule (%GC tends to be uniform and species specific) that, within a species, oligonucleotides of the same GC% will be at approximately equal quantities in single stranded DNA. Thus, for example, while quantities of CAT and ATG (reverse complements) will be closely correlated because of both of the above Chargaff rules, CAT and GTA (forward complements) will show some correlation only because of the latter rule. The need for complete genomic sequences in bioinformatic analyses may have been somewhat overplayed.

Base Sequence↗

Exploiting uniqueness: seed-chain-extend alignment on elastic founder graphs.

SUMMARY: Sequence-to-graph alignment is a central challenge of computational pangenomics. To overcome the theoretical hardness of the problem, state-of-the-art tools use seed-and-extend or seed-chain-extend heuristics to alignment. We implement a complete seed-chain-extend alignment workflow based on indexable elastic founder graphs (iEFGs) that support linear-time exact searches unlike general graphs. We show how to construct iEFGs, find high-quality seeds, chain, and extend them at the scale of a telomere-to-telomere assembled human chromosome. AVAILABILITY AND IMPLEMENTATION: Our sequence-to-graph alignment tool and the scripts to replicate our experiments are available in https://github.com/algbio/SRFAligner.

Software↗

scnanoseq: an nf-core pipeline for Oxford Nanopore single-cell RNA-sequencing.

MOTIVATION: Recent advancements in long-read single-cell RNA sequencing (scRNA-seq) have facilitated the quantification of full-length transcripts and isoforms at the single-cell level. Historically, long-read data would need to be complemented with short-read single-cell data in order to overcome the higher sequencing errors to correctly identify cellular barcodes and unique molecular identifiers. Improvements in Oxford Nanopore sequencing, and development of novel computational methods have removed this requirement. Though these methods now exist, the limited availability of modular and portable workflows remains a challenge. RESULTS: Here, we present, nf-core/scnanoseq, a secondary analysis pipeline for long-read single-cell and single-nuclei RNA that delivers gene and transcript-level quantification. The scnanoseq pipeline is implemented using Nextflow and is built upon the nf-core framework, enabling portability across computational environments, scalability and reproducibility of results across pipeline runs. The nf-core/scnanoseq workflow follows best practices for analyzing single-cell and single-nuclei data, performing barcode detection and correction, genome and transcriptome read alignment, unique molecular identifier deduplication, gene and transcript quantification, and extensive quality control reporting. AVAILABILITY AND IMPLEMENTATION: The source code, and detailed documentation are freely available at https://github.com/nf-core/scnanoseq and https://nf-co.re/scnanoseq under the MIT License. Documentation for the version of nf-core/scnanoseq used for this paper, including default parameters and descriptions of output files are available at https://nf-co.re/scnanoseq/1.1.0.

Single-Cell Analysis↗

GABI-Kat SimpleSearch: a flanking sequence tag (FST) database for the identification of T-DNA insertion mutants in Arabidopsis thaliana.

SUMMARY: GABI-Kat SimpleSearch is a database of flanking sequence tags (FSTs) of T-DNA mutagenized Arabidopsis thaliana lines that were generated by the GABI-Kat project. Sequences flanking the T-DNA insertion sites were aligned to the A.thaliana genome sequence, annotated with information about the FST, the insertion site and the line from which the FST was derived. A web interface permits text-based as well as sequence-based searches for relevant insertions. GABI-Kat SimpleSearch aims to help biologists to quickly find T-DNA insertion mutants for their research. AVAILABILITY: http://www.mpiz-koeln.mpg.de/GABI-Kat/

Arabidopsis↗

SynView: a GBrowse-compatible approach to visualizing comparative genome data.

UNLABELLED: We present SynView, a simple and generic approach to dynamically visualize multi-species comparative genome data. It is a light-weight application based on the popular and configurable web-based GBrowse framework. It can be used with a variety of databases and provides the user with a high degree of interactivity. The tool is written in Perl and runs on top of the GBrowse framework. It is in use in the PlasmoDB (http://www.PlasmoDB.org) and the CryptoDB (http://www.CryptoDB.org) projects and can be easily integrated into other cross-species comparative genome projects. AVAILABILITY: The program and instructions are freely available at http://www.ApiDB.org/apps/SynView/ CONTACT: jkissing@uga.edu.

Algorithms↗

Deriving the genomic tree of life in the presence of horizontal gene transfer: conditioned reconstruction.

The horizontal gene transfer (HGT) being inferred within prokaryotic genomes appears to be sufficiently massive that many scientists think it may have effectively obscured much of the history of life recorded in DNA. Here, we demonstrate that the tree of life can be reconstructed even in the presence of extensive HGT, provided the processes of genome evolution are properly modeled. We show that the dynamic deletions and insertions of genes that occur during genome evolution, including those introduced by HGT, may be modeled using techniques similar to those used to model nucleotide substitutions that occur during sequence evolution. In particular, we show that appropriately designed general Markov models are reasonable tools for reconstructing genome evolution. These studies indicate that, provided genomes contain sufficiently many genes and that the Markov assumptions are met, it is possible to reconstruct the tree of life. We also consider the fusion of genomes, a process not encountered in gene sequence evolution, and derive a method for the identification and reconstruction of genome fusion events. Genomic reconstructions of a well-defined classical four-genome problem, the root of the multicellular animals, show that the method, when used in conjunction with paralinear/logdet distances, performs remarkably well and is relatively unaffected by the recently discovered big genome artifact.

Computational Biology↗

PolyA_DB 2: mRNA polyadenylation sites in vertebrate genes.

Polyadenylation of nascent transcripts is one of the key mRNA processing events in eukaryotic cells. A large number of human and mouse genes have alternative polyadenylation sites, or poly(A) sites, leading to mRNA variants with different protein products and/or 3'-untranslated regions (3'-UTRs). PolyA_DB 2 contains poly(A) sites identified for genes in several vertebrate species, including human, mouse, rat, chicken and zebrafish, using alignments between cDNA/ESTs and genome sequences. Several new features have been added to the database since its last release, including syntenic genome regions for human poly(A) sites in seven other vertebrates and cis-element information adjacent to poly(A) sites. Trace sequences are used to provide additional evidence for poly(A/T) tails in cDNA/ESTs. The updated database is intended to broaden poly(A) site coverage in vertebrate genomes, and provide means to assess the authenticity of poly(A) sites identified by bioinformatics. The URL for this database is http://polya.umdnj.edu/PolyA_DB2.

Animals↗

An active foamy virus integrase is required for virus replication.

Foamy viruses (FVs) make use of a replication strategy which is unique among retroviruses and shows analogies to hepadnaviruses. The presence of an integrase (IN) and obligate provirus integration distinguish retroviruses from hepadnaviruses. To clarify whether a functional IN is required for FV replication, a mutant in the highly conserved DD35E motif of the active centre was analysed. This mutant was found to be able to express Gag and Pol protein precursors and cleavage products and to generate and deliver cDNA. However, this mutant was replication-deficient. The junctions of individual foamy proviruses with cellular DNA were sequenced. The findings suggest that FV integration is asymmetrical, because the proviruses started with what is believed to be the U3 end of the free linear DNA to generate the conventional TG dinucleotide, while apparently two nucleotides from the U5 end were cleaved to create the complementary CA dinucleotide. Alignment of known FV genome sequences indicated that this mechanism of integration is not restricted to the two FV isolates from which integrates were studied, but appears to be a common feature of this retrovirus subfamily. In conclusion, with respect to the necessity of a functionally active IN for virus replication FVs behave like other retroviruses; their mechanism of integration, however, is probably unique.

Animals↗

Databases and software for the comparison of prokaryotic genomes.

The explosion in the number of complete genomes over the past decade has spawned a new and exciting discipline, that of comparative genomics. To exploit the full potential of this approach requires the development of novel algorithms, databases and software which are sophisticated enough to draw meaningful comparisons between complete genome sequences and are widely accessible to the scientific community at large. This article reviews progress towards the development of computational tools and databases for organizing and extracting biological meaning from the comparison of large collections of genomes.

Computational Biology↗

Detection and visualization of compositionally similar cis-regulatory element clusters in orthologous and coordinately controlled genes.

Evolutionarily conserved noncoding genomic sequences represent a potentially rich source for the discovery of gene regulatory regions. However, detecting and visualizing compositionally similar cis-element clusters in the context of conserved sequences is challenging. We have explored potential solutions and developed an algorithm and visualization method that combines the results of conserved sequence analyses (BLASTZ) with those of transcription factor binding site analyses (MatInspector) (http://trafac.chmcc.org). We define hits as the density of co-occurring cis-element transcription factor (TF)-binding sites measured within a 200-bp moving average window through phylogenetically conserved regions. The results are depicted as a Regulogram, in which the hit count is plotted as a function of position within each of the two genomic regions of the aligned orthologs. Within a high-scoring region, the relative arrangement of shared cis-elements within compositionally similar TF-binding site clusters is depicted in a Trafacgram. On the basis of analyses of several training data sets, the approach also allows for the detection of similarities in composition and relative arrangement of cis-element clusters within nonorthologous genes, promoters, and enhancers that exhibit coordinate regulatory properties. Known functional regulatory regions of nonorthologous and less-conserved orthologous genes frequently showed cis-element shuffling, demonstrating that compositional similarity can be more sensitive than sequence similarity. These results show that combining sequence similarity with cis-element compositional similarity provides a powerful aid for the identification of potential control regions.

Animals↗

Sequence assembly with CAFTOOLS.

Large-scale genomic sequencing requires a software infrastructure to support and integrate applications that are not directly compatible. We describe a suite of software tools built around the Common Assembly Format (CAF), a comprehensive representation of a sequence assembly as a text file. These tools form the backbone of sequencing informatics at the Sanger Centre and the Genome Sequencing Center. The CAF format is intentionally flexible, and our Perl and C libraries, which parse and manipulate it, provide powerful tools for creating new applications as well as wrappers to incorporate other software. The tools are available free by anonymous FTP from ftp://ftp.sanger.ac.uk/pub/badger/.

Algorithms↗

Chromatin immunoprecipitation-mediated target identification proved aquaporin 5 is regulated directly by estrogen in the uterus.

Estrogens play a central role in the reproduction of vertebrates and affect a variety of biological processes. The major target molecules of estrogens are nuclear estrogen receptors (ERs), which have been studied extensively at the molecular level. In contrast, our knowledge of the genes that are regulated directly by ERs remains limited, especially at the level of the whole organism rather than cultured cells. In order to identify genes that are regulated directly by ERs in vivo, we used estrogen treated mouse uterus and performed chromatin immunoprecipitation. Sequence analysis of a precipitated DNA fragment enabled alignment with the mouse genomic sequence and revealed that the promoter region of the gene encoding aquaporin 5 (AQP5) was precipitated with antibody against ER alpha. Quantitative PCR and DNA microarray analyses confirmed that AQP5 is activated soon after administration of estrogen. In addition, the promoter region of AQP5 contained a functional estrogen response element that was activated directly by estrogen. Although several AQP genes are expressed in the uterus, only direct activation of AQP5 could be detected following treatment with estrogen. This chromatin immunoprecipitation-mediated target identification may be applicable to the study of other transcription factor networks.

Animals↗

Exhaustive matching of the entire protein sequence database.

The entire protein sequence database has been exhaustively matched. Definitive mutation matrices and models for scoring gaps were obtained from the matching and used to organize the sequence database as sets of evolutionarily connected components. The methods developed are general and can be used to manage sequence data generated by major genome sequencing projects. The alignments made possible by the exhaustive matching are the starting point for successful de novo prediction of the folded structures of proteins, for reconstructing sequences of ancient proteins and metabolisms in ancient organisms, and for obtaining new perspectives in structural biochemistry.

Amino Acid Sequence↗

Aligning awareness, systems and policy to increase equitable access to genomically driven cancer care.

Genomic testing has the potential to transform cancer care across the patient pathway. However, its benefits remain unevenly realised across populations and health systems. Precision oncology is characterised by a strong promissory discourse, with expectations of improved outcomes and cost-effectiveness, yet real-world implementation remains variable and context dependent. This review examines how patient and public awareness interacts with, and is constrained by, structural, organisational and political-economic factors that shape equitable access to genomically driven cancer care across five themes: (1) the power of patient and public advocacy; (2) learning from the patient perspective; (3) culturally responsive communication; (4) structural and personal barriers and facilitators and (5) political economy of health. Examples are mapped across global regions to highlight how health system, structural and societal factors continue to limit the universal realisation of genomic medicine's benefits. We present four recommendations to strengthen the translation of awareness into equitable access and clinical impact: (1) expand and adequately power genomic studies in underserved populations; (2) improve risk communication and decision-making across the cancer pathway; (3) equitable validation and interpretation of emerging genomic technologies and (4) generate real-world evidence on access, uptake and outcomes of genome-matched therapies. Embedding awareness, trust, access and equity in future initiatives is imperative to realise the promise of precision medicine for all patients.

Biomarkers↗

Chicken genomics resource: sequencing and annotation of 35,407 ESTs from single and multiple tissue cDNA libraries and CAP3 assembly of a chicken gene index.

Its accessibility, unique evolutionary position, and recently assembled genome sequence have advanced the chicken to the forefront of comparative genomics and developmental biology research as a model organism. Several chicken expressed sequence tag (EST) projects have placed the chicken in 10th place for accrued ESTs among all organisms in GenBank. We have completed the single-pass 5'-end sequencing of 37,557 chicken cDNA clones from several single and multiple tissue cDNA libraries and have entered 35,407 EST sequences into GenBank. Our chicken EST sequences and those found in public databases (on July 1, 2004) provided a total of 517,727 public chicken ESTs and mRNAs. These sequences were used in the CAP3 assembly of a chicken gene index composed of 40,850 contigs and 79,192 unassembled singlets. The CAP3 contigs show a 96.7% match to the chicken genome sequence. The University of Delaware (UD) EST collection (43,928 clones) was assembled into 19,237 nonredundant sequences (13,495 contigs and 5,742 unassembled singlets). The UD collection contains 6,223 unique sequences that are not found in other public EST collections but show a 76% match to the chicken genome sequence. Our chicken contig and singlet sequences were annotated according to the highest BlastX and/or BlastN hits. The UD CAP3 contig assemblies and singlets are searchable by nucleotide sequence or key word (http://cogburn.dbi.udel.edu), and the cDNA clones are readily available for distribution from the chick EST website and clone repository (http://www.chickest.udel.edu). The present paper describes the construction and normalization of single and multiple tissue chicken cDNA libraries, high-throughput EST sequencing from these libraries, the CAP3 assembly of a chicken gene index from all public ESTs, and the identification of several nonredundant chicken gene sets for production of custom DNA microarrays.

Animals↗

No statistical support for correlation between the positions of protein interaction sites and alternatively spliced regions.

BACKGROUND: Alternative splicing is an efficient mechanism for increasing the variety of functions fulfilled by proteins in a living cell. It has been previously demonstrated that alternatively spliced regions often comprise functionally important and conserved sequence motifs. The objective of this work was to test the hypothesis that alternative splicing is correlated with contact regions of protein-protein interactions. RESULTS: Protein sequence spans involved in contacts with an interaction partner were delineated from atomic structures of transient interaction complexes and juxtaposed with the location of alternatively spliced regions detected by comparative genome analysis and spliced alignment. The total of 42 alternatively spliced isoforms were identified in 21 amino acid chains involved in biomolecular interactions. Using this limited dataset and a variety of sophisticated counting procedures we were not able to establish a statistically significant correlation between the positions of protein interaction sites and alternatively spliced regions. CONCLUSIONS: This finding contradicts a naïve hypothesis that alternatively spliced regions would correlate with points of contact. One possible explanation for that could be that all alternative splicing events change the spatial structure of the interacting domain to a sufficient degree to preclude interaction. This is indirectly supported by the observed lack of difference in the behaviour of relatively short regions affected by alternative splicing and cases when large portions of proteins are removed. More structural data on complexes of interacting proteins, including structures of alternative isoforms, are needed to test this conjecture.

Alternative Splicing↗

Effect of 5'UTR introns on gene expression in Arabidopsis thaliana.

BACKGROUND: The majority of introns in gene transcripts are found within the coding sequences (CDSs). A small but significant fraction of introns are also found to reside within the untranslated regions (5'UTRs and 3'UTRs) of expressed sequences. Alignment of the whole genome and expressed sequence tags (ESTs) of the model plant Arabidopsis thaliana has identified introns residing in both coding and non-coding regions of the genome. RESULTS: A bioinformatic analysis revealed some interesting observations: (1) the density of introns in 5'UTRs is similar to that in CDSs but much higher than that in 3'UTRs; (2) the 5'UTR introns are preferentially located close to the initiating ATG codon; (3) introns in the 5'UTRs are, on average, longer than introns in the CDSs and 3'UTRs; and (4) 5'UTR introns have a different nucleotide composition to that of CDS and 3'UTR introns. Furthermore, we show that the 5'UTR intron of the A. thaliana EF1alpha-A3 gene affects the gene expression and the size of the 5'UTR intron influences the level of gene expression. CONCLUSION: Introns within the 5'UTR show specific features that distinguish them from introns that reside within the coding sequence and the 3'UTR. In the EF1alpha-A3 gene, the presence of a long intron in the 5'UTR is sufficient to enhance gene expression in plants in a size dependent manner.

3' Untranslated Regions↗

Choice, methodology, and characterization of focal ischemic stroke models: the search for clinical relevance.

To develop novel neuroprotective or neurorestorative agents for clinical application, the appropriate selection and characterization of preclinical focal stroke models is required to provide confidence in predicting therapeutic efficacy. Compelling evidence for novel therapies derived from the pathological and functional consequences of models of cerebral ischemia in the rat (and higher species) is an essential prerequisite before large expensive clinical trials are begun. This chapter provides an overview of focal ischemic models, with an emphasis on objective functional assessment of pathological mechanisms and efficacy of novel therapeutic strategies. The ability to predict functional consequences from structural abnormalities is a critical theme that can be extrapolated from the preclinical to the clinical setting, in that certain brain regions are inextricably linked to specific behavioral functions. This underlying approach is highly relevant, as monitoring the dynamic pathological and functional changes attributed to focal stroke will reveal new insights into novel mechanisms and targets that play a role in the evolution of cell death and impaired function. The utility of novel genomic technologies that are aligned with methods to determine structure-function relationships in preclinical models will facilitate a greater understanding of the pathophysiological process and potentially generate new targets that may ultimately be used to predict or offer clinical benefit.

Animals↗