Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Regulatory factor interactions and somatic silencing of the germ cell-specific ALF gene.

Germ cell-specific genes are active in oocytes and spermatocytes but are silent in all other cell types. To understand the basis for this seemingly simple pattern of regulation, we characterized factors that recognize the promoter-proximal region of the germ cell-specific TFIIA alpha/beta-like factor (ALF) gene. Two of the protein-DNA complexes formed with liver extracts (C4 and C5) are due to the zinc finger proteins Sp1 and Sp3, respectively, whereas another complex (C6) is due to the transcription factor RFX1. Two additional complexes (C1 and C3) are due to the multivalent zinc finger protein CTCF, a factor that plays a role in gene silencing and chromatin insulation. An investigation of CTCF binding revealed a recognition site of only 17 bp that overlaps with the Sp1/Sp3 site. This site is predictive of other genomic CTCF sites and can be aligned to create a functional consensus. Studies on the activity of the ALF promoter in somatic 293 cells revealed mutations that result in increased reporter activity. In addition, RNAi-mediated down-regulation of CTCF is associated with activation of the endogenous ALF gene, and both CTCF and Sp3 repress the promoter in transient transfection assays. Overall, the results suggest a role for several factors, including the multivalent zinc finger chromatin insulator protein CTCF, in mediating somatic repression of the ALF gene. Release of such repression, perhaps in conjunction with other members of the CTCF, RFX, and Sp1 families of transcription factors, could be an important aspect of germ cell gene activation.

Animals↗

Chromosomal breakpoint reuse in genome sequence rearrangement.

In order to apply gene-order rearrangement algorithms to the comparison of genome sequences, Pevzner and Tesler bypass gene finding and ortholog identification and use the order of homologous blocks of unannotated sequence as input. The method excludes blocks shorter than a threshold length. Here we investigate possible biases introduced by eliminating short blocks, focusing on the notion of breakpoint reuse introduced by these authors. Analytic and simulation methods show that reuse is very sensitive to the proportion of blocks excluded. As is pertinent to the comparison of mammalian genomes, this exclusion risks randomizing the comparison partially or entirely.

Algorithms↗

Comparative genome assembly.

One of the most complex and computationally intensive tasks of genome sequence analysis is genome assembly. Even today, few centres have the resources, in both software and hardware, to assemble a genome from the thousands or millions of individual sequences generated in a whole-genome shotgun sequencing project. With the rapid growth in the number of sequenced genomes has come an increase in the number of organisms for which two or more closely related species have been sequenced. This has created the possibility of building a comparative genome assembly algorithm, which can assemble a newly sequenced genome by mapping it onto a reference genome. We describe here a novel algorithm for comparative genome assembly that can accurately assemble a typical bacterial genome in less than four minutes on a standard desktop computer. The software is available as part of the open-source AMOS project.

Algorithms↗

Computational space reduction and parallelization of a new clustering approach for large groups of sequences.

MOTIVATION: The explosive growth of the biological sequences databases stimulated by genome projects has modified the framework of several applications in the biological sequence analysis area. In most cases, this new scenario is characterized by studies on large sets of sequences, suggesting the need for effective and automatic methods for their clustering. A more effective clustering of the database could be followed by the application of common family analysis schemes to the groups so formed. RESULTS: In this work, we present a new strategy to reduce the computational cost associated with the clustering of large sets of sequences which are expected to contain several families. The strategy is based on the grouping of the sequences into families by using a dynamic threshold on a pairwise sequence similarity criterion. Routine clustering of large data sets can now be done very efficiently. The method developed here achieves a computational space reduction of about an order of magnitude over more traditional ones of all-versus-all comparisons. The outcome of this approach produces family groupings that reproduce closely already accepted biological results. Our work includes a parallel implementation for distributed memory multiprocessors with a dynamic scheduling strategy for performance optimization. AVAILABILITY: By anonymous ftp at ftp.ac.uma.es (/pub/ots/pCluster directory), or from our Web site http://www.cnb. uam.es/www/software/software_index.html CONTACT: ots@ac.uma.es

Algorithms↗

A case study in genome-level fragment assembly.

MOTIVATION: We use the fact of two teams independently sequencing the one megabase genome of Borrelia burgdorferi as an opportunity to study the accuracy of genome-level assembly. RESULTS: We compare the results of three different assembly programs (PHRAP, TIGR Assembler, and STROLL) on the DNA fragments used in both the Brookhaven and TIGR sequencing projects. We also describe the algorithms and data structures used in our assembly program STROLL, which was used in the Brookhaven Borrelia project.

Algorithms↗

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny↗

Decipher RNA isoform combinations from minigene splicing assays and massive parallel sequencing with MAGIC.

SUMMARY: Functional testing of RNA using minigene splicing assays is increasingly being realized to demonstrate the effects of variants on splicing. In complex cases, variant pathogenicity is assessed by Sanger sequencing, which can be time consuming and may be replaced by short read sequencing. Moreover, strategies based on long read sequencing of the amplified minigene construct are promising and allow the isoforms to be fully characterized. We introduce MAGIC, a user-friendly tool that first generates the artificial construction genome files required to then perform alignment, assembly and annotation of the isoforms obtained by either short or long read minigene splicing assay sequencing. AVAILABILITY AND IMPLEMENTATION: MAGIC is available at https://github.com/LBGC-CFB/MAGIC. Zenodo DOI: 10.5281/zenodo.17052752.

High-Throughput Nucleotide Sequencing↗

PRIMEX: rapid identification of oligonucleotide matches in whole genomes.

SUMMARY: PRIMEX (PRImer Match EXtractor) can detect oligonucleotide sequences in whole genomes, allowing for mismatches. Using a word lookup table and server functionality, PRIMEX accepts queries from client software and returns matches rapidly. We find it faster and more sensitive than currently available tools. AVAILABILITY: Running applications and source code have been made available at http://bioinformatics.cribi.unipd.it/primex

Algorithms↗

Development of joint application strategies for two microbial gene finders.

MOTIVATION: As a starting point in annotation of bacterial genomes, gene finding programs are used for the prediction of functional elements in the DNA sequence. Due to the faster pace and increasing number of genome projects currently underway, it is becoming especially important to have performant methods for this task. RESULTS: This study describes the development of joint application strategies that combine the strengths of two microbial gene finders to improve the overall gene finding performance. Critica is very specific in the detection of similarity-supported genes as it uses a comparative sequence analysis-based approach. Glimmer employs a very sophisticated model of genomic sequence properties and is sensitive also in the detection of organism-specific genes. Based on a data set of 113 microbial genome sequences, we optimized a combined application approach using different parameters with relevance to the gene finding problem. This results in a significant improvement in specificity while there is similarity in sensitivity to Glimmer. The improvement is especially pronounced for GC rich genomes. The method is currently being applied for the annotation of several microbial genomes. AVAILABILITY: The methods described have been implemented within the gene prediction component of the GenDB genome annotation system.

Algorithms↗

A System for Automated Bacterial (genome) Integrated Annotation--SABIA.

UNLABELLED: A web-based software suite, SABIA (System for Automated Bacterial Integrated Annotation), is described that provides a comprehensive computational support for the assembly and annotation of whole bacterial genomes from the data derived from sequencing projects. AVAILABILITY: Both SABIA and supplementary materials are available at http://www.sabia.lncc.br

Algorithms↗

EST clustering error evaluation and correction.

MOTIVATION: The gene expression intensity information conveyed by (EST) Expressed Sequence Tag data can be used to infer important cDNA library properties, such as gene number and expression patterns. However, EST clustering errors, which often lead to greatly inflated estimates of obtained unique genes, have become a major obstacle in the analyses. The EST clustering error structure, the relationship between clustering error and clustering criteria, and possible error correction methods need to be systematically investigated. RESULTS: We identify and quantify two types of EST clustering error, namely, Type I and II in EST clustering using CAP3 assembling program. A Type I error occurs when ESTs from the same gene do not form a cluster whereas a Type II error occurs when ESTs from distinct genes are falsely clustered together. While the Type II error rate is <1.5% for both 5' and 3' EST clustering, the Type I error in the 5' EST case is approximately 10 times higher than the 3' EST case (30% versus 3%). An over-stringent identity rule, e.g., P >/= 95%, may even inflate the Type I error in both cases. We demonstrate that approximately 80% of the Type I error is due to insufficient overlap among sibling ESTs (ISO error) in 5' EST clustering. A novel statistical approach is proposed to correct ISO error to provide more accurate estimates of the true gene cluster profile.

Algorithms↗

Divide-and-conquer approach for the exemplar breakpoint distance.

MOTIVATION: A one-to-one correspondence between the sets of genes in the two genomes being compared is necessary for the notions of breakpoint and reversal distances. To compare genomes where there are paralogous genes, Sankoff formulated the exemplar distance problem as a general version of the genome rearrangement problem. Unfortunately, the problem is NP-hard even for the breakpoint distance. RESULTS: This paper proposes a divide-and-conquer approach for calculating the exemplar breakpoint distance between two genomes with multiple gene families. The combination of our approach and Sankoff's branch-and-bound technique leads to a practical program to answer this question. Tests with both simulated and real datasets show that our program is much more efficient than the existing program that is based only on the branch-and-bound technique. AVAILABILITY: Code for the program is available from the authors.

Algorithms↗

BlastXtract--a new way of exploring translated searches.

SUMMARY: Searches of translated, unannotated genomic DNA sequences against protein databases is a useful early-stage method for discovering protein homologues encoded by the sequence, but generates huge amounts of output data that quickly become impregnable. BlastXtract is a web-based tool for managing and visualizing results from large translated BLAST and FastA searches. It combines the speed and storage benefits of relational database management systems with an easy-to-use graphical navigation map, and greatly facilitates the early exploration of genomic sequence. AVAILABILITY: BlastXtract can be downloaded from http://bioinfo.ucc.ie/blastxtract/.

Amino Acid Sequence↗

GenColors: accelerated comparative analysis and annotation of prokaryotic genomes at various stages of completeness.

SUMMARY: GenColors is a new web-based software/database system aimed at an improved and accelerated annotation of prokaryotic genomes, considering information on related genomes and making extensive use of genome comparison. It offers a seamless integration of data from ongoing sequencing projects and annotated genomic sequences obtained from GenBank. The genome comparison tools determine, for example, best-bidirectional hits, gene conservation, syntenies and gene core sets. Swiss-Prot/TrEMBL hits allow annotations in an effective manner. To further support the annotation base-specific quality data can also be displayed if available. With GenColors dedicated genome browsers containing a group of related genomes can be easily set up and maintained. It has been efficiently used for Borrelia garinii and is currently applied to various ongoing genome projects. AVAILABILITY: Detailed information on GenColors is available at http://gencolors.imb-jena.de. Online usage of GenColors-based genome browsers is the preferred application mode. The system is also available upon request for local installation.

Borrelia↗

Current methods of gene prediction, their strengths and weaknesses.

While the genomes of many organisms have been sequenced over the last few years, transforming such raw sequence data into knowledge remains a hard task. A great number of prediction programs have been developed that try to address one part of this problem, which consists of locating the genes along a genome. This paper reviews the existing approaches to predicting genes in eukaryotic genomes and underlines their intrinsic advantages and limitations. The main mathematical models and computational algorithms adopted are also briefly described and the resulting software classified according to both the method and the type of evidence used. Finally, the several difficulties and pitfalls encountered by the programs are detailed, showing that improvements are needed and that new directions must be considered.

Algorithms↗

ASePCR: alternative splicing electronic RT-PCR in multiple tissues and organs.

RT-PCR is one of the most powerful and direct methods to detect transcript variants due to alternative splicing (AS) that increase transcript diversity significantly in vertebrates. ASePCR is an efficient web-based application that emulates RT-PCR in various tissues. It estimates the amplicon size for a given primer pair based on the transcript models identified by the reverse e-PCR program of the NCBI. The tissue specificity of each PCR band is deduced from the tissue information of expressed sequence tag (EST) sequences compatible with each transcript structure. The output page shows PCR bands like a gel electrophoresis in various tissues. Each band in the output picture represents a putative isoform that could happen in a tissue-specific manner. It also shows the EST alignment and tissue information in the genome browser. Furthermore, the user can compare the AS patterns of orthologous genes in other species. The ASePCR, available at http://genome.ewha.ac.kr/ASePCR/, supports the transcriptome models of the RefSeq, Ensembl, ECgene and AceView for human, mouse, rat and chicken genomes. It will be a valuable web resource to explore the transcriptome diversity associated with different tissues and organs in multiple species.

Algorithms↗

Genome Information Broker for Viruses (GIB-V): database for comparative analysis of virus genomes.

Genome Information Broker for Viruses (GIB-V) is a comprehensive virus genome/segment database. We extracted 18 418 complete virus genomes/segments from the International Nucleotide Sequence Database Collaboration (INSDC, http://www.insdc.org/) by DNA Data Bank of Japan (DDBJ), EMBL and GenBank and stored them in our system. The list of registered viruses is arranged hierarchically according to taxonomy. Keyword searches can be performed for genome/segment data or biological features of any virus stored in GIB-V. GIB-V is equipped with a BLAST search function, and search results are displayed graphically or in list form. Moreover, the BLAST results can be used online with the ClustalW feature of the DDBJ. All available virus genome/segment data can be collected by the GIB-V download function. GIB-V can be accessed at no charge at http://gib-v.genes.nig.ac.jp/.

Databases, Nucleic Acid↗