Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Integrating alternative splicing detection into gene prediction.

BACKGROUND: Alternative splicing (AS) is now considered as a major actor in transcriptome/proteome diversity and it cannot be neglected in the annotation process of a new genome. Despite considerable progresses in term of accuracy in computational gene prediction, the ability to reliably predict AS variants when there is local experimental evidence of it remains an open challenge for gene finders. RESULTS: We have used a new integrative approach that allows to incorporate AS detection into ab initio gene prediction. This method relies on the analysis of genomically aligned transcript sequences (ESTs and/or cDNAs), and has been implemented in the dynamic programming algorithm of the graph-based gene finder EuGENE. Given a genomic sequence and a set of aligned transcripts, this new version identifies the set of transcripts carrying evidence of alternative splicing events, and provides, in addition to the classical optimal gene prediction, alternative optimal predictions (among those which are consistent with the AS events detected). This allows for multiple annotations of a single gene in a way such that each predicted variant is supported by a transcript evidence (but not necessarily with a full-length coverage). CONCLUSIONS: This automatic combination of experimental data analysis and ab initio gene finding offers an ideal integration of alternatively spliced gene prediction inside a single annotation pipeline.

Algorithms↗

Comparative genomics: genome-wide analysis in metazoan eukaryotes.

The increasing number of complete and nearly complete metazoan genome sequences provides a significant amount of material for large-scale comparative genomic analysis. Finding new effective methods to analyse such enormous datasets has been the object of intense research. Three main areas in comparative genomics have recently shown important developments: whole-genome alignment, gene prediction and regulatory-region prediction. Each of these areas improves the methods of deciphering long genomic sequences and uncovering what lies hidden in them.

Animals↗

The G protein-coupled receptor subset of the chicken genome.

G protein-coupled receptors (GPCRs) are one of the largest families of proteins, and here we scan the recently sequenced chicken genome for GPCRs. We use a homology-based approach, utilizing comparisons with all human GPCRs, to detect and verify chicken GPCRs from translated genomic alignments and Genscan predictions. We present 557 manually curated sequences for GPCRs from the chicken genome, of which 455 were previously not annotated. More than 60% of the chicken Genscan gene predictions with a human ortholog needed curation, which drastically changed the average percentage identity between the human-chicken orthologous pairs (from 56.3% to 72.9%). Of the non-olfactory chicken GPCRs, 79% had a one-to-one orthologous relationship to a human GPCR. The Frizzled, Secretin, and subgroups of the Rhodopsin families have high proportions of orthologous pairs, although the percentage of amino acid identity varies. Other groups show large differences, such as the Adhesion family and GPCRs that bind exogenous ligands. The chicken has only three bitter Taste 2 receptors, and it also lacks an ortholog to human TAS1R2 (one of three GPCRs in the human genome in the Taste 1 receptor family [TAS1R]), implying that the chicken's ability and mode of detecting both bitter and sweet taste may differ from the human's. The chicken genome contains at least 229 olfactory receptors, and the majority of these (218) originate from a chicken-specific expansion. To our knowledge, this dataset of chicken GPCRs is the largest curated dataset from a single gene family from a non-mammalian vertebrate. Both the updated human GPCR dataset, as well the chicken GPCR dataset, are available for download.

Animals↗

Analysis of intrastrain recombination in herpes simplex virus type 1 strain 17 and herpes simplex virus type 2 strain HG52 using restriction endonuclease sites as unselected markers and temperature-sensitive lesions as selected markers.

The viral and host factors involved in herpes simplex virus (HSV) recombination are little understood. To identify features of the process, recombination in HSV-1 and HSV-2 has been studied by analysing the segregation of unselected markers in the form of restriction endonuclease (RE) sites. By confining parental interactions to only one strain of virus of each serotype, restrictions imposed by non-homology are overcome and differential growth phenotypes can be discounted. The analysis of unselected and selected recombinants using RE sites in conjunction with temperature-sensitive mutations is consistent with (i) HSV being highly recombinogenic, (ii) parental and progeny molecules taking part in the process, (iii) the four genomic isomers participating in recombination, (iv) genome alignment being part of the recombination process and (v) cellular factors in conjunction with genome homology influencing the efficiency of recombination.

Animals↗

Maximal sequence length of exact match between members from a gene family during early evolution.

Mutation (substitution, deletion, insertion, etc.) in nucleotide acid causes the maximal sequence lengths of exact match (MALE) between paralogous members from a duplicate event to become shorter during evolution. In this work, MALE changes between members of 26 gene families from four representative species (Arabidopsis thaliana, Oryza sativa, Mus musculus and Homo sapiens) were investigated. Comparative study of paralogous' MALE and amino acid substitution rate (d(A)<0.5) indicated that a close relationship existed between them. The results suggested that MALE could be a sound evolutionary scale for the divergent time for paralogous genes during their early evolution. A reference table between MALE and divergent time for the four species was set up, which would be useful widely, for large-scale genome alignment and comparison. As an example, detection of large-scale duplication events of rice genome based on the table was illustrated.

Amino Acid Sequence↗

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms↗

Identification and analysis of cis-regulatory elements in development using comparative genomics with the pufferfish, Fugu rubripes.

The control of vertebrate development is facilitated by cis-regulatory sequences hardwired into the genome. Given that many developmental processes are strikingly similar across all backboned animals, it is reasonable to expect these sequences to be conserved at the nucleotide level, their potential for mutation being constrained by their function. Comparison between the genomes of highly divergent organisms allows such sequences to be identified and some of the most successful approaches have compared regions from the pufferfish, Fugu rubripes, with its distant mammalian relatives, rodents and humans. This review describes progress made in this kind of comparison, from small regions of individual genes, to whole genome alignments.

Animals↗

Comparative analysis of hepatitis C virus phylogenies from coding and non-coding regions: the 5' untranslated region (UTR) fails to classify subtypes.

BACKGROUND: The duration of treatment for HCV infection is partly indicated by the genotype of the virus. For studies of disease transmission, vaccine design, and surveillance for novel variants, subtype-level classification is also needed. This study used the Shimodaira-Hasegawa test and related statistical techniques to compare phylogenetic trees obtained from coding and non-coding regions of a whole-genome alignment for the reliability of subtyping in different regions. RESULTS: Different regions of the HCV genome yield inconsistent phylogenies, which can lead to erroneous conclusions about classification of a given infection. In particular, the highly conserved 5' untranslated region (UTR) yields phylogenetic trees with topologies that differ from the HCV polyprotein and complete genome phylogenies. Phylogenetic trees from the NS5B gene reliably cluster related subtypes, and yield topologies consistent with those of the whole genome and polyprotein. CONCLUSION: These results extend those from previous studies and indicate that, unlike the NS5B gene, the 5' UTR contains insufficient variation to resolve HCV classifications to the level of viral subtype, and fails to distinguish genotypes reliably. Use of the 5' UTR for clinical tests to characterize HCV infection should be replaced by a subtype-informative test.

5' Untranslated Regions↗

Comparative genomic analysis of vertebrate Hox3 and Hox4 genes.

We used a comparative genomic approach to identify putative cis-acting regulatory sequences of the zebrafish hoxb3a and hoxb4a genes. We aligned genomic sequences spanning the clustered Hoxb1 to Hoxb5 genes from pufferfish, mice, and humans with the zebrafish hoxba and hoxbb cluster sequences. We identified multiple blocks of conserved sequences in non-coding regions within and surrounding the Hoxb3/b4 gene locus; a subset of these blocks are conserved in the zebrafish hoxbb cluster, despite loss of hoxb3/b4 genes. Overall, we find that the architecture of the Hoxb3/b4 loci and of the conserved sequence elements is very similar in teleosts and mammals. Our analyses also revealed two alternative transcripts of the zebrafish hoxb3a gene and an exon sequence unusually located 10 kb upstream of adjacent hoxb4a; an equivalent murine Hoxb3 exon has not yet been confirmed. We show that many of the Hoxb3/b4 conserved non-coding sequences correlate with functional neural enhancers previously described in the mouse. Further, within the conserved non-coding sequences we have identified binding sites for transcription factors, including Kreisler/Valentino, Krox20, Hox, and Pbx, some of which had not been previously described for the mouse. Finally, we demonstrate that the regulatory sequences of zebrafish hoxa3a are divergent with respect to the mouse ortholog Hoxa3, or the paralog hoxb3a. Despite limited conservation of regulatory sequences, zebrafish hoxa3a and hoxb3a genes share very similar expression profiles.

Animals↗

Improving quality control of microbial agri-inputs by confirming strain identity with an easy and low-cost PCR-multiplex: A study case with Azospirillum brasilense.

The first commercial product containing the Azospirillum brasilense elite strains Ab-V5 and Ab-V6 was launched in Brazil in 2009. These strains have demonstrated agronomic efficiency in grasses and in legume co-inoculation, accounting for approximately 43 million doses in 2024. Official identification of these strains is currently performed by rep-PCR, a reliable but time-consuming and laborious method. In this study, a multiplex PCR assay was developed for the simultaneous identification of Ab-V5 and Ab-V6 in a single reaction using strain-specific SNPs. Forward primers were designed so that the terminal nucleotide at the 3' end corresponded to a strain-specific SNP unique to each target strain. To further enhance specificity, artificial mismatches were introduced at the fourth nucleotide from the 3' end of the forward primers. SNPs were identified using Snippy based on genomic alignments between Ab-V5 and Ab-V6 and confirmed by local BLASTn against the genomes of other Azospirillum species. In the multiplex assay, simultaneous and specific amplification of both strains was observed in a single reaction, without non-specific amplification. Primer specificity was also experimentally evaluated against other A. brasilense strains (Ab-V1, Ab-V2, Ab-V4, Ab-V7, Ab-V8, and Sp7T), in silico against bacteria from different genera associated with agricultural inoculants, and in commercial inoculant samples containing Ab-V5 and Ab-V6. The results confirmed the high specificity of the primers for Ab-V5 and Ab-V6 and demonstrated that the assay was capable of identifying the strains in commercial inoculants. This assay facilitates inoculant quality control by enabling strain confirmation using a simple, rapid, and low-cost method.

Azospirillum↗

SynBrowse: a synteny browser for comparative sequence analysis.

MOTIVATION: The recent efforts of various sequence projects to sequence deeply into various phylogenies provide great resources for comparative sequence analysis. A generic and portable tool is essential for scientists to visualize and analyze sequence comparisons. RESULTS: We have developed SynBrowse, a synteny browser for visualizing and analyzing genome alignments both within and between species. It is intended to help scientists study macrosynteny, microsynteny and homologous genes between sequences. It can also aid with the identification of uncharacterized genes, putative regulatory elements and novel structural features of a species. SynBrowse is a GBrowse (the Generic Genome Browser) family software tool that runs on top of the open source BioPerl modules. It consists of two components: a web-based front end and a set of relational database back ends. Each database stores pre-computed alignments from a focus sequence to reference sequences in addition to the genome annotations of the focus sequence. The user interface lets end users select a key comparative alignment type and search for syntenic blocks between two sequences and zoom in to view the relationships among the corresponding genome annotations in detail. SynBrowse is portable with simple installation, flexible configuration, convenient data input and easy integration with other components of a model organism system. AVAILABILITY: The software is available at http://www.gmod.org CONTACT: vbrendel@iastate.edu

Algorithms↗

Postgenomic bioinformatic analysis of yeast artificial chromosome sequence.

The free availability of multiple genomic sequences represents one of the greatest advances in biology of the new millennium, and promises to revolutionize our ability to determine and treat the causes of human disease. This chapter highlights a number of basic, freely available, and user-friendly bioinformatic techniques that can be used to predict the functional genetic contents of specific yeast artificial chromosome (YAC) clones. The content of this chapter is written for the level of graduate students, who may be relatively inexperienced with the use of computers for analyzing DNA sequences. The basic instructions that allow the identification of the genomic sequence of interest and to download this sequence onto a personal computer from an online database are presented. Simple instructions are also given on how to perform basic sequence manipulations, how to use online tools to design polymerase chain reaction primers, and how to map restriction sites. Also described are more complicated programs that rapidly and efficiently perform genome alignments that, in addition to predicting the location of protein coding sequences, allow the prediction of functional genomic sequences, such as cis regulatory elements and scaffold/matrix attachment sites. The availability of genomic sequences and the rapidly expanding numbers of predictive programs that allow the predictive analysis of these sequences promises to greatly facilitate the use of YAC clones in the search for the causes of disease.

Chromosomes, Artificial, Yeast↗

A comparative analysis of numt evolution in human and chimpanzee.

Mitochondrial DNA sequences are frequently transferred into the nuclear genome, giving rise to numts (nuclear DNA sequences of mitochondrial origin). So far, the evolutionary history of numts has largely been studied by using single genomes. Here, we present the first attempt to study numt evolution in a comparative manner by using a pairwise genomic alignment. The total number of numts was estimated to be 452 in human and 469 in chimpanzee. numts that were found in both genomes at identical loci were deemed to be orthologous; 391 numts (>80%) were classified as such. The preponderance of orthologous numts is due to the very short divergence time between the 2 hominoids. The rest of numts were deemed to be nonorthologous. Nonorthologous numts were subdivided into 1) ancestral numts that have lost an ortholog in one species through deletion (12 in human and 11 in chimpanzee), 2) new numts acquired by the insertion of a mitochondrial sequence after the divergence of the 2 species (34 in human and 46 in chimpanzee), and 3) paralogous numts created by the tandem duplication of a preexisting numt (2 in human). This approach also enabled us to reconstruct the numt repertoire in the common ancestor of humans and chimpanzees (409 numts). Our comparative approach is also useful in identifying the exact boundaries of numts.

Animals↗

GRIL: genome rearrangement and inversion locator.

UNLABELLED: GRIL is a tool to automatically identify collinear regions in a set of bacterial-size genome sequences. GRIL uses three basic steps. First, regions of high sequence identity are located. Second, some of these regions are filtered based on user-specified criteria. Finally, the remaining regions of sequence identity are used to define significant collinear regions among the sequences. By locating collinear regions of sequence, GRIL provides a basis for multiple genome alignment using current alignment systems. GRIL also provides a basis for using current inversion distance tools to infer phylogeny. AVAILABILITY: GRIL is implemented in C++ and runs on any x86-based Linux or Windows platform. It is available from http://asap.ahabs.wisc.edu/gril

Chromosome Inversion↗

Versatile and open software for comparing large genomes.

The newest version of MUMmer easily handles comparisons of large eukaryotic genomes at varying evolutionary distances, as demonstrated by applications to multiple genomes. Two new graphical viewing tools provide alternative ways to analyze genome alignments. The new system is the first version of MUMmer to be released as open-source software. This allows other developers to contribute to the code base and freely redistribute the code. The MUMmer sources are available at http://www.tigr.org/software/mummer.

Animals↗

REMI-RFLP mapping in the Dictyostelium genome.

A set of 147 Dictyostelium discoideum strains was constructed by random integration of a vector containing rare restriction sites. The strains were generated by transformation using restriction enzyme-mediated integration (REMI) which results in the integration of linear DNA fragments into randomly distributed genomic restriction sites. Restriction fragment length polymorphism (RFLP) was generated in a single genomic site in each strain. These REMI-RFLP strains were used to confirm gene linkages previously supported by two other physical mapping techniques: yeast artificial chromosome (YAC) contig construction, and megabase-scale restriction mapping. New linkages were uncovered when two or more hybridization probes identified the same RFLP fragments. Probes for 100 genes have marked 53% of the RFLPs, representing greater than 22 Mb of the 40 Mb Dictyostelium genome. Alignment of these and other large fragments along each chromosome should lead to a complete physical map of the Dictyostelium genome.

Animals↗

Genomic features in the breakpoint regions between syntenic blocks.

MOTIVATION: We study the largely unaligned regions between the syntenic blocks conserved in humans and mice, based on data extracted from the UCSC genome browser. These regions contain evolutionary breakpoints caused by inversion, translocation and other processes. RESULTS: We suggest explanations for the limited amount of genomic alignment in the neighbourhoods of breakpoints. We discount inferences of extensive breakpoint reuse as artefacts introduced during the reconstruction of syntenic blocks. We find that the number, size and distribution of small aligned fragments in the breakpoint regions depend on the origin of the neighbouring blocks and the other blocks on the same chromosome. We account for this and for the generalized loss of alignment in the regions partially by artefacts due to alignment protocols and partially by mutational processes operative only after the rearrangement event. These results are consistent with breakpoints occurring randomly over virtually the entire genome.

Algorithms↗