Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Evolution of vertebrate genes related to prion and Shadoo proteins--clues from comparative genomic analysis.

Recent findings of new genes in fish related to the prion protein (PrP) gene PRNP, including our recent report of SPRN coding for Shadoo (Sho) protein found also in mammals, raise issues of their function and evolution. Here we report additional novel fish genes found in public databases, including a duplicated SPRN gene, SPRNB, in Fugu, Tetraodon, carp, and zebrafish encoding the Sho2 protein, and we use comparative genomic analysis to analyze the evolutionary relationships and to infer evolutionary trajectories of the complete data set. Phylogenetic footprinting performed on aligned human, mouse, and Fugu SPRN genes to define candidate regulatory promoter regions, detected 16 conserved motifs, three of which are known transcription factor-binding sites for a receptor and transcription factors specific to or associated with expression in brain. This result and other homology-based (VISTA global genomic alignment; protein sequence alignment and phylogenetics) and context-dependent (genomic context; relative gene order and orientation) criteria indicate fish and mammalian SPRN genes are orthologous and suggest a strongly conserved basic function in brain. Whereas tetrapod PRNPs share context with the analogous stPrP-2-coding gene in fish, their sequences are diverged, suggesting that the tetrapod and fish genes are likely to have significantly different functions. Phylogenetic analysis predicts the SPRN/SPRNB duplication occurred before divergence of fish from tetrapods, whereas that of stPrP-1 and stPrP-2 occurred in fish. Whereas Sho appears to have a conserved function in vertebrate brain, PrP seems to have an adaptive role fine-tuned in a lineage-specific fashion. An evolutionary model consistent with our findings and literature knowledge is proposed that has an ancestral prevertebrate SPRN-like gene leading to all vertebrate PrP-related and Sho-related genes. This provides a new framework for exploring the evolution of this unusual family of proteins and for searching for members in other fish branches and intermediate vertebrate groups.

Animals↗

SinicView: a visualization environment for comparisons of multiple nucleotide sequence alignment tools.

BACKGROUND: Deluged by the rate and complexity of completed genomic sequences, the need to align longer sequences becomes more urgent, and many more tools have thus been developed. In the initial stage of genomic sequence analysis, a biologist is usually faced with the questions of how to choose the best tool to align sequences of interest and how to analyze and visualize the alignment results, and then with the question of whether poorly aligned regions produced by the tool are indeed not homologous or are just results due to inappropriate alignment tools or scoring systems used. Although several systematic evaluations of multiple sequence alignment (MSA) programs have been proposed, they may not provide a standard-bearer for most biologists because those poorly aligned regions in these evaluations are never discussed. Thus, a tool that allows cross comparison of the alignment results obtained by different tools simultaneously could help a biologist evaluate their correctness and accuracy. RESULTS: In this paper, we present a versatile alignment visualization system, called SinicView, (for Sequence-aligning INnovative and Interactive Comparison VIEWer), which allows the user to efficiently compare and evaluate assorted nucleotide alignment results obtained by different tools. SinicView calculates similarity of the alignment outputs under a fixed window using the sum-of-pairs method and provides scoring profiles of each set of aligned sequences. The user can visually compare alignment results either in graphic scoring profiles or in plain text format of the aligned nucleotides along with the annotations information. We illustrate the capabilities of our visualization system by comparing alignment results obtained by MLAGAN, MAVID, and MULTIZ, respectively. CONCLUSION: With SinicView, users can use their own data sequences to compare various alignment tools or scoring systems and select the most suitable one to perform alignment in the initial stage of sequence analysis.

Algorithms↗

Gene prediction and verification in a compact genome with numerous small introns.

The genomes of clusters of related eukaryotes are now being sequenced at an increasing rate, creating a need for accurate, low-cost annotation of exon-intron structures. In this paper, we demonstrate that reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on predicted gene structures satisfy this need, at least for single-celled eukaryotes. The TWINSCAN gene prediction algorithm was adapted for the fungal pathogen Cryptococcus neoformans by using a precise model of intron lengths in combination with ungapped alignments between the genome sequences of the two closely related Cryptococcus varieties. This approach resulted in approximately 60% of known genes being predicted exactly right at every coding base and splice site. When previously unannotated TWINSCAN predictions were tested by RT-PCR and direct sequencing, 75% of targets spanning two predicted introns were amplified and produced high-quality sequence. When targets spanning the complete predicted open reading frame were tested, 72% of them amplified and produced high-quality sequence. We conclude that sequencing a small number of expressed sequence tags (ESTs) to provide training data, running TWINSCAN on an entire genome, and then performing RT-PCR and direct sequencing on all of its predictions would be a cost-effective method for obtaining an experimentally verified genome annotation.

Algorithms↗

Alternative transcripts of rat slc19a1: cloning, genomic organisation, tissue specific promoters and alternative splicing.

Recently, the rat genome project revealed the genomic sequence of slc19a1, coding for the methotrexate carrier-1, identical to the reduced folate carrier-1 of humans, on rat chromosome 20. At the same time, we have cloned and analysed the complete or partial cDNAs of now at least six different transcripts from rat liver and kidneys. Alignment with the genomic sequence revealed seven exons. The first two non-coding exons, exon I and Ia were used alternatively in kidneys and liver, respectively, suggesting usage of alternative promoters. Three minor mRNA forms resulted from absent splicing of intron III, a shortened exon III (exon IIIa), and a shortened exon IV (exon IVa). The minor transcripts were predicted to result in translation products with 7 or 6 instead of 12 transmembrane domains (TMDs) and a peptide mass of 38, 39 and 40 kDa instead of 58 kDa.

Alternative Splicing↗

A comparison of the genomes of capripoxvirus isolates of sheep, goats, and cattle.

HindIII, PstI, AvaI, and SalI sites were mapped on the genomes of six isolates of capripoxvirus from sheep, goats, and cattle. Genome pairs were aligned by the alignment of cross-hybridizing HindIII fragments and the introduction of padding fragments (pads) at specific locations. The majority of these pads represent the approximate positions of relative deletions or insertions. Three possible phylogenetic networks were generated for seven capripoxvirus isolates by Wagner parsimony analysis of the nonconserved HindIII sites on their genomes, and confidence limits were calculated for the network nodes. Nucleotide sequence divergence values, calculated from the numbers of nonconserved HindIII, PstI, AvaI, and SalI sites on the genomes of typical sheep, goat, and cattle isolates, indicated that the typical sheep and cattle isolates are more closely related to one another than to the typical goat isolate. Nonconserved HindIII, PstI, AvaI, and SalI sites were shown to be distributed throughout the genomes. Evidence that isolates YG-1 and OS-1 are descended from an isolate whose genome arose by recombination is discussed. Terminally repeated regions were identified on each of the capripoxvirus genomes mapped here. By mapping BamHI, ClaI, EcoRI, HindII, HindIII, and SalI sites present within the terminal 10 kb of the genome of isolate InS-1, the terminal repeats of this genome were shown to be between 2.25 and 3.40 kb in length and inverted with respect to one another.

Animals↗

Four basic symmetry types in the universal 7-cluster structure of microbial genomic sequences.

Coding information is the main source of heterogeneity (non-randomness) in the sequences of microbial genomes. The heterogeneity corresponds to a cluster structure in triplet distributions of relatively short genomic fragments (200-400 bp). We found a universal 7-cluster structure in microbial genomic sequences and explained its properties. We show that codon usage of bacterial genomes is a multi-linear function of their genomic G+C-content with high accuracy. Based on the analysis of 143 completely sequenced bacterial genomes available in Genbank in August 2004, we show that there are four "pure" types of the 7-cluster structure observed. All 143 cluster animated 3D-scatters are collected in a database which is made available on our web-site (http://www.ihes.fr/~zinovyev/7clusters). The findings can be readily introduced into software for gene prediction, sequence alignment or microbial genomes classification.

Codon↗

ChimerDB--a knowledgebase for fusion sequences.

Chromosome translocation and gene fusion are frequent events in the human genome and are often the cause of many types of tumor. ChimerDB is the database of fusion sequences encompassing bioinformatics analysis of mRNA and expressed sequence tag (EST) sequences in the GenBank, manual collection of literature data and integration with other known database such as OMIM. Our bioinformatics analysis identifies the fusion transcripts that have non-overlapping alignments at multiple genomic loci. Fusion events at exon-exon borders are selected to filter out the cloning artifacts in cDNA library preparation. The result is classified into two groups--genuine chromosome translocation and fusion between neighboring genes owing to intergenic splicing. We also integrated manually collected literature and OMIM data for chromosome translocation as an aid to assess the validity of each fusion event. The database is available at http://genome.ewha.ac.kr/ChimerDB/ for human, mouse and rat genomes.

Animals↗

Characterization of the porcine alpha interferon multigene family.

The availability of data on the pig genome sequence prompted us to characterize the porcine IFN-alpha (PoIFN-alpha) multigene family. Fourteen functional PoIFN-alpha genes and two PoIFN-alpha pseudogenes were detected in the porcine genome. Multiple sequence alignment revealed a C-terminal deletion of eight residues in six subtypes. A phylogenetic tree of the porcine IFN-alpha gene family defined the evolutionary relationship of the various subtypes. In addition, analysis of the evolutionary rate and the effect of positive selection suggested that the C-terminal deletion is a strategy for preservation in the genome. Eight PoIFN-alpha subtypes were isolated from the porcine liver genome and expressed in BHK-21 cells line. We detected the level of transcription by real-time quantitative RT-PCR analysis. The antiviral activities of the products were determined by WISH cells/Vesicular Stomatitis Virus (VSV) and PK 15 cells/Pseudorabies Virus (PRV) respectively. We found the antiviral activities of intact PoIFN-alpha genes are approximately 2-50 times higher than those of the subtypes with C-terminal deletions in WISH cells and 15-55 times higher in PK 15 cells. There was no obvious difference between the subtypes with and without C-terminal deletion on acid susceptibility.

Amino Acid Sequence↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗

An initial strategy for the systematic identification of functional elements in the human genome by low-redundancy comparative sequencing.

With the recent completion of a high-quality sequence of the human genome, the challenge is now to understand the functional elements that it encodes. Comparative genomic analysis offers a powerful approach for finding such elements by identifying sequences that have been highly conserved during evolution. Here, we propose an initial strategy for detecting such regions by generating low-redundancy sequence from a collection of 16 eutherian mammals, beyond the 7 for which genome sequence data are already available. We show that such sequence can be accurately aligned to the human genome and used to identify most of the highly conserved regions. Although not a long-term substitute for generating high-quality genomic sequences from many mammalian species, this strategy represents a practical initial approach for rapidly annotating the most evolutionarily conserved sequences in the human genome, providing a key resource for the systematic study of human genome function.

Animals↗

Fast and systematic genome-wide discovery of conserved regulatory elements using a non-alignment based approach.

We describe a powerful new approach for discovering globally conserved regulatory elements between two genomes. The method is fast, simple and comprehensive, without requiring alignments. Its application to pairs of yeasts, worms, flies and mammals yields a large number of known and novel putative regulatory elements. Many of these are validated by independent biological observations, have spatial and/or orientation biases, are co-conserved with other elements and show surprising conservation across large phylogenetic distances.

Animals↗

Utilization of long-read sequencing for the detection of structural rearrangements with AgileStructure.

MOTIVATION: Changes in genome organisation contribute to genetic disease when they disrupt gene function or regulation. Structural rearrangements may interrupt coding sequence or alter expression through promoter loss or gain, chromatin changes, copy-number variation, or disruption of short-range regulatory elements. Although short-read sequencing excels at detecting small variants, it performs poorly at resolving breakpoints of large rearrangements, especially in repetitive or low-complexity regions. Long-read sequencing overcomes these limitations, but analytical tools have not kept pace, making accurate identification and annotation of large structural variants challenging. RESULTS: We developed AgileStructure, a desktop application for locating and annotating large‑scale genomic rearrangements using aligned long‑read data. The software enables user‑guided exploration of breakpoint‑spanning reads, supporting accurate interpretation of complex events and filling a key gap in current structural variant analysis workflows. AVAILABILITY AND IMPLEMENTATION: Source code, binaries, user guide, and example aligned read data, are available on GitHub: https://github.com/msjimc/AgileStructure. An archived version is also available on Zenodo at https://doi.org/10.5281/zenodo.18610110.

Software↗

Cross-species overgo hybridization and comparative physical mapping within avian genomes.

The chicken genome sequence facilitates comparative genomics within other avian species. We performed cross-species hybridizations using overgo probes designed from chicken genomic and zebra finch expressed sequence tags (ESTs) to turkey and zebra finch BAC libraries. As a result, 3772 turkey BACs were assigned to 336 markers or genes, and 1662 zebra finch BACs were assigned to 164 genes. As expected, cross-hybridization was more successful with overgos within coding sequences than within untranslated region, intron or flanking sequences and between chicken and turkey, when compared with chicken-zebra finch or zebra finch-turkey cross-hybridization. These data contribute to the comparative alignment of avian genome maps using a 'one sequence, multiple genomes' strategy.

Animals↗

The NEIBank project for ocular genomics: data-mining gene expression in human and rodent eye tissues.

NEIBank is a project to gather and organize genomic resources for eye research. The first phase of this project covers the construction and sequence analysis of cDNA libraries from human and animal model eye tissues to develop an overview of the repertoire of genes expressed in the eye and a resource of cDNA clones for further studies. The sequence data are grouped and identified using the tools of bioinformatics and the results are displayed through a web site where they can be interrogated by keyword search, chromosome location, by Blast (sequence comparison) or by alignment on completed genomes. Many novel proteins and novel splice forms of known genes have already emerged from analysis of the accumulating data. This review provides an overview of the current state of the database for human eye tissues, with specific comparisons to some parallel data from mouse and rat, and with illustrative examples of the kinds of insights and discoveries these data can produce. One of the major themes that emerges is that at the molecular level human eye tissues have significant differences from those of rodents, encompassing species specific genes, alternative splice forms and great variation in levels of gene expression. These point to specific adaptations and mechanisms in the human eye and emphasize that care needs to be taken in the application of appropriate animal model systems.

Amino Acid Sequence↗

A recombinational event in the history of luteoviruses probably induced by base-pairing between the genomes of two distinct viruses.

Alignments of luteovirus readthrough protein amino acid sequences show they consist of two distinct regions, here named the N domain and the C domain. N domain sequences were classified, and comparison of this gene phylogeny to phylogenies of other luteovirus genes revealed an anomaly in the relationships between beet western yellows luteovirus, cucurbit aphidborne yellows luteovirus (CABYV), and pea enation mosaic RNA1 (PEMV1). Together with alignments of virion protein and readthrough protein amino acid sequences, these gene phylogenies indicate the anomaly to be the result of two recombinational events, probably between ancestors of CABYV and PEMV1 and leading to the transfer of RNA coding for the N domain to an ancestor of CABYV. Two likely recombination sites were identified from the alignments, one at the 5' end of the readthrough protein gene and the other at the 5' end of the sequence coding for the C domain. Alignments of the nucleotide sequences encompassing the probable recombination sites suggest that base-pairing between the genomes of the two ancestral luteoviruses, resulting from local sequence similarity at the 5' end of the readthrough protein gene, probably induced one of the interspecies recombinational events.

Amino Acid Sequence↗

Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.

We have conducted a comprehensive search for conserved elements in vertebrate genomes, using genome-wide multiple alignments of five vertebrate species (human, mouse, rat, chicken, and Fugu rubripes). Parallel searches have been performed with multiple alignments of four insect species (three species of Drosophila and Anopheles gambiae), two species of Caenorhabditis, and seven species of Saccharomyces. Conserved elements were identified with a computer program called phastCons, which is based on a two-state phylogenetic hidden Markov model (phylo-HMM). PhastCons works by fitting a phylo-HMM to the data by maximum likelihood, subject to constraints designed to calibrate the model across species groups, and then predicting conserved elements based on this model. The predicted elements cover roughly 3%-8% of the human genome (depending on the details of the calibration procedure) and substantially higher fractions of the more compact Drosophila melanogaster (37%-53%), Caenorhabditis elegans (18%-37%), and Saccharaomyces cerevisiae (47%-68%) genomes. From yeasts to vertebrates, in order of increasing genome size and general biological complexity, increasing fractions of conserved bases are found to lie outside of the exons of known protein-coding genes. In all groups, the most highly conserved elements (HCEs), by log-odds score, are hundreds or thousands of bases long. These elements share certain properties with ultraconserved elements, but they tend to be longer and less perfectly conserved, and they overlap genes of somewhat different functional categories. In vertebrates, HCEs are associated with the 3' UTRs of regulatory genes, stable gene deserts, and megabase-sized regions rich in moderately conserved noncoding sequences. Noncoding HCEs also show strong statistical evidence of an enrichment for RNA secondary structure.

3' Untranslated Regions↗

The NS5 gene location of two turkey meningoencephalitis virus genomic sequences.

Two new turkey meningoencephalitis virus (TMEV) nucleotide sequences were aligned to complete sequences of genomes of the flaviviruses that were available at present in the GeneBank. It was found that the both TMEV sequences represent different NS5 locations; the sequence with Acc. No. AF098456 is located downstream of that with Acc. No. AF013377 on the TMEV NS5 gene. This finding provides further insight into the TMEV NS5 gene structure and shows that the two sequences are located on the NS5 gene separately.

Animals↗

Human-ovine comparative sequencing of a 250-kb imprinted domain encompassing the callipyge (clpg) locus and identification of six imprinted transcripts: DLK1, DAT, GTL2, PEG11, antiPEG11, and MEG8.

Two ovine BAC clones and a connecting long-range PCR product, jointly spanning approximately 250 kb and representing most of the MULGE5-OY3 marker interval known to contain the clpg locus, were completely sequenced. The resulting genomic sequence was aligned with its human ortholog and extensively annotated. Six transcripts, four of which were novel, were predicted to originate from within the analyzed region and their existence confirmed experimentally: DLK1, DAT, GTL2, PEG11, antiPEG11, and MEG8. RT-PCR experiments performed on a range of tissues sampled from an 8-wk-old animal demonstrated the preferential expression of all six transcripts in skeletal muscle, which suggests that they are under control of common regulatory elements. The six transcripts were also shown to be subject to parental imprinting: DLK1, DAT, and PEG11 were shown to be paternally expressed and GTL2, antiPEG11, and MEG8 to be maternally expressed.

Animals↗