Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Leveraging human genomic information to identify nonhuman primate sequences for expression array development.

BACKGROUND: Nonhuman primates (NHPs) are essential for biomedical research due to their similarities to humans. The utility of NHPs will be greatly increased by the application of genomics-based approaches such as gene expression profiling. Sequence information from the 3' end of genes is the key resource needed to create oligonucleotide expression arrays. RESULTS: We have developed the algorithms and procedures necessary to quickly acquire sequence information from the 3' end of nonhuman primate orthologs of human genes. To accomplish this, we identified terminal exons of over 15,000 human genes by aligning mRNA sequences with genomic sequence. We found the mean length of complete last exons to be approximately 1,400 bp, significantly longer than previous estimates. We designed primers to amplify genomic DNA, which included at least 300 bp of the terminal exon. We cloned and sequenced the PCR products representing over 5,500 Macaca mulatta (rhesus monkey) orthologs of human genes. This sequence information has been used to select probes for rhesus gene expression profiling. We have also tested 10 sets of primers with genomic DNA from Macaca fascicularis (Cynomolgus monkey), Papio hamadryas (Baboon), and Chlorocebus aethiops (African green monkey, vervet). The results indicate that the primers developed for this study will be useful for acquiring sequence from the 3' end of genes for other nonhuman primate species. CONCLUSION: This study demonstrates that human genomic DNA sequence can be leveraged to obtain sequence from the 3' end of NHP orthologs and that this sequence can then be used to generate NHP oligonucleotide microarrays. Affymetrix and Agilent used sequences obtained with this approach in the design of their rhesus macaque oligonucleotide microarrays.

Algorithms↗

Computational identification of operons in microbial genomes.

By applying graph representations to biochemical pathways, a new computational pipeline is proposed to find potential operons in microbial genomes. The algorithm relies on the fact that enzyme genes in operons tend to catalyze successive reactions in metabolic pathways. We applied this algorithm to 42 microbial genomes to identify putative operon structures. The predicted operons from Escherichia coli were compared with a selected metabolism-related operon dataset from the RegulonDB database, yielding a prediction sensitivity (89%) and specificity (87%) relative to this dataset. Several examples of detected operons are given and analyzed. Modular gene cluster transfer and operon fusion are observed. A further use of predicted operon data to assign function to putative genes was suggested and, as an example, a previous putative gene (MJ1604) from Methanococcus jannaschii is now annotated as a phosphofructokinase, which was regarded previously as a missing enzyme in this organism. GC content changes in the operon region and nonoperon region were examined. The results reveal a clear GC content transition at the boundaries of putative operons. We looked further into the conservation of operons across genomes. A trp operon alignment is analyzed in depth to show gene loss and rearrangement in different organisms during operon evolution.

Algorithms↗

[Initial analysis of complete genome sequences of SARS coronavirus].

Multiple sequence alignment among 12 complete SARS coronavirus (SARS-CoV) sequences reveals that the major parts of 29708 b of the genomes have 99.82% identical bases. Forty two nucleotide mismatches were found in addition to the five and six gaps in two genomes. Among them, 28 mismatches result in changes of amino acid in the encoded proteins. Analysis of the changes implies possible effect on the Spike and Membrane protein of the virus, while most of the other changes seem not very significant to alter the structure and function of the proteins. These results have been released on the anti-sars web site maintained by the Centre of Bioinformatics, Peking University (antisars.cbi.pku.edu.cn) and may be of help for further experimental study.

Amino Acid Sequence↗

Phylogenetic shadowing and computational identification of human microRNA genes.

We sequenced 122 miRNAs in 10 primate species to reveal conservation characteristics of miRNA genes. Strong conservation is observed in stems of miRNA hairpins and increased variation in loop sequences. Interestingly, a striking drop in conservation was found for sequences immediately flanking the miRNA hairpins. This characteristic profile was employed to predict novel miRNAs using cross-species comparisons. Nine hundred and seventy-six candidate miRNAs were identified by scanning whole-genome human/mouse and human/rat alignments. Most of the novel candidates are conserved also in other vertebrates (dog, cow, chicken, opossum, zebrafish). Northern blot analysis confirmed the expression of mature miRNAs for 16 out of 69 representative candidates. Additional support for the expression of 179 novel candidates can be found in public databases, their presence in gene clusters, and literature that appeared after these predictions were made. Taken together, these results suggest the presence of significantly higher numbers of miRNAs in the human genome than previously estimated.

Computational Biology↗

Evidence for symmetric chromosomal inversions around the replication origin in bacteria.

BACKGROUND: Whole-genome comparisons can provide great insight into many aspects of biology. Until recently, however, comparisons were mainly possible only between distantly related species. Complete genome sequences are now becoming available from multiple sets of closely related strains or species. RESULTS: By comparing the recently completed genome sequences of Vibrio cholerae, Streptococcus pneumoniae and Mycobacterium tuberculosis to those of closely related species - Escherichia coli, Streptococcus pyogenes and Mycobacterium leprae, respectively - we have identified an unusual and previously unobserved feature of bacterial genome structure. Scatterplots of the conserved sequences (both DNA and protein) between each pair of species produce a distinct X-shaped pattern, which we call an X-alignment. The key feature of these alignments is that they have symmetry around the replication origin and terminus; that is, the distance of a particular conserved feature (DNA or protein) from the replication origin (or terminus) is conserved between closely related pairs of species. Statistically significant X-alignments are also found within some genomes, indicating that there is symmetry about the replication origin for paralogous features as well. CONCLUSIONS: The most likely mechanism of generation of X-alignments involves large chromosomal inversions that reverse the genomic sequence symmetrically around the origin of replication. The finding of these X-alignments between many pairs of species suggests that chromosomal inversions around the origin are a common feature of bacterial genome evolution.

Bacteria↗

Intron-flanking EST-PCR markers: from genetic marker development to gene structure analysis in Rhododendron.

With a long-term goal of constructing a linkage map of Rhododendron enriched with gene-specific markers, we utilized Rhododendron catawbiense ESTs for the development of high-efficiency (in terms of generating polymorphism frequency) PCR-based markers. Using the gene-sequence alignment between Rhododendron ESTs and the genomic sequences of Arabidopsis homologs, we developed 'intron-flanking' EST-PCR-based primers that would anneal in conserved exon regions and amplify across the more highly diverged introns. These primers resulted in increased efficiency (61% vs. 13%; 4.7-fold) of polymorphism-detection compared with conventional EST-PCR methods, supporting the assumption that intron regions are more diverged than exons. Significantly, this study demonstrates that Arabidopsis genome database can be useful in developing gene-specific PCR-based markers for other non-model plant species for which the EST data are available but genomic sequences are not. The comparative analysis of intron sizes between Rhododendron and Arabidopsis (made possible in this study by aligning of Rhododendron ESTs with Arabidopsis genomic sequences and the sequencing of Rhododendron genomic PCR products) provides the first insight into the gene structure of Rhododendron.

Arabidopsis↗

SyMAP: A system for discovering and viewing syntenic regions of FPC maps.

Previous approaches to comparing gene and chromosome organization between two genomes have been based on genetic maps or genomic sequences. We have developed a system to align an FPC-based physical map to a genomic sequence based on BAC end sequences and sequence-tagged hybridization markers and to align two FPC maps to one another based on shared markers and fingerprints. The system, called SyMAP (Synteny Mapping and Analysis Program), consists of an algorithm to compute synteny blocks and Web-based graphics to visualize the results. The approach to calculating the anchors (corresponding elements on the respective maps) maximizes the inclusion of anchors with different rates of divergence. Chains (putative syntenic sets of anchors) are computed using a dynamic programming algorithm, which includes off-diagonal anchors that result from map coordinate errors and small inversions. As the gap parameters (the distances allowed between anchors in a chain) can vary over different data sets and be difficult to set manually, they are automatically computed per data set. The criterion for a chain to be acceptable is based on the number of anchors and the Pearson correlation coefficient. Neighboring chains are merged into synteny blocks for display. This algorithm has been tested with three data sets that vary in the number of BACs, BAC end sequences, hybridization markers, distance between anchors, and number and antiquity of genome duplication events. The Web-based graphics uses Java for a highly interactive display that allows the user to interrogate the evidence of synteny.

Algorithms↗

Detection of chromosomal regions showing differential gene expression in human skeletal muscle and in alveolar rhabdomyosarcoma.

BACKGROUND: Rhabdomyosarcoma is a relatively common tumour of the soft tissue, probably due to regulatory disruption of growth and differentiation of skeletal muscle stem cells. Identification of genes differentially expressed in normal skeletal muscle and in rhabdomyosarcoma may help in understanding mechanisms of tumour development, in discovering diagnostic and prognostic markers and in identifying novel targets for drug therapy. RESULTS: A Perl-code web client was developed to automatically obtain genome map positions of large sets of genes. The software, based on automatic search on Human Genome Browser by sequence alignment, only requires availability of a single transcribed sequence for each gene. In this way, we obtained tissue-specific chromosomal maps of genes expressed in rhabdomyosarcoma or skeletal muscle. Subsequently, Perl software was developed to calculate gene density along chromosomes, by using a sliding window. Thirty-three chromosomal regions harbouring genes mostly expressed in rhabdomyosarcoma were identified. Similarly, 48 chromosomal regions were detected including genes possibly related to function of differentiated skeletal muscle, but silenced in rhabdomyosarcoma. CONCLUSION: In this study we developed a method and the associated software for the comparative analysis of genomic expression in tissues and we identified chromosomal segments showing differential gene expression in human skeletal muscle and in alveolar rhabdomyosarcoma, appearing as candidate regions for harbouring genes involved in origin of alveolar rhabdomyosarcoma representing possible targets for drug treatment and/or development of tumor markers.

Chromosome Mapping↗

ThurGood: evaluating assembly-to-assembly mapping.

The alignment and mapping of large genomic sequences is the focus of much recent research. However, relatively little has been done so far about testing and validating alignment methods. We introduce criteria and new tools we have developed for alignment evaluation. These tools have already proved useful in the evaluation and ranking of several methods for assembly-to-assembly mapping, which were recently used to map multiple versions of the human genome to each other (Istrail et aL, 2004).

Algorithms↗

igv-reports: embedding interactive genomic visualizations in HTML reports to aid variant review.

SUMMARY: We present igv-reports, a command-line tool to create standalone HTML pages embedding interactive genomic visualizations of read alignments and associated annotations to support variant inspection workflows. The reports contain all data and code required for visualization of the variant sites, with no dependencies on the input data files. AVAILABILITY AND IMPLEMENTATION: igv-reports is a command-line application written in Python. It is freely available at https://github.com/igvteam/igv-reports under an MIT license.

Software↗

Towards a reliable objective function for multiple sequence alignments.

Multiple sequence alignment is a fundamental tool in a number of different domains in modern molecular biology, including functional and evolutionary studies of a protein family. Multiple alignments also play an essential role in the new integrated systems for genome annotation and analysis. Thus, the development of new multiple alignment scores and statistics is essential, in the spirit of the work dedicated to the evaluation of pairwise sequence alignments for database searching techniques. We present here norMD, a new objective scoring function for multiple sequence alignments. NorMD combines the advantages of the column-scoring techniques with the sensitivity of methods incorporating residue similarity scores. In addition, norMD incorporates ab initio sequence information, such as the number, length and similarity of the sequences to be aligned. The sensitivity and reliability of the norMD objective function is demonstrated using structural alignments in the SCOP and BAliBASE databases. The norMD scores are then applied to the multiple alignments of the complete sequences (MACS) detected by BlastP with E-value<10, for a set of 734 hypothetical proteins encoded by the Vibrio cholerae genome. Unrelated or badly aligned sequences were automatically removed from the MACS, leaving a high-quality multiple alignment which could be reliably exploited in a subsequent functional and/or structural annotation process. After removal of unreliable sequences, 176 (24 %) of the alignments contained at least one sequence with a functional annotation. 103 of these new matches were supported by significant hits to the Interpro domain and motif database.

Amino Acid Motifs↗

New bioinformatics tools for viral genome analyses at Viral Bioinformatics--Canada.

Viruses are much smaller than prokaryotes and eukaryotes, and it is now practical to sequence closely related members of virus families, strains, or even different isolates recovered during the course of an outbreak. However, comparative analysis of viral genomes requires the development of novel bioinformatics tools that allow us to align, edit, compare and interact with these genomes at all levels, from whole genome, to gene family, to single nucleotide polymorphisms. Comparative viral genomics can lead to the identification of the core characteristics that define a virus family, as well as the unique properties of viral species or isolates that contribute to variations in pathogenesis. This paper describes a number of tools, mainly developed for Viral Bioinformatics--Canada, that can be used for annotation and comparative genomic analysis of poxviruses. Nonetheless, these tools are also broadly applicable to other virus families.

Base Sequence↗

The dog genome: survey sequencing and comparative analysis.

A survey of the dog genome sequence (6.22 million sequence reads; 1.5x coverage) demonstrates the power of sample sequencing for comparative analysis of mammalian genomes and the generation of species-specific resources. More than 650 million base pairs (>25%) of dog sequence align uniquely to the human genome, including fragments of putative orthologs for 18,473 of 24,567 annotated human genes. Mutation rates, conserved synteny, repeat content, and phylogeny can be compared among human, mouse, and dog. A variety of polymorphic elements are identified that will be valuable for mapping the genetic basis of diseases and traits in the dog.

Animals↗

Identification of the linker-SH2 domain of STAT as the origin of the SH2 domain using two-dimensional structural alignment.

The availability of large volumes of genomic sequences presents an unprecedented proteomic challenge to characterize the structure and function of various protein motifs. Primary structural alignment is often unable to accurately identify a given motif due to sequence divergence; however, with the aid of secondary structural prediction for analysis, it becomes feasible to explore protein motifs on a proteome-wide scale. Here we report the use of secondary structural alignment to characterize the Src homology 2 (SH2) domains of both conventional and divergent sequences and divide them into two groups, Src-type and STAT-type. In addition to the basic "alphabetabetabetaalpha" structure (betaBeta), the Src-type SH2 domain contains an extra beta-strand (betaE or betaE-betaF motif). Alternatively, the linker domain-conjugated SH2 domain in STAT contains the alphaB' motif. Combining BLAST data from betaBeta core motif sequences with predicted secondary structural alignment, we have screened for SH2 domains in various eukaryotic model systems including Arabidopsis, Dictyostelium, and Saccharomyces. Two novel genes carrying the linker-SH2 domain of STAT were discovered and subsequently cloned from Arabidopsis. These genes, designated as STAT-type linker-SH2 domain factors (STATL), are found in a wide array of vascular and nonvascular plants, suggesting that the linker-SH2 domain evolved prior to the divergence of plants and animals. Using this approach, we expanded the number of putative SH2 domain-bearing genes in Dictyostelium and comparatively studied the secondary structural profiles of both typical and atypical SH2 domains. Our results indicate that the linker-SH2 domain of the transcription factor STAT is one of the most ancient and fully developed functional domains, serving as a template for the continuing evolution of the SH2 domain essential for phosphotyrosine signal transduction.

Amino Acid Sequence↗

Inference of bacterial microevolution using multilocus sequence data.

We describe a model-based method for using multilocus sequence data to infer the clonal relationships of bacteria and the chromosomal position of homologous recombination events that disrupt a clonal pattern of inheritance. The key assumption of our model is that recombination events introduce a constant rate of substitutions to a contiguous region of sequence. The method is applicable both to multilocus sequence typing (MLST) data from a few loci and to alignments of multiple bacterial genomes. It can be used to decide whether a subset of isolates share common ancestry, to estimate the age of the common ancestor, and hence to address a variety of epidemiological and ecological questions that hinge on the pattern of bacterial spread. It should also be useful in associating particular genetic events with the changes in phenotype that they cause. We show that the model outperforms existing methods of subdividing recombinogenic bacteria using MLST data and provide examples from Salmonella and Bacillus. The software used in this article, ClonalFrame, is available from http://bacteria.stats.ox.ac.uk/.

Bacteria↗

Quantifying single gene copy number by measuring fluorescent probe lengths on combed genomic DNA.

An approach was developed for the quantification of subtle gains and losses of genomic DNA. The approach relies on a process called molecular combing. Molecular combing consists of the extension and alignment of purified molecules of genomic DNA on a glass coverslip. It has the advantage that a large number of genomes can be combed per coverslip, which allows for a statistically adequate number of measurements to be made on the combed DNA. Consequently, a high-resolution approach to mapping and quantifying genomic alterations is possible. The approach consists of applying fluorescence hybridization to the combed DNA by using probes to identify the amplified region. Measurements then are made on the linear hybridization signals to ascertain the region's exact size. The reliability of the approach first was tested for low copy number amplifications by determining the copy number of chromosome 21 in a normal and trisomy 21 cell line. It then was tested for high copy number amplifications by quantifying the copy number of an oncogene amplified in the tumor cell line GTL-16. These results demonstrate that a wide range of amplifications can be accurately and reliably quantified. The sensitivity and resolution of the approach likewise was assessed by determining the copy number of a single allele (160 kb) alteration.

Bacteriophage lambda↗

Gene recognition in eukaryotic DNA by comparison of genomic sequences.

MOTIVATION: Sequencing of complete eukaryotic genomes and large syntenic fragments of genomes makes it possible to apply genomic comparison for gene recognition. RESULTS: This paper describes a spliced alignment algorithm that aligns candidate exon chains of two homologous genomic sequence fragments from different species. The algorithm is implemented in Pro-Gen software. Unlike other algorithms, Pro-Gen does not assume conservation of the exon-intron structure. Amino acid sequences obtained by the formal translation of candidate exons are aligned instead of nucleotide sequences, which allows for distant comparisons. The algorithm was tested on a sample of human-mammal (mouse), human-vertebrate (Xenopus ) and human-invertebrate (Drosophila ) gene pairs. Surprisingly, the best results, 97-98% correlation between the actual and predicted genes, were obtained for more distant comparisons, whereas the correlation on the human-mouse sample was only 93%. The latter value increases to 95% if conservation of the exon-intron structure is assumed. This is caused by a large amount of sequence conservation in non-coding regions of the human and mouse genes probably due to regulatory elements. AVAILABILITY: Pro-Gen v. 3.0 is available to academic researchers free of charge at http://www.anchorgen.com/pro_gen/pro_gen.html.

Algorithms↗

Genome trees and the nature of genome evolution.

Genome trees are a means to capture the overwhelming amount of phylogenetic information that is present in genomes. Different formalisms have been introduced to reconstruct genome trees on the basis of various aspects of the genome. On the basis of these aspects, we separate genome trees into five classes: (a) alignment-free trees based on statistic properties of the genome, (b) gene content trees based on the presence and absence of genes, (c) trees based on chromosomal gene order, (d) trees based on average sequence similarity, and (e) phylogenomics-based genome trees. Despite their recent development, genome tree methods have already had some impact on the phylogenetic classification of bacterial species. However, their main impact so far has been on our understanding of the nature of genome evolution and the role of horizontal gene transfer therein. An ideal genome tree method should be capable of using all gene families, including those containing paralogs, in a phylogenomics framework capitalizing on existing methods in conventional phylogenetic reconstruction. We expect such sophisticated methods to help us resolve the branching order between the main bacterial phyla.

Evolution, Molecular↗