Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

ECgene: genome-based EST clustering and gene modeling for alternative splicing.

With the availability of the human genome map and fast algorithms for sequence alignment, genome-based EST clustering became a viable method for gene modeling. We developed a novel gene-modeling method, ECgene (Gene modeling by EST Clustering), which combines genome-based EST clustering and the transcript assembly procedure in a coherent and consistent fashion. Specifically, ECgene takes alternative splicing events into consideration. The position of splice sites (i.e., exon-intron boundaries) in the genome map is utilized as the critical information in the whole procedure. Sequences that share any splice sites are grouped together to define an EST cluster in a manner similar to that of the genome-based version of the UniGene algorithm. Transcript assembly is achieved using graph theory that represents the exon connectivity in each cluster as a directed acyclic graph (DAG). Distinct paths along exons correspond to possible gene models encompassing all alternative splicing events. EST sequences in each cluster are subclustered further according to the compatibility with gene structure of each splice variant, and they can be regarded as clone evidence for the corresponding isoform. The reliability of each isoform is assessed from the nature of cluster members and from the minimum number of clones required to reconstruct all exons in the transcript.

Algorithms↗

The SUPERFAMILY database in 2004: additions and improvements.

The SUPERFAMILY database provides structural assignments to protein sequences and a framework for analysis of the results. At the core of the database is a library of profile Hidden Markov Models that represent all proteins of known structure. The library is based on the SCOP classification of proteins: each model corresponds to a SCOP domain and aims to represent an entire superfamily. We have applied the library to predicted proteins from all completely sequenced genomes (currently 154), the Swiss-Prot and TrEMBL databases and other sequence collections. Close to 60% of all proteins have at least one match, and one half of all residues are covered by assignments. All models and full results are available for download and online browsing at http://supfam.org. Users can study the distribution of their superfamily of interest across all completely sequenced genomes, investigate with which other superfamilies it combines and retrieve proteins in which it occurs. Alternatively, concentrating on a particular genome as a whole, it is possible first, to find out its superfamily composition, and secondly, to compare it with that of other genomes to detect superfamilies that are over- or under-represented. In addition, the webserver provides the following standard services: sequence search; keyword search for genomes, superfamilies and sequence identifiers; and multiple alignment of genomic, PDB and custom sequences.

Animals↗

Phylogeny of 33 ribosomal and six other proteins encoded in an ancient gene cluster that is conserved across prokaryotic genomes: influence of excluding poorly alignable sites from analysis.

Thirty-nine proteins encoded in a large gene cluster that is well-conserved in gene content and gene order across 18 sequenced prokaryotic genomes were extracted, aligned and subjected to phylogenetic analysis. In individual analyses of the alignments, only two probable examples of lateral gene transfer between archaea and eubacteria were detected, involving the genes for ribosomal protein Rpl23 and adenylate kinase. Amino acid sequences for 35 of the 39 proteins were concatenated to yield a data set of 9087 amino acid positions per genome. Many of these proteins, 33 of which are ribosomal proteins, are not highly conserved across distantly related organisms and thus contain many regions that are difficult to align. Phylogenetic analyses were performed with subsets of the concatenated data from which the most highly variable sites had been iteratively removed, using the number of different amino acids that occur at a given site as a criterion of variability. Glycine, which has a strong influence on protein structure, tended to be more frequent at the most conserved (least polymorphic) sites. With most subsets of the data, the proteins from the cyanobacterium Synechocystis tended to branch with their homologues from gram-positive bacteria. The results indicate that excluding only a few percentage of poorly alignable sites from phylogenetic analysis can have a severe impact upon the phylogeny inferred and that bootstrap support for branches can fluctuate substantially, depending upon which sites are excluded.

Adenylate Kinase↗

Gene for gene alignment between the Brassica and Arabidopsis genomes by direct transcriptome mapping.

We report a global gene for gene alignment of the genomes of Brassica oleracea and Arabidopsis thaliana by construction of a transcriptome map based on B. oleracea cDNAs obtained from leaf tissue. cDNAs were synthesized from total RNA extracted from individual F2s of a mapping population resulting from crossing double-haploids of broccoli and cauliflower. The map consisted of 247 cDNA markers obtained by the SRAP technique. After sequencing 190 of the polymorphic cDNA bands, FASTA detected 169 sequences with similarity to genes reported in Arabidopsis. There was extensive colinearity between the two genomes for chromosomal segments rather than for whole chromosomes, often showing inversions and deletions/insertions. Large-scale duplications were observed in the B. oleracea genome, but were unevenly distributed, arguing against ancient triplication of the entire genome. The most duplicated segments corresponded to those found on Arabidopsis chromosomes 1 and 5, whereas chromosomes 2 and 4 were the least represented in Brassica. Clear differences in the similarity score value of related sequences allowed the identification of orthologs. Transcriptome mapping is an efficient approach that allows gene-for-gene alignment between a fully sequenced and a poorly characterized genome.

Arabidopsis↗

MAP: searching large genome databases.

A number of biological applications require comparison of large genome strings. Current techniques suffer from both disk I/O and computational cost because of extensive memory requirements and large candidate sets. We propose an efficient technique for alignment of large genome strings. Our technique precomputes the associations between the database strings and the query string. These associations are used to prune the database-query substring pairs that do not contain similar regions. We use a hash table to compare the unpruned regions of the query and database strings. The cost of the ensuing search is determined by how the hash table is constructed. We present a dynamic strategy that optimizes the random disk I/O needed for accessing the hash table. It also provides the user a coarse grain visualization of the similarity pattern quickly before the actual search. The experimental results show that our technique aligns genome strings up to 97 times faster than BLAST.

Algorithms↗

Localization and characterization of mouse-human alignments within the human genome. Does evolutionary conservation suggest functional importance?

In an attempt to validate the use of evolutionary conservation as a method to identify putative regulatory elements, we have quantified the frequency of Single Nucleotide Polymorphisms (SNPs) within the most tightly conserved regions across the entire Human Genome. Our results show that conserved non-coding sequences have a significantly lower SNP frequency than their exonic counterparts, which suggests that these regions are functionally important.

Animals↗

A computer program for aligning a cDNA sequence with a genomic DNA sequence.

We address the problem of efficiently aligning a transcribed and spliced DNA sequence with a genomic sequence containing that gene, allowing for introns in the genomic sequence and a relatively small number of sequencing errors. A freely available computer program, described herein, solves the problem for a 100-kb genomic sequence in a few seconds on a workstation.

Algorithms↗

Efficient large-scale sequence comparison by locality-sensitive hashing.

MOTIVATION: Comparison of multimegabase genomic DNA sequences is a popular technique for finding and annotating conserved genome features. Performing such comparisons entails finding many short local alignments between sequences up to tens of megabases in length. To process such long sequences efficiently, existing algorithms find alignments by expanding around short runs of matching bases with no substitutions or other differences. Unfortunately, exact matches that are short enough to occur often in significant alignments also occur frequently by chance in the background sequence. Thus, these algorithms must trade off between efficiency and sensitivity to features without long exact matches. RESULTS: We introduce a new algorithm, LSH-ALL-PAIRS, to find ungapped local alignments in genomic sequence with up to a specified fraction of substitutions. The length and substitution rate of these alignments can be chosen so that they appear frequently in significant similarities yet still remain rare in the background sequence. The algorithm finds ungapped alignments efficiently using a randomized search technique, locality-sensitive hashing. We have found LSH-ALL-PAIRS to be both efficient and sensitive for finding local similarities with as little as 63% identity in mammalian genomic sequences up to tens of megabases in length

Algorithms↗

Polymorphix: a sequence polymorphism database.

Within-species sequence variation data are of special interest since they contain information about recent population/species history, and the molecular evolutionary forces currently in action in natural populations. These data, however, are presently dispersed within generalist databases, and are difficult to access. To solve this problem, we have developed Polymorphix, a database dedicated to sequence polymorphism. It contains within-species homologous sequence families built using EMBL/GenBank under suitable similarity and bibliographic criteria. Polymorphix is an ACNUC structured database allowing both simple and complex queries for population genomic studies. Alignments within families as well as phylogenetic trees can be download. When available, outgroups are included in the alignment. Polymorphix contains sequences from the nuclear, mitochondrial and chloroplastic genomes of every eukaryote species represented in EMBL. It can be accessed by a web interface (http://pbil.univ-lyon1.fr/polymorphix/query.php).

Animals↗

Glocal alignment: finding rearrangements during alignment.

MOTIVATION: To compare entire genomes from different species, biologists increasingly need alignment methods that are efficient enough to handle long sequences, and accurate enough to correctly align the conserved biological features between distant species. The two main classes of pairwise alignments are global alignment, where one string is transformed into the other, and local alignment, where all locations of similarity between the two strings are returned. Global alignments are less prone to demonstrating false homology as each letter of one sequence is constrained to being aligned to only one letter of the other. Local alignments, on the other hand, can cope with rearrangements between non-syntenic, orthologous sequences by identifying similar regions in sequences; this, however, comes at the expense of a higher false positive rate due to the inability of local aligners to take into account overall conservation maps. RESULTS: In this paper we introduce the notion of glocal alignment, a combination of global and local methods, where one creates a map that transforms one sequence into the other while allowing for rearrangement events. We present Shuffle-LAGAN, a glocal alignment algorithm that is based on the CHAOS local alignment algorithm and the LAGAN global aligner, and is able to align long genomic sequences. To test Shuffle-LAGAN we split the mouse genome into BAC-sized pieces, and aligned these pieces to the human genome. We demonstrate that Shuffle-LAGAN compares favorably in terms of sensitivity and specificity with standard local and global aligners. From the alignments we conclude that about 9% of human/mouse homology may be attributed to small rearrangements, 63% of which are duplications.

Algorithms↗

MAVID: constrained ancestral alignment of multiple sequences.

We describe a new global multiple-alignment program capable of aligning a large number of genomic regions. Our progressive-alignment approach incorporates the following ideas: maximum-likelihood inference of ancestral sequences, automatic guide-tree construction, protein-based anchoring of ab-initio gene predictions, and constraints derived from a global homology map of the sequences. We have implemented these ideas in the MAVID program, which is able to accurately align multiple genomic regions up to megabases long. MAVID is able to effectively align divergent sequences, as well as incomplete unfinished sequences. We demonstrate the capabilities of the program on the benchmark CFTR region, which consists of 1.8 Mb of human sequence and 20 orthologous regions in marsupials, birds, fish, and mammals. Finally, we describe two large MAVID alignments, an alignment of all the available HIV genomes and a multiple alignment of the entire human, mouse, and rat genomes.

Animals↗

Complete comparative genomic analysis of two field isolates of Mamestra configurata nucleopolyhedrovirus-A.

A second genotype of Mamestra configurata nucleopolyhedrovirus-A (MacoNPV-A), variant 90/4 (v90/4), was identified due to its altered restriction endonuclease profile and reduced virulence for the host insect, M. configurata, relative to the archetypal genotype, MacoNPV-A variant 90/2 (v90/2). To investigate the genetic differences between these two variants, the genome of v90/4 was sequenced completely. The MacoNPV-A v90/4 genome is 153 656 bp in size, 1404 bp smaller than the v90/2 genome. Sequence alignment showed that there was 99.5 % nucleotide sequence identity between the genomes of v90/4 and v90/2. However, the v90/4 genome has 521 point mutations and numerous deletions and insertions when compared to the genome of v90/2. Gene content and organization in the genome of v90/4 is identical to that in v90/2, except for an additional bro gene that is found in the v90/2 genome. The region between hr1 and orf31 shows the greatest divergence between the two genomes. This region contains three bro genes, which are among the most variable baculovirus genes. These results, together with other published data, suggest that bro genes may influence baculovirus genome diversity and may be involved in recombination between baculovirus genomes. Many ambiguous residues found in the v90/4 sequence also reveal the presence of 214 sequence polymorphisms. Sequence analysis of cloned HindIII fragments of the original MacoNPV field isolate that the 90/4 variant was derived from indicates that v90/4 is an authentic variant and may represent approximately 25 % of the genotypes in the field isolate. These results provide evidence of extensive sequence variation among the individual genomes comprising a natural baculovirus outbreak in a continuous host population.

Animals↗

Characterization of intron loss events in mammals.

The exon/intron structure of eukaryotic genes differs extensively across species, but the mechanisms and relative rates of intron loss and gain are still poorly understood. Here, we used whole-genome sequence alignments of human, mouse, rat, and dog to perform a genome-wide analysis of intron loss and gain events in >17,000 mammalian genes. We found no evidence for intron gain and 122 cases of intron loss, most of which occurred within the rodent lineage. The majority (68%) of the deleted introns were extremely small (<150 bp), significantly smaller than average. The intron losses occurred almost exclusively within highly expressed, housekeeping genes, supporting the hypothesis that intron loss is mediated via germline recombination of genomic DNA with intronless cDNA. This study constitutes the largest scale analysis for intron dynamics in vertebrates to date and allows us to confirm and extend several hypotheses previously based on much smaller samples. Our results in mammals show that intron gain has not been a factor in the evolution of gene structure during the past 95 Myr and has likely been restricted to more ancient history.

Animals↗

Evolutionarily conserved coding sequences in the dpy-20-unc-22 region of Caenorhabditis elegans.

Caenorhabditis elegans provides an excellent opportunity to study the organization of a complex genome. The alignment of the genetic and molecular maps over a large stretch of the genome is an essential part of this study. The objective of this paper was the identification and characterization of coding regions in four cosmids containing DNA from the interval between dpy-20 and unc-22 on linkage group IV. These cosmids were characterized with regard to the map position and the developmental patterns of expression of coding sequences. Since an extensive genetic map already exists for this region, this detailed description of the coding sequences in the dpy-20-unc-22 region will make possible alignment of the molecular and genetic maps for this portion of the C. elegans genome. In this study, we have used interspecies cross-hybridization to localize and identify potential coding elements. We have investigated four cosmids containing approximately 150 kb of C. elegans genome adjacent to the well-characterized muscle gene, unc-22(IV). Fragments subcloned from the four cosmids were hybridized at moderate stringency to the genome of the related species, Caenorhabditis briggsae. In this way nine potential coding regions were identified. Seven of these nine fragments also hybridized to mRNA transcripts on Northern blots. Five of the seven showed maximal hybridization to RNA from L2-stage animals, a pattern that resembles that of actin transcription. It is speculated that the functions of these five may in some way be related to one another, and perhaps also to that of unc-22, which is itself a muscle gene.

Animals↗

Coupled analysis of gene expression and chromosomal location.

Microarray technology can be used to assess simultaneously global changes in expression of mRNA or genomic DNA copy number among thousands of genes in different biological states. In many cases, it is desirable to determine if altered patterns of gene expression correlate with chromosomal abnormalities or assess expression of genes that are contiguous in the genome. We describe a method, differential gene locus mapping (DIGMAP), which aligns the known chromosomal location of a gene to its expression value deduced by microarray analysis. The method partitions microarray data into subsets by chromosomal location for each gene interrogated by an array. Microarray data in an individual subset can then be clustered by physical location of genes at a subchromosomal level based upon ordered alignment in genome sequence. A graphical display is generated by representing each genomic locus with a colored cell that quantitatively reflects its differential expression value. The clustered patterns can be viewed and compared based on their expression signatures as defined by differential values between control and experimental samples. In this study, DIGMAP was tested using previously published studies of breast cancer analyzed by comparative genomic hybridization (CGH) and prostate cancer gene expression profiles assessed by cDNA microarray experiments. Analysis of the breast cancer CGH data demonstrated the ability of DIGMAP to deduce gene amplifications and deletions. Application of the DIGMAP method to the prostate data revealed several carcinoma-related loci, including one at 16q13 with marked differential expression encompassing 19 known genes including 9 encoding metallothionein proteins. We conclude that DIGMAP is a powerful computational tool enabling the coupled analysis of microarray data with genome location.

Chromosome Mapping↗

Using multiple alignments to improve seeded local alignment algorithms.

Multiple alignments among genomes are becoming increasingly prevalent. This trend motivates the development of tools for efficient homology search between a query sequence and a database of multiple alignments. In this paper, we present an algorithm that uses the information implicit in a multiple alignment to dynamically build an index that is weighted most heavily towards the promising regions of the multiple alignment. We have implemented Typhon, a local alignment tool that incorporates our indexing algorithm, which our test results show to be more sensitive than algorithms that index only a sequence. This suggests that when applied on a whole-genome scale, Typhon should provide improved homology searches in time comparable to existing algorithms.

Algorithms↗

Exploring transcription factor binding properties of several non-coding DNA sequence elements in the human NF-IL6 gene.

We examined several DNA segments upstream of the transcription start site of the human NF-IL6 gene to evaluate the predictions of two computational models developed to identify potential regulatory elements in the non-coding regions of genes. One model, comparative genomics, is based on the hypothesis that functional regulatory sequences can be localized in alignments of genomic DNA from several species. The other model is based on the hypothesis that protein-binding sites in genomic DNA may include sequence elements that occur frequently in proximal promoters of genes. The segments selected for DNA binding and functional evaluations included: (1) two conserved regions identified in multi-species sequence alignments; (2) a region containing several localized hits with 9-mers that ranked highly in studies of proximal promoters of human genes; and (3) two regions that were either GC-rich and/or contained tracts of G. The assays were done under nearly identical experimental conditions, using a cell line (U937) representing human monocytes/macrophages. The experiments also aimed at evaluating what effect, if any, cellular stimulation could have on the interactions of nuclear proteins with naturally occurring GC-rich elements in a human genomic DNA. In DNA binding assays, several complexes were formed with the conserved regions identified in multi-species sequence alignment. Furthermore, these regions were active in functional assays. The region containing several matches with 9-mers derived from proximal promoters of human genes was not conserved but formed several complexes with nuclear proteins including Sp1, Egr-1, and an unidentified protein. In addition, this region was active in functional assays and responded to cellular stimulations. Overall, the results of the assays suggest an important role for the sequence context of genomic DNA in protein binding and selection.

Animals↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗