Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

FBSA: feature-based sequence alignment technique for very large sequences.

The ability to align pairs of very large molecular sequences is essential for a range of comparative genomic studies. However, given the complexity of genomic sequences, it has been difficult to devise a systematic method that can align - even within the same species - pairs of large sequences. Most existing approaches typically attempt to align nucleotide sequences while ignoring valuable features contained within them, eg they filter out low-complexity regions and retroelements before aligning the sequences. However, features are then added post-alignment for visualisation and analysis purposes. We argue that repetitive elements and other features (such as genes, exons and regulatory elements) should be part of the alignment process. A hierarchical approach that aligns the biologically relevant features before aligning the detailed nucleotide sequences has a number of interesting characteristics: (1) features define 'alignment anchor points' that can guide meaningful nucleotide alignment; (2) features can be weighted; (3) a hierarchical approach would identify only meaningful regions to be aligned; (4) nucleotide sequences can be described as sequences of features and non-features, providing a natural mechanism to divide the sequences for processing; and (5) computational speed is significantly faster than other approaches. In this paper, we describe and discuss a feature-based approach to aligning large genome sequences. We refer to this as 'feature-based sequence alignment'.

Algorithms↗

MGView: an alignment and visualization tool to enhance gap closure of microbial genomes.

Gap closure is a challenging phase in microbial random shotgun genome sequencing projects, particularly since genome assemblies are often complicated by the presence of repeat elements, insertion sequences and other similar factors that contribute to sequence misassemblies. While it is well recognized that the conservation of genetic information between microbial genomes, combined with the exponential increase in available microbial sequences, can be exploited to increase the efficiency of gap closure, we lack the computational tools to aid in this process. We describe here a new tool, MGView, which was developed to create a graphical depiction of the alignment of a set of microbial contigs against a completed microbial genome. The results of our assembly of the Staphylococcus aureus RF122 genome show that MGView enables a considerable reduction in time and economic cost associated with closure. Together, the results also show that the application of MGView not only enables a reduction in fold-coverage requirements of the random shotgun sequence phase, but also provides interesting insights into differences in gene content and organization between finished and unfinished microbial genomes.

Base Sequence↗

The many faces of sequence alignment.

Starting with the sequencing of the mouse genome in 2002, we have entered a period where the main focus of genomics will be to compare multiple genomes in order to learn about human biology and evolution at the DNA level. Alignment methods are the main computational component of this endeavour. This short review aims to summarise the current status of research in alignments, emphasising large-scale genomic comparisons and suggesting possible directions that will be explored in the near future.

Algorithms↗

SLAM: cross-species gene finding and alignment with a generalized pair hidden Markov model.

Comparative-based gene recognition is driven by the principle that conserved regions between related organisms are more likely than divergent regions to be coding. We describe a probabilistic framework for gene structure and alignment that can be used to simultaneously find both the gene structure and alignment of two syntenic genomic regions. A key feature of the method is the ability to enhance gene predictions by finding the best alignment between two syntenic sequences, while at the same time finding biologically meaningful alignments that preserve the correspondence between coding exons. Our probabilistic framework is the generalized pair hidden Markov model, a hybrid of (1). generalized hidden Markov models, which have been used previously for gene finding, and (2). pair hidden Markov models, which have applications to sequence alignment. We have built a gene finding and alignment program called SLAM, which aligns and identifies complete exon/intron structures of genes in two related but unannotated sequences of DNA. SLAM is able to reliably predict gene structures for any suitably related pair of organisms, most notably with fewer false-positive predictions compared to previous methods (examples are provided for Homo sapiens/Mus musculus and Plasmodium falciparum/Plasmodium vivax comparisons). Accuracy is obtained by distinguishing conserved noncoding sequence (CNS) from conserved coding sequence. CNS annotation is a novel feature of SLAM and may be useful for the annotation of UTRs, regulatory elements, and other noncoding features.

Animals↗

Sequence-based alignment of sorghum chromosome 3 and rice chromosome 1 reveals extensive conservation of gene order and one major chromosomal rearrangement.

The completed rice genome sequence will accelerate progress on the identification and functional classification of biologically important genes and serve as an invaluable resource for the comparative analysis of grass genomes. In this study, methods were developed for sequence-based alignment of sorghum and rice chromosomes and for refining the sorghum genetic/physical map based on the rice genome sequence. A framework of 135 BAC contigs spanning approximately 33 Mbp was anchored to sorghum chromosome 3. A limited number of sequences were collected from 118 of the BACs and subjected to BLASTX analysis to identify putative genes and BLASTN analysis to identify sequence matches to the rice genome. Extensive conservation of gene content and order between sorghum chromosome 3 and the homeologous rice chromosome 1 was observed. One large-scale rearrangement was detected involving the inversion of an approximately 59 cM block of the short arm of sorghum chromosome 3. Several small-scale changes in gene collinearity were detected, indicating that single genes and/or small clusters of genes have moved since the divergence of sorghum and rice. Additionally, the alignment of the sorghum physical map to the rice genome sequence allowed sequence-assisted assembly of an approximately 1.6 Mbp sorghum BAC contig. This streamlined approach to high-resolution genome alignment and map building will yield important information about the relationships between rice and sorghum genes and genomic segments and ultimately enhance our understanding of cereal genome structure and evolution.

Base Sequence↗

CpG Island microarray probe sequences derived from a physical library are representative of CpG Islands annotated on the human genome.

An effective tool for the global analysis of both DNA methylation status and protein-chromatin interactions is a microarray constructed with sequences containing regulatory elements. One type of array suited for this purpose takes advantage of the strong association between CpG Islands (CGIs) and gene regulatory regions. We have obtained 20,736 clones from a CGI Library and used these to construct CGI arrays. The utility of this library requires proper annotation and assessment of the clones, including CpG content, genomic origin and proximity to neighboring genes. Alignment of clone sequences to the human genome (UCSC hg17) identified 9595 distinct genomic loci; 64% were defined by a single clone while the remaining 36% were represented by multiple, redundant clones. Approximately 68% of the loci were located near a transcription start site. The distribution of these loci covered all 23 chromosomes, with 63% overlapping a bioinformatically identified CGI. The high representation of genomic CGI in this rich collection of clones supports the utilization of microarrays produced with this library for the study of global epigenetic mechanisms and protein-chromatin interactions. A browsable database is available on-line to facilitate exploration of the CGIs in this library and their association with annotated genes or promoter elements.

Base Sequence↗

Assembly, annotation, and integration of UNIGENE clusters into the human genome draft.

The recent release of the first draft of the human genome provides an unprecedented opportunity to integrate human genes and their functions in a complete positional context. However, at least three significant technical hurdles remain: first, to assemble a complete and nonredundant human transcript index; second, to accurately place the individual transcript indices on the human genome; and third, to functionally annotate all human genes. Here, we report the extension of the UNIGENE database through the assembly of its sequence clusters into nonredundant sequence contigs. Each resulting consensus was aligned to the human genome draft. A unique location for each transcript within the human genome was determined by the integration of the restriction fingerprint, assembled genomic contig, and radiation hybrid (RH) maps. A total of 59,500 UNIGENE clusters were mapped on the basis of at least three independent criteria as compared with the 30,000 human genes/ESTs currently mapped in Genemap'99. Finally, the extension of the human transcript consensus in this study enabled a greater number of putative functional assignments than the 11,000 annotated entries in UNIGENE. This study reports a draft physical map with annotations for a majority of the human transcripts, called the Human Index of Nonredundant Transcripts (HINT). Such information can be immediately applied to the discovery of new genes and the identification of candidate genes for positional cloning.

Alleles↗

CREMSA: compressed indexing of (ultra) large multiple sequence alignments.

MOTIVATION: Recent viral outbreaks motivate the systematic collection of pathogenic genomes in order to accelerate their study and monitor the apparition/spread of variants. Due to their limited length and temporal proximity of their sequencing, viral genomes are usually organized, and analyzed as oversized Multiple Sequence Alignments (MSAs). Such MSAs are largely ungapped, and mostly homogeneous on a column-wise level but not at a sequential level due to local variations, hindering the performances of sequential compression algorithms. RESULTS: In order to enable an efficient handling of MSAs, including subsequent statistical analyses, we introduce CREMSA (Column-wise Run-length Encoding for MSAs), a new index that builds on sparse bitvector representations to compress an existing or streamed MSA, all the while allowing for an expressive set of accelerated requests to query the alignment without prior decompression. Using CREMSA, a 65 GB MSA consisting of 1.9M SARS-CoV 2 genomes could be compressed into 22 MB using less than half a gigabyte of main memory, while executing access requests in the order of 100 ns. Such a speed up enables a comprehensive analysis of covariation over this very large MSA. We further assess the impact of the sequence ordering on the compressibility of MSAs and propose a resorting strategy that, despite the proven NP-hardness of an optimal sort, induces greatly increased compression ratios at a marginal computational cost. AVAILABILITY AND IMPLEMENTATION: CREMSA is freely accessible at https://gitlab.univ-lille.fr/cremsa/cremsa. The Snakemake workflow for the benchmarks is available at: https://gitlab.univ-lille.fr/cremsa/bench. The data used in the paper is on Zenodo at https://zenodo.org/records/14698859 and https://zenodo.org/records/15100011.

SARS-CoV-2↗

GeneTrees: a phylogenomics resource for prokaryotes.

The GeneTrees phylogenomics system pursues comparative genomic analyses from the perspective of gene phylogenies for individual genes. The GeneTrees project has the goal of providing detailed evolutionary models for all protein-coding gene components of the fully sequenced genomes. Currently, a database of alignments and trees for all protein sequences for 325 fully sequenced and annotated prokaryote genomes is available. The prokaryote database contains 890,000 protein sequences organized into over 100,000 alignments, each described by a phylogenetic tree. An original homology group discovery tool assembles sets of related proteins from all versus all pairwise alignments. Multiple alignments for each homology group are stored and subjected to phylogenetic tree inference. A graphical web interface provides visual exploration of the GeneTrees database. Homology groups can be queried by sequence identifiers or annotation terms. Genomes can be browsed visually on a gene map of each chromosome or plasmid. Phylogenetic trees with support values are displayed in conjunction with the associated sequence alignment. A variety of classes of information can be selected to label the tree tips to aid in visual evaluation of annotation and gene function. This web interface is available at http://genetrees.vbi.vt.edu.

Bacterial Proteins↗

Heat shock protein 60 sequence comparisons: duplications, lateral transfer, and mitochondrial evolution.

Heat shock proteins 60 (GroEL) are highly expressed essential proteins in eubacterial genomes and in eukaryotic organelles. These chaperone proteins have been advanced as propitious marker sequences for tracing the evolution of mitochondrial (Mt) genomes. Similarities among HSP60 sequences based on significant segment pair alignment calculations are used to deduce associations of sequences taking into account GroEL functional/structural domain differences and to relate HSP60 duplications pervasive in alpha-proteobacterial lineages to the dynamics of lateral transfer and plasmid integration. Multiple alignments with consensuses are determined for 10 natural groups. The group consensuses sharpen the similarity contrasts among individual sequences. In particular, the Mt group matches best with the classical alpha-proteobacteria and closely with Rickettsia but significantly worse with the rickettsial groups Ehrlichia and Orientia. However, across broad protein sequence comparisons, there appears to be no consistent prokaryote whose protein sequences align best with animal Mt genomes. There are plausible scenarios indicating that the nuclear-encoded HSP60 (and HSP70) sequences functioning in Mt are results of lateral transfer and are probably derived from an alpha-proteobacterium. This hypothesis relates to the plethora of duplicated HSP60 sequences among the classical alpha-proteobacteria contrasted with no duplications of HSP60 among other clades of proteobacterial genomes. Evolutionary relations are confounded by differential selection pressures, convergence, variable mutational rates, site variability, and lateral gene transfer.

Bacteria↗

Molecular structure of the Frankia spp. nifD-K intergenic spacer and design of Frankia genus compatible primer.

The nifD-K intergenic spacer (IGS) of ArI3 and ACoN24d were found to have a length 265 and 199 nucleotides, respectively. They are markedly less conserved than the two neighbouring genes and have, in some instances, a repeated structure reminiscent of an insertion event. The repeated sequence and the IGSs have no detectable homology with sequences in DNA databanks. The IGS has a stem-loop structure with a low folding energy, lower than that between nifH and nifD. No convincing alignment of IGS sequences could be obtained among Frankia strains. Only between ACoN24d and ArI3, which belong to the same genomic species, was the alignment good enough to permit detection of a doubly repeated structure. No promoter could be detected in the IGSs. The putative nifK open reading frame (ORF) in Frankia strain ArI3 has a length of 1587 nucleotides, starting with a GTG codon, preceded by a ribosome binding site of a structure similar to that of nifH (GGAGGN7). The codon usage was similar to that of previously sequenced Frankia genes with a strong bias toward G- and C-ending codons except in the case of glycine where GGT is frequent. Alignment of the three Frankia nifK sequences (EUN1f; ArI3 and ACoN24d) with those of other nitrogen-fixing bacteria permitted detection of a sequence conserved among the three Frankia strains but absent in the other sequences. A primer targeted to that region in combination with FGPD807-85 amplified the nifD-KIGS sequences of all Frankia strains (except the non-nitrogen-fixing Frankia strains CN3 and AgB1-9) and yet failed to amplify DNA of all other nitrogen-fixing bacteria.(ABSTRACT TRUNCATED AT 250 WORDS)

Actinomycetales↗

Computational analysis of alternative splicing using EST tissue information.

Expressed sequence tags (ESTs) from normal and tumor tissues have been deposited in public databases. These ESTs and all mRNA sequences were aligned with the human genome sequence using LEADS, Compugen's alternative splicing modeling platform. We developed a novel computational approach to analyze tissue information of aligned ESTs in order to identify cancer-specific alternative splicing and gene segments highly expressed in particular cancers. Several genes, including one encoding a possible pre-mRNA splicing factor, displayed cancer-specific alternative splicing. In addition, multiple candidate gene segments highly expressed in colon cancers were identified.

Alternative Splicing↗

Optimal classification of protein sequences and selection of representative sets from multiple alignments: application to homologous families and lessons for structural genomics.

Hierarchical classification is probably the most popular approach to group related proteins. However, there are a number of problems associated with its use for this purpose. One is that the resulting tree showing a nested sequence of groups may not be the most suitable representation of the data. Another is that visual inspection is the most common method to decide the most appropriate number of subsets from a tree. In fact, classification of proteins in general is bedevilled with the need for subjective thresholds to define group membership (e.g., 'significant' sequence identity for homologous families). Such arbitrariness is not only intellectually unsatisfying but also has important practical consequences. For instance, it hinders meaningful identification of protein targets for structural genomics. I describe an alternative approach to cluster related proteins without the need for an a priori threshold: one, through its use of dynamic programming, which is guaranteed to produce globally optimal solutions at all levels of partition granularity. Grouping proteins according to weights assigned to their aligned sequences makes it possible to delineate dynamically a 'core-periphery' structure within families. The 'core' of a protein family comprises the most typical sequences while the 'periphery' consists of the atypical ones. Further, a new sequence weighting scheme that combines the information in all the multiply aligned positions of an alignment in a novel way is put forward. Instead of averaging over all positions, this procedure takes into account directly the distribution of sequence variability along an alignment. The relationships between sequence weights and sequence identity are investigated for 168 families taken from HOMSTRAD, a database of protein structure alignments for homologous families. An exact solution is presented for the problem of how to select the most representative pair of sequences for a protein family. Extension of this approach by a greedy algorithm allows automatic identification of a minimal set of aligned sequences. The results of this analysis are available on the Web at http://mathbio.nimr.mrc.ac.uk/~amay.

Algorithms↗

Numerous novel annotations of the human genome sequence supported by a 5'-end-enriched cDNA collection.

A collection of 90,000 human cDNA clones generated to increase the fraction of "full-length" cDNAs available was analyzed by sequence alignment on the human genome assembly. Five hundred fifty-two gene models not found in LocusLink, with coding regions of at least 300 bp, were defined by using this collection. Exon composition proposed for novel genes showed an average of 4.7 exons per gene. In 20% of the cases, at least half of the exons predicted for new genes coincided with evolutionary conserved regions defined by sequence comparisons with the pufferfish Tetraodon nigroviridis. Among this subset, CpG islands were observed at the 5' end of 75%. In-frame stop codons upstream of the initiator ATG were present in 49% of the new genes, and 16% contained a coding region comprising at least 50% of the cDNA sequence. This cDNA resource also provided candidate small protein-coding genes, usually not included in genome annotations. In addition, analysis of a sample from this cDNA collection indicates that approximately 380 gene models described in LocusLink could be extended at their 5' end by at least one new exon. Finally, this cDNA resource provided an experimental support for annotations based exclusively on predictions, thus representing a resource substantially improving the human genome annotation.

5' Untranslated Regions↗

A comparison of expressed sequence tags (ESTs) to human genomic sequences.

The Expressed Sequence Tag (EST) division of GenBank, dbEST, is a large repository of the data being generated by human genome sequencing centers. ESTs are short, single pass cDNA sequences generated from randomly selected library clones. The approximately 415 000 human ESTs represent a valuable, low priced, and easily accessible biological reagent. As many ESTs are derived from yet uncharacterized genes, dbEST is a prime starting point for the identification of novel mRNAs. Conversely, other genes are represented by hundreds of ESTs, a redundancy which may provide data about rare mRNA isoforms. Here we present an analysis of >1000 ESTs generated by the WashU-Merck EST project. These ESTs were collected by querying dbEST with the genomic sequences of 15 human genes. When we aligned the matching ESTs to the genomic sequences, we found that in one gene, 73% of the ESTs which derive from spliced or partially spliced transcripts either contain intron sequences or are spliced at previously unreported sites; other genes have lower percentages of such ESTs, and some have none. This finding suggests that ESTs could provide researchers with novel information about alternative splicing in certain genes. In a related analysis of pairs of ESTs which are reported to derive from a single gene, we found that as many as 26% of the pairs do not BOTH align with the sequence of the same gene. We suspect that some of these unusual ESTs result from artifacts in EST generation, and caution researchers that they may find such clones while analyzing sequences in dbEST.

Alternative Splicing↗

Molecular cloning and genomic organization of a novel receptor from Drosophila melanogaster structurally related to mammalian galanin receptors.

We screened the Berkeley "Drosophila Genome Project" database with "electronic probes" corresponding to conserved amino acid sequences from the five known rat somatostatin receptors. This yielded alignment with a Drosophila genomic clone that contained a DNA sequence coding for a protein, having amino acid sequence identities with the rat galanin receptors. Using PCR with Drosophila cDNA as a template, and oligonucleotide probes coding for the exons of the presumed Drosophila gene, we were able to clone the cDNA for this receptor. The Drosophila receptor has most amino acid sequence identity with the three mammalian galanin receptors (37% identity with the rat galanin receptor type-1, 32% identity with type-2, and 29% identity with type-3). Less sequence identity exists with the mammalian opioid/nociceptin-orphanin FQ receptors (26% identity with the rat micro opioid receptor), and mammalian somatostatin receptors (25% identity with the rat somatostatin receptor type-2). The novel Drosophila receptor gene contains ten introns and eleven exons and is located at the distal end of the X chromosome.

Amino Acid Sequence↗

New in silico insight into the synteny between rice (Oryza sativa L.) and maize (Zea mays L.) highlights reshuffling and identifies new duplications in the rice genome.

A unigene set of 1411 contigs was constructed from 2629 redundant maize expressed sequence tags (ESTs) mapped on the maizeDB genetic map. Rice orthologous sequences were identified by blast alignment against the rice genomic sequence. A total of 1046 (74%) maize contigs were associated with their corresponding homologues in the rice genome and 656 (47%) defined as potential orthologous relationships. One hundred and seventeen (8%) maize EST contigs mapped to two distinct loci on the maize genetic map, reflecting the tetraploid nature of the maize genome. Among 492 mono-locus contigs, 344 (484 redundant ESTs) identify collinear blocks between maize chromosomes 2 and 4 and a single rice chromosome, defining six new collinear regions. Fine-scale analysis of collinearity between rice chromosomes 1 and 5 with maize chromosomes 3, 6 and 8 shows the presence of internal rearrangements within collinear regions. Mapping of maize contigs to two distinct loci on the rice sequence identifies five new duplication events in rice. Detailed analysis of a duplication between rice chromosomes 1 and 5 shows that 11% of the annotated genes from the chromosome 1 locus are found duplicated on the chromosome 5 paralogous counterpart, indicating a high degree of re-organisations. The implications of these findings for map-based cloning in collinear regions are discussed.

Chromosome Mapping↗