Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Dynamite: a flexible code generating language for dynamic programming methods used in sequence comparison.

We have developed a code generating language, called Dynamite, specialised for the production and subsequent manipulation of complex dynamic programming methods for biological sequence comparison. From a relatively simple text definition file Dynamite will produce a variety of implementations of a dynamic programming method, including database searches and linear space alignments. The speed of the generated code is comparable to hand written code, and the additional flexibility has proved invaluable in designing and testing new algorithms. An innovation is a flexible labelling system, which can be used to annotate the original sequences with biological information. We illustrate the Dynamite syntax and flexibility by showing definitions for dynamic programming routines (i) to align two protein sequences under the assumption that they are both poly-topic transmembrane proteins, with the simultaneous assignment of transmembrane helices and (ii) to align protein information to genomic DNA, allowing for introns and sequencing error.

Algorithms↗

Comparative genomics.

The genomes from three mammals (human, mouse, and rat), two worms, and several yeasts have been sequenced, and more genomes will be completed in the near future for comparison with those of the major model organisms. Scientists have used various methods to align and compare the sequenced genomes to address critical issues in genome function and evolution. This review covers some of the major new insights about gene content, gene regulation, and the fraction of mammalian genomes that are under purifying selection and presumed functional. We review the evolutionary processes that shape genomes, with particular attention to variation in rates within genomes and along different lineages. Internet resources for accessing and analyzing the treasure trove of sequence alignments and annotations are reviewed, and we discuss critical problems to address in new bioinformatic developments in comparative genomics.

Animals↗

Phylogenetic position of Salinibacter ruber based on concatenated protein alignments.

A total of 22 genes from the genome of Salinibacter ruber strain M31 were selected in order to study the phylogenetic position of this species based on protein alignments. The selection of the genes was based on their essential function for the organism, dispersion within the genome, and sufficient informative length of the final alignment. For each gene, an individual phylogenetic analysis was performed and compared with the resulting tree based on the concatenation of the 22 genes, which rendered a single alignment of 10,757 homologous positions. In addition to the manually chosen genes, an automatically selected data set of 74 orthologous genes was used to reconstruct a tree based on 17,149 homologous positions. Although single genes supported different topologies, the tree topology of both concatenated data sets was shown to be identical to that previously observed based on small subunit (SSU) rRNA gene analysis, in which S. ruber was placed together with Bacteroidetes. In both concatenated data sets the bootstrap was very high, but an analysis with a gradually lower number of genes indicated that the bootstrap was greatly reduced with less than 12 genes. The results indicate that tree reconstructions based on concatenating large numbers of protein coding genes seem to produce tree topologies with similar resolution to that of the single 16S rRNA gene trees. For classification purposes, 16S rRNA gene analysis may remain as the most pragmatic approach to infer genealogic relationships.

Algorithms↗

Genetic instability and fragmentation of a stealth viral genome.

Partial sequencing was performed on cloned DNA obtained from cultures of a stealth virus isolated from a patient with the chronic fatigue syndrome. The results extend earlier findings showing regions of homology to cytomegalovirus (CMV). Although the virus is much more closely related to simian CMV than to human CMV, many of the cloned viral segments could be aligned with the human CMV genome. The aggregate size of the aligned segments exceeds 100 kilobase pairs (kbp). Undigested viral DNA has a mobility in agarose gel electrophoresis corresponding to approximately 20 kbp. The virus, therefore, apparently exists in multiple fragments. Considerable sequence variation exists between individual clones which overlap to similar regions of the human CMV genome. The fragmented genome and sequence microheterogeneity suggest that both the processivity and the fidelity of replication of the viral genome are defective. An unstable viral genome may provide a potential mechanism of recovery from stealth viral illness.

Base Sequence↗

Comparative genomics of prokaryotic GTP-binding proteins (the Era, Obg, EngA, ThdF (TrmE), YchF and YihA families) and their relationship to eukaryotic GTP-binding proteins (the DRG, ARF, RAB, RAN, RAS and RHO families).

Several GTP-binding proteins with poorly defined functions were previously identified in Escherichia coli (i.e. Era, ThdF (TrmE)), Bacillus subtilis (i.e. Obg) and Neisseria gonorrhoeae (i.e. EngA). In these species, every individual protein is encoded by an essential gene. BLAST searches were used to detect orthologs in genomes of various organisms. Alignments of orthologous sequences allowed the construction of phylogenetic trees and the definition of protein families. The BLAST searches also resulted in the identification of two additional families, the YchF and YihA families, named after the ychF and yihA genes of E. coli. Most families are not present in archaeal genomes, but representatives of each family were also detected in eukaryotic genomes. Only representatives of the YchF family are present in every genome sequenced to date, suggesting that YchF-like proteins might be involved in a fundamental life process. The GTP1/DRG family consisting of eukaryotic and archaeal proteins is related to the YchF family of GTP-binding proteins. The relationship of the six prokaryotic families of GTP-binding proteins and the GTP1/DRG family to eukaryotic GTPase families was also investigated: With the exception of the ARF family, a clear separation of the six prokaryotic families and the GTP1/DRG family with respect to eukaryotic (RAB, RAN, RAS and RHO) GTPases was observed.

Animals↗

Alu-containing exons are alternatively spliced.

Alu repetitive elements are found in approximately 1.4 million copies in the human genome, comprising more than one-tenth of it. Numerous studies describe exonizations of Alu elements, that is, splicing-mediated insertions of parts of Alu sequences into mature mRNAs. To study the connection between the exonization of Alu elements and alternative splicing, we used a database of ESTs and cDNAs aligned to the human genome. We compiled two exon sets, one of 1176 alternatively spliced internal exons, and another of 4151 constitutively spliced internal exons. Sixty one alternatively spliced internal exons (5.2%) had a significant BLAST hit to an Alu sequence, but none of the constitutively spliced internal exons had such a hit. The vast majority (84%) of the Alu-containing exons that appeared within the coding region of mRNAs caused a frame-shift or a premature termination codon. Alu-containing exons were included in transcripts at lower frequencies than alternatively spliced exons that do not contain an Alu sequence. These results indicate that internal exons that contain an Alu sequence are predominantly, if not exclusively, alternatively spliced. Presumably, evolutionary events that cause a constitutive insertion of an Alu sequence into an mRNA are deleterious and selected against.

Alternative Splicing↗

Characterization, nucleotide sequence and genome organization of leek white stripe virus, a putative new species of the genus Necrovirus.

White stripe is a disease affecting leek in France with which an isometric virus c. 30 nm in diameter is associated. The most evident symptom is the presence of white stripes on the leaves extending to the stem. Attempts to demonstrate transmission through the soil by sowing or transplanting leek in contaminated soil were unsuccessful. The virus was transmitted by sap inoculation to a narrow range of herbaceous hosts, all of which were infected only locally. Virus purification was from infected leek tissues, where it accumulated in large amounts, as demonstrated by ultrastructural observations. RNA was extracted from purified virus preparations and cDNA clones were prepared. The complete nucleotide sequence of the viral RNA was determined: The genome is 3,662 nucleotides long and contains five open reading frames (ORFs). The first (ORF 1) encodes a putative translation product of M(r) 23,803 (p24) and read through of its amber stop codon results in a protein of M(r) 82,625 (p83) (ORF 2). ORF 3 and ORF 4 encode two small polypeptides of M(r) 11,280 (p11) and M(r) 6,261 (p6), respectively. ORF 5 encodes the capsid protein of M(r) 27,460 (p27). The genome organization and sequence alignments with the corresponding products of necroviruses suggest that the virus isolated from leek is a new species in the genus Necrovirus, for which the name of leek white stripe virus (LWSV) is proposed.

Amino Acid Sequence↗

Using homolog groups to create a whole-genomic tree of free-living organisms: an update.

Genomic trees have been constructed based on the presence and absence of families of protein-encoding genes observed in 27 complete genomes, including genomes of 15 free-living organisms. This method does not rely on the identification of suspected orthologs in each genome, nor the specific alignment used to compare gene sequences because the protein-encoding gene families are formed by grouping any protein with a pairwise similarity score greater than a preset value. Because of this all inclusive grouping, this method is resilient to some effects of lateral gene transfer because transfers of genes are masked when the recipient genome already has a homolog (not necessarily an ortholog) of the incoming gene. Of 71 genes suspected to have been laterally transferred to the genome of Aeropyrum pernix, only approximately 7 to 15 represent genes where a lateral gene transfer appears to have generated homoplasy in our character dataset. The genomic tree of the 15 free-living taxa includes six different bacterial orders, six different archaeal orders, and two different eukaryotic kingdoms. The results are remarkably similar to results obtained by analysis of rRNA. Inclusion of the other 12 genomes resulted in a tree only broadly similar to that suggested by rRNA with at least some of the differences due to artifacts caused by the small genome size of many of these species. Very small genomes, such as those of the two Mycoplasma genomes included, fall to the base of the Bacterial domain, a result expected due to the substantial gene loss inherent to these lineages. Finally, artificial "partial genomes" were generated by randomly selecting ORFs from the complete genomes in order to test our ability to recover the tree generated by the whole genome sequences when only partial data are available. The results indicated that partial genomic data, when sampled randomly, could robustly recover the tree generated by the whole genome sequences.

Animals↗

Differential annotation of tRNA genes with anticodon CAT in bacterial genomes.

We have developed three strategies to discriminate among the three types of tRNA genes with anticodon CAT (tRNA(Ile), elongator tRNA(Met) and initiator tRNA(fMet)) in bacterial genomes. With these strategies, we have classified the tRNA genes from 234 bacterial and several organellar genomes. These sequences, in an aligned or unaligned format, may be used for the identification and annotation of tRNA (CAT) genes in other genomes. The first strategy is based on the position of the problem sequences in a phenogram (a tree-like network), the second on the minimum average number of differences against the tRNA sequences of the three types and the third on the search for the highest score value against the profiles of the three types of tRNA genes. The species with the maximum number of tRNA(fMet) and tRNA(Met) was Photobacterium profundum, whereas the genome of one Escherichia coli strain presented the maximum number of tRNA(Ile) (CAT) genes. This last tRNA gene and tilS, encoding an RNA-modifying enzyme, are not essential in bacteria. The acquisition of a tRNA(Ile) (TAT) gene by Mycoplasma mobile has led to the loss of both the tRNA(Ile) (CAT) and the tilS genes. The new tRNA has appropriated the function of decoding AUA codons.

Anticodon↗

Prokaryote phylogeny without sequence alignment: from avoidance signature to composition distance.

A new and essentially simple method to reconstruct prokaryotic phylogenetic trees from their complete genome data without using sequence alignment is proposed. It is based on the appearance frequency of oligopeptides of a fixed length (up to K = 6) in their proteomes. This is a method without fine adjustment and choice of genes. It can incorporate the effect of lateral gene transfer to some extent and leads to results comparable with the bacteriologists' systematics as reflected in the latest 2001 edition of the Bergey's Manual of Systematic Bacteriology [1, 2]. A key point in our approach is subtraction of a random background by using a Markovian model of order K - 1 from the composition vectors to highlight the shaping role of natural selection.

Base Composition↗

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH↗

European elk papillomavirus: characterization of the genome, induction of tumors in animals, and transformation in vitro.

The European elk papillomavirus (EEPV) genome was cloned in the BamHI cleavage site of the pBR322 vector. The cloned genome was used for construction of a physical map, employing restriction endonucleases BamHI, BglII, HindIII, PvuII, SacI, and XhoI. The sequence homology between the EEPV and bovine papillomavirus type 1 genomes was elucidated by performing hybridizations in different concentrations of formamide. Sequence homology could only be revealed under less stringent conditions, i.e., Tm - 43 degrees C. Nucleotide sequence information was also collected from the regions which lie adjacent to the three HindIII sites that are present in the EEPV genome. The results made it possible to align the EEPV and bovine papillomavirus type 1 genomes. Transformation by EEPV was demonstrated with the C127 mouse cell line, and fibrosarcomas were induced in young hamsters after subcutaneous injection. The transformed cells and the tumors contain multiple, nonintegrated copies of the EEPV genome. Virus particles could not be detected either in tumors or in transformed cells.

Animals↗

A second generation radiation hybrid map to aid the assembly of the bovine genome sequence.

BACKGROUND: Several approaches can be used to determine the order of loci on chromosomes and hence develop maps of the genome. However, all mapping approaches are prone to errors either arising from technical deficiencies or lack of statistical support to distinguish between alternative orders of loci. The accuracy of the genome maps could be improved, in principle, if information from different sources was combined to produce integrated maps. The publicly available bovine genomic sequence assembly with 6x coverage (Btau_2.0) is based on whole genome shotgun sequence data and limited mapping data however, it is recognised that this assembly is a draft that contains errors. Correcting the sequence assembly requires extensive additional mapping information to improve the reliability of the ordering of sequence scaffolds on chromosomes. The radiation hybrid (RH) map described here has been contributed to the international sequencing project to aid this process. RESULTS: An RH map for the 30 bovine chromosomes is presented. The map was built using the Roslin 3000-rad RH panel (BovGen RH map) and contains 3966 markers including 2473 new loci in addition to 262 amplified fragment-length polymorphisms (AFLP) and 1231 markers previously published with the first generation RH map. Sequences of the mapped loci were aligned with published bovine genome maps to identify inconsistencies. In addition to differences in the order of loci, several cases were observed where the chromosomal assignment of loci differed between maps. All the chromosome maps were aligned with the current 6x bovine assembly (Btau_2.0) and 2898 loci were unambiguously located in the bovine sequence. The order of loci on the RH map for BTA 5, 7, 16, 22, 25 and 29 differed substantially from the assembled bovine sequence. From the 2898 loci unambiguously identified in the bovine sequence assembly, 131 mapped to different chromosomes in the BovGen RH map. CONCLUSION: Alignment of the BovGen RH map with other published RH and genetic maps showed higher consistency in marker order and chromosome assignment than with the current 6x sequence assembly. This suggests that the bovine sequence assembly could be significantly improved by incorporating additional independent mapping information.

Animals↗

Comparative genomics for identification of clone-specific sequence blocks in Streptococcus pneumoniae.

The partial genome sequences of a serotype 3 and a serotype 2 pneumococcal strain were compared to the complete type 4 pneumococcal genome. Over 500000 and 150000 base pairs of the partial genome data, obtained from published patents, were analysed respectively. Global alignment showed that nearly the whole genome is highly conserved in accordance with data of multilocus sequence typing of housekeeping genes. The search for clone-specific genes revealed 17 new open reading frames in the type 3 strain, while no new open reading frame was detected in the type 2 strain. Allelic variation of genes was restricted by the use of crude sequence data, but still permitted identification of some new alleles and the observation that all surface proteins present in the partial genome data were highly conserved. In both strains we observed also a variety of chromosomal rearrangements and variations due to mobile genetic elements. All together, this comparative genomic approach gives a genome-based overview of strain relatedness and a prospective on what could be expected when sequencing other pneumococcal strains.

Alleles↗

Identifying uniformly mutated segments within repeats.

Given a long string of characters from a constant size alphabet we present an algorithm to determine whether its characters have been generated by a single i.i.d. random source. More specifically, consider all possible n-coin models for generating a binary string S, where each bit of S is generated via an independent toss of one of the n coins in the model. The choice of which coin to toss is decided by a random walk on the set of coins where the probability of a coin change is much lower than the probability of using the same coin repeatedly. We present a procedure to evaluate the likelihood of a n-coin model for given S, subject a uniform prior distribution over the parameters of the model (that represent mutation rates and probabilities of copying events). In the absence of detailed prior knowledge of these parameters, the algorithm can be used to determine whether the a posteriori probability for n=1 is higher than for any other n>1. Our algorithm runs in time O(l4logl), where l is the length of S, through a dynamic programming approach which exploits the assumed convexity of the a posteriori probability for n. Our test can be used in the analysis of long alignments between pairs of genomic sequences in a number of ways. For example, functional regions in genome sequences exhibit much lower mutation rates than non-functional regions. Because our test provides means for determining variations in the mutation rate, it may be used to distinguish functional regions from non-functional ones. Another application is in determining whether two highly similar, thus evolutionarily related, genome segments are the result of a single copy event or of a complex series of copy events. This is particularly an issue in evolutionary studies of genome regions rich with repeat segments (especially tandemly repeated segments).

Algorithms↗

Comparison of methods for genomic localization of gene trap sequences.

BACKGROUND: Gene knockouts in a model organism such as mouse provide a valuable resource for the study of basic biology and human disease. Determining which gene has been inactivated by an untargeted gene trapping event poses a challenging annotation problem because gene trap sequence tags, which represent sequence near the vector insertion site of a trapped gene, are typically short and often contain unresolved residues. To understand better the localization of these sequences on the mouse genome, we compared stand-alone versions of the alignment programs BLAT, SSAHA, and MegaBLAST. A set of 3,369 sequence tags was aligned to build 34 of the mouse genome using default parameters for each algorithm. Known genome coordinates for the cognate set of full-length genes (1,659 sequences) were used to evaluate localization results. RESULTS: In general, all three programs performed well in terms of localizing sequences to a general region of the genome, with only relatively subtle errors identified for a small proportion of the sequence tags. However, large differences in performance were noted with regard to correctly identifying exon boundaries. BLAT correctly identified the vast majority of exon boundaries, while SSAHA and MegaBLAST missed the majority of exon boundaries. SSAHA consistently reported the fewest false positives and is the fastest algorithm. MegaBLAST was comparable to BLAT in speed, but was the most susceptible to localizing sequence tags incorrectly to pseudogenes. CONCLUSION: The differences in performance for sequence tags and full-length reference sequences were surprisingly small. Characteristic variations in localization results for each program were noted that affect the localization of sequence at exon boundaries, in particular.

Algorithms↗

Comprehensive copy number profiles of breast cancer cell model genomes.

INTRODUCTION: Breast cancer is the most commonly diagnosed cancer in women worldwide and consequently has been extensively investigated in terms of histopathology, immunochemistry and familial history. Advances in genome-wide approaches have contributed to molecular classification with respect to genomic changes and their subsequent effects on gene expression. Cell lines have provided a renewable resource that is readily used as model systems for breast cancer cell biology. A thorough characterization of their genomes to identify regions of segmental DNA loss (potential tumor-suppressor-containing loci) and gain (potential oncogenic loci) would greatly facilitate the interpretation of biological data derived from such cells. In this study we characterized the genomes of seven of the most commonly used breast cancer model cell lines at unprecedented resolution using a newly developed whole-genome tiling path genomic DNA array. METHODS: Breast cancer model cell lines MCF-7, BT-474, MDA-MB-231, T47D, SK-BR-3, UACC-893 and ZR-75-30 were investigated for genomic alterations with the submegabase-resolution tiling array (SMRT) array comparative genomic hybridization (CGH) platform. SMRT array CGH provides tiling coverage of the human genome permitting break-point detection at about 80 kilobases resolution. Two novel discrete alterations identified by array CGH were verified by fluorescence in situ hybridization. RESULTS: Whole-genome tiling path array CGH analysis identified novel high-level alterations and fine-mapped previously reported regions yielding candidate genes. In brief, 75 high-level gains and 48 losses were observed and their respective boundaries were documented. Complex alterations involving multiple levels of change were observed on chromosome arms 1p, 8q, 9p, 11q, 15q, 17q and 20q. Furthermore, alignment of whole-genome profiles enabled simultaneous assessment of copy number status of multiple components of the same biological pathway. Investigation of about 60 loci containing genes associated with the epidermal growth factor family (epidermal growth factor receptor, HER2, HER3 and HER4) revealed that all seven cell lines harbor copy number changes to multiple genes in these pathways. CONCLUSION: The intrinsic genetic differences between these cell lines will influence their biologic and pharmacologic response as an experimental model. Knowledge of segmental changes in these genomes deduced from our study will facilitate the interpretation of biological data derived from such cells.

Breast Neoplasms↗

The International Gene Trap Consortium Website: a portal to all publicly available gene trap cell lines in mouse.

Gene trapping is a method of generating murine embryonic stem (ES) cell lines containing insertional mutations in known and novel genes. A number of international groups have used this approach to create sizeable public cell line repositories available to the scientific community for the generation of mutant mouse strains. The major gene trapping groups worldwide have recently joined together to centralize access to all publicly available gene trap lines by developing a user-oriented Website for the International Gene Trap Consortium (IGTC). This collaboration provides an impressive public informatics resource comprising approximately 45 000 well-characterized ES cell lines which currently represent approximately 40% of known mouse genes, all freely available for the creation of knockout mice on a non-collaborative basis. To standardize annotation and provide high confidence data for gene trap lines, a rigorous identification and annotation pipeline has been developed combining genomic localization and transcript alignment of gene trap sequence tags to identify trapped loci. This information is stored in a new bioinformatics database accessible through the IGTC Website interface. The IGTC Website (www.genetrap.org) allows users to browse and search the database for trapped genes, BLAST sequences against gene trap sequence tags, and view trapped genes within biological pathways. In addition, IGTC data have been integrated into major genome browsers and bioinformatics sites to provide users with outside portals for viewing this data. The development of the IGTC Website marks a major advance by providing the research community with the data and tools necessary to effectively use public gene trap resources for the large-scale characterization of mammalian gene function.

Animals↗