Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Identification of a novel gene family that includes the interferon-inducible human genes 6-16 and ISG12.

BACKGROUND: The human 6-16 and ISG12 genes are transcriptionally upregulated in a variety of cell types in response to type I interferon (IFN). The predicted products of these genes are small (12.9 and 11.5 kDa respectively), hydrophobic proteins that share 36% overall amino acid identity. Gene disruption and over-expression studies have so far failed to reveal any biochemical or cellular roles for these proteins. RESULTS: We have used in silico analyses to identify a novel family of genes (the ISG12 gene family) related to both the human 6-16 and ISG12 genes. Each ISG12 family member codes for a small hydrophobic protein containing a conserved ~80 amino-acid motif (the ISG12 motif). So far we have detected 46 family members in 25 organisms, ranging from unicellular eukaryotes to humans. Humans have four ISG12 genes: the 6-16 gene at chromosome 1p35 and three genes (ISG12(a), ISG12(b) and ISG12(c)) clustered at chromosome 14q32. Mice have three family members (ISG12(a), ISG12(b1) and ISG12(b2)) clustered at chromosome 12F1 (syntenic with human chromosome 14q32). There does not appear to be a murine 6-16 gene. On the basis of phylogenetic analyses, genomic organisation and intron-alignments we suggest that this family has arisen through divergent inter- and intra-chromosomal gene duplication events. The transcripts from human and mouse genes are detectable, all but two (human ISG12(b) and ISG12(c)) being upregulated in response to type I IFN in the cell lines tested. CONCLUSIONS: Members of the eukaryotic ISG12 gene family encode a small hydrophobic protein with at least one copy of a newly defined motif of approximately 80 amino-acids (the ISG12 motif). In higher eukaryotes, many of the genes have acquired a responsiveness to type I IFN during evolution suggesting that a role in resisting cellular or environmental stress may be a unifying property of all family members. Analysis of gene-function in higher eukaryotes is complicated by the possibility of functional redundancy between family-members. Genetic studies in organisms (e.g. Dictyostelium discoideum) with just one family member so far identified may be particularly helpful in this respect.

Amino Acid Sequence↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

Sequencing of 42kb of the APO E-C2 gene cluster reveals a new gene: PEREC1.

Through the sequencing of a 42kb cosmid clone we describe a new gene, designated PEREC1, located approximately 1.5kb centromeric of the human apolipoprotein (APO) E-C2 cluster. The combination of dotplot analysis, predicted coding potential and interrogation of the Expressed Sequence Tag (EST) database determined the genomic organisation of PEREC1. Sequence alignment with multiple overlapping ESTs confirmed the predicted splice sites. The predicted cDNA and amino acid sequences of PEREC1 have extensive similarity to the Caenorhabditis elegans protein, C18E9.6. Conserved structural and functional motifs have been defined by combining nucleotide and amino acid analyses to identify third base degeneracy and therefore selection at the protein level. The Poliovirus Receptor Related Protein2 gene (PRR2), previously mapped to chromosome 19q13.2 by Fluorescent In-Situ Hybridisation, has also been located approximately 17kb centromeric of APO E.

Alzheimer Disease↗

Automatic detection of conserved gene clusters in multiple genomes by graph comparison and P-quasi grouping.

We previously reported two graph algorithms for analysis of genomic information: a graph comparison algorithm to detect locally similar regions called correlated clusters and an algorithm to find a graph feature called P-quasi complete linkage. Based on these algorithms we have developed an automatic procedure to detect conserved gene clusters and align orthologous gene orders in multiple genomes. In the first step, the graph comparison is applied to pairwise genome comparisons, where the genome is considered as a one-dimensionally connected graph with genes as its nodes, and correlated clusters of genes that share sequence similarities are identified. In the next step, the P-quasi complete linkage analysis is applied to grouping of related clusters and conserved gene clusters in multiple genomes are identified. In the last step, orthologous relations of genes are established among each conserved cluster. We analyzed 17 completely sequenced microbial genomes and obtained 2313 clusters when the completeness parameter P: was 40%. About one quarter contained at least two genes that appeared in the metabolic and regulatory pathways in the KEGG database. This collection of conserved gene clusters is used to refine and augment ortholog group tables in KEGG and also to define ortholog identifiers as an extension of EC numbers.

Algorithms↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

The PIR-International Protein Sequence Database.

The Protein Information Resource (PIR; http://www-nbrf.georgetown. edu/pir/) supports research on molecular evolution, functional genomics, and computational biology by maintaining a comprehensive, non-redundant, well-organized and freely available protein sequence database. Since 1988 the database has been maintained collaboratively by PIR-International, an international association of data collection centers cooperating to develop this resource during a period of explosive growth in new sequence data and new computer technologies. The PIR Protein Sequence Database entries are classified into superfamilies, families and homology domains, for which sequence alignments are available. Full-scale family classification supports comparative genomics research, aids sequence annotation, assists database organization and improves database integrity. The PIR WWW server supports direct on-line sequence similarity searches, information retrieval, and knowledge discovery by providing the Protein Sequence Database and other supplementary databases. Sequence entries are extensively cross-referenced and hypertext-linked to major nucleic acid, literature, genome, structure, sequence alignment and family databases. The weekly release of the Protein Sequence Database can be accessed through the PIR Web site. The quarterly release of the database is freely available from our anonymous FTP server and is also available on CD-ROM with the accompanying ATLAS database search program.

Amino Acid Sequence↗

SwissRegulon: a database of genome-wide annotations of regulatory sites.

SwissRegulon (http://www.swissregulon.unibas.ch) is a database containing genome-wide annotations of regulatory sites in the intergenic regions of genomes. The regulatory site annotations are produced using a number of recently developed algorithms that operate on multiple alignments of orthologous intergenic regions from related genomes in combination with, whenever available, known sites from the literature, and ChIP-on-chip binding data. Currently SwissRegulon contains annotations for yeast and 17 prokaryotic genomes. The database provides information about the sequence, location, orientation, posterior probability and, whenever available, binding factor of each annotated site. To enable easy viewing of the regulatory site annotations in the context of other features annotated on the genomes, the sites are displayed using the GBrowse genome browser interface and can be queried based on any annotated genomic feature. The database can also be queried for regulons, i.e. sites bound by a common factor.

Algorithms↗

Efficient parameterized algorithms for biopolymer structure-sequence alignment.

Computational alignment of a biopolymer sequence (e.g., an RNA or a protein) to a structure is an effective approach to predict and search for the structure of new sequences. To identify the structure of remote homologs, the structure-sequence alignment has to consider not only sequence similarity, but also spatially conserved conformations caused by residue interactions and, consequently, is computationally intractable. It is difficult to cope with the inefficiency without compromising alignment accuracy, especially for structure search in genomes or large databases. This paper introduces a novel method and a parameterized algorithm for structure-sequence alignment. Both the structure and the sequence are represented as graphs, where, in general, the graph for a biopolymer structure has a naturally small tree width. The algorithm constructs an optimal alignment by finding in the sequence graph the maximum valued subgraph isomorphic to the structure graph. It has the computational time complexity O[k(t)N(2)] for the structure of N residues and its tree decomposition of width t. Parameter k, small in nature, is determined by a statistical cutoff for the correspondence between the structure and the sequence. This paper demonstrates a successful application of the algorithm to RNA structure search used for noncoding RNA identification. An application to protein threading is also discussed.

Algorithms↗

Genome-specific primer sets for starch biosynthesis genes in wheat.

Common wheat (Triticum aestivum L.,2n=6x=42) is an allohexaploid composed of three closely related genomes, designated A, B, and D. Genetic analysis in wheat is complicated, as most genes are present in triplicated sets located in the same chromosomal regions of homoeologous chromosomes. The goal of this report was to use genomic information gathered from wheat-rice sequence comparison to develop genome-specific primer sets for five genes involved in starch biosynthesis. Intron locations in wheat were inferred through the alignment of wheat cDNA sequences with rice genomic sequence.Exon-anchored primers, which amplify across introns,allowed the sequencing of introns from the three genomes for each gene. Sequence variation within introns among the three wheat genomes provided the basis for genome-specific primer design. For three genes, ADP-glucose pyrophosphorylase (Agp-L), sucrose transporter (SUT),and waxy (Wx), genome-specific primer sets were developed for all three genomes. Genome-specific primers were developed for two of the three genomes for Agp-S and starch synthase I (Ssl). Thus, 13 of 15 possible genome-specific primer sets were developed using this strategy. Seven genome-specific primer combinations were used to amplify alleles in hexaploid wheat lines for sequence comparison. Three single nucleotide polymorphisms(SNPs) were identified in a comparison of 5,093 bp among a minimum of ten wheat accessions. Two of theseSNPs could be converted into cleaved amplified polymorphism sequence (CAPS) markers. Our results indicated that the design of genome-specific primer sets using intron-based sequence differences has a high probability of success, while the identification of polymorphism among alleles within a genome may be a challenge.

Base Sequence↗

The Atlantic salmon prepro-gonadotropin releasing hormone gene and mRNA.

Screening for the gene encoding salmon gonadotropin releasing hormone (sGnRH) in an Atlantic salmon (Salmo salar) genomic library resulted in isolation of a positive clone designated lambda sGnRH-1. An anchor polymerase chain reaction (PCR) technique was used to amplify GnRH cDNA derived from salmon hypothalamic mRNA. The cDNA sequence was aligned to the 7607 base pair genomic sequence which was shown to encode the entire prepro-GnRH gene. The cDNA proved that the cloned gene is expressed in the hypothalamus of mature salmon. The coding domain of sGnRH differs from the mammalian GnRH by six nucleotide changes which allow the two amino acid differences between the two GnRH variants. Salmon GnRH associated peptide (GAP) differs extensively in sequence and size from the mammalian counterpart. Compared to the GnRH cDNA of a cichlid species the similarity is 69.3% in the protein coding sequence.

Amino Acid Sequence↗

ShiBASE: an integrated database for comparative genomics of Shigella.

Among the major enteric bacterial pathogens, Shigella is found to display extreme genome diversity and dynamics, which imposes a challenge in comparative genomic studies. To facilitate further studies in this area, we have constructed an integrated online database, ShiBASE (http://www.mgc.ac.cn/ShiBASE/),which contains Shigella genomic sequences of four species and additional comparative genomic hybridization (CGH) data of 43 serotypes. ShiBASE offers online comparative analysis on DNA sequences, gene orders, metabolic pathways and virulence factors. In addition, ShiBASE has a newly developed online comparative visualization service, Shi-align, which enables the alignment of any query sequence with the reference genome sequences.

Databases, Nucleic Acid↗

From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.

A database of functional sites for proteins with known structures, SITE, is constructed and used in conjunction with a simple pattern matching program SiteMatch to evaluate possible function conservation in a recently constructed database of fold predictions for Escherichia coli proteins (Rychlewski L et al., 1999, Protein Sci 8:614-624). In this and other prediction databases, fold predictions are based on algorithms that can recognize weak sequence similarities and putatively assign new proteins into already characterized protein families. It is not clear whether such sequence similarities arise from distant homologies or general similarity of physicochemical features along the sequence. Leaving aside the important question of nature of relations within fold superfamilies, it is possible to assess possible function conservation by looking at the pattern of conservation of crucial functional residues. SITE consists of a multilevel function description based on structure annotations and structure analyses. In particular, active site residues, ligand binding residues, and patterns of hydrophobic residues on the protein surface are used to describe different functional features. SiteMatch, a simple pattern matching program, is designed to check the conservation of residues involved in protein activity in alignments generated by any alignment method. Here, this procedure is used to study conservation of functional features in alignments between protein sequences from the E. coli genome and their optimal structural templates. The optimal templates were identified and alignments taken from the database of genomic structural predictions was described in a previous publication (Rychlewski L et al., 1999, Protein Sci 8:614-624). An automated assessment of function conservation is used to analyze the relation between fold and function similarity for a large number of fold predictions. For instance, it is shown that identifying low significance predictions with a high level of functional residue conservations can be used to extend the prediction sensitivity for fold prediction methods. Over 100 new fold/function predictions in this class were obtained in the E. coli genome. At the same time, about 30% of our previous fold predictions are not confirmed as function predictions, further highlighting the problem of function divergence in fold superfamilies.

Algorithms↗

Tree decomposition based fast search of RNA structures including pseudoknots in genomes.

Searching genomes for RNA secondary structure with computational methods has become an important approach to the annotation of non-coding RNAs. However, due to the lack of efficient algorithms for accurate RNA structure-sequence alignment, computer programs capable of fast and effectively searching genomes for RNA secondary structures have not been available. In this paper, a novel RNA structure profiling model is introduced based on the notion of a conformational graph to specify the consensus structure of an RNA family. Tree decomposition yields a small tree width t for such conformation graphs (e.g., t = 2 for stem loops and only a slight increase for pseudo-knots). Within this modelling framework, the optimal alignment of a sequence to the structure model corresponds to finding a maximum valued isomorphic subgraph and consequently can be accomplished through dynamic programming on the tree decomposition of the conformational graph in time O(k(t)N(2)), where k is a small parameter; and N is the size of the projiled RNA structure. Experiments show that the application of the alignment algorithm to search in genomes yields the same search accuracy as methods based on a Covariance model with a significant reduction in computation time. In particular; very accurate searches of tmRNAs in bacteria genomes and of telomerase RNAs in yeast genomes can be accomplished in days, as opposed to months required by other methods. The tree decomposition based searching tool is free upon request and can be downloaded at our site h t t p ://w.uga.edu/RNA-informatics/software/index.php.

Algorithms↗

A method for finding single-nucleotide polymorphisms with allele frequencies in sequences of deep coverage.

BACKGROUND: The allele frequencies of single-nucleotide polymorphisms (SNPs) are needed to select an optimal subset of common SNPs for use in association studies. Sequence-based methods for finding SNPs with allele frequencies may need to handle thousands of sequences from the same genome location (sequences of deep coverage). RESULTS: We describe a computational method for finding common SNPs with allele frequencies in single-pass sequences of deep coverage. The method enhances a widely used program named PolyBayes in several aspects. We present results from our method and PolyBayes on eighteen data sets of human expressed sequence tags (ESTs) with deep coverage. The results indicate that our method used almost all single-pass sequences in computation of the allele frequencies of SNPs. CONCLUSION: The new method is able to handle single-pass sequences of deep coverage efficiently. Our work shows that it is possible to analyze sequences of deep coverage by using pairwise alignments of the sequences with the finished genome sequence, instead of multiple sequence alignments.

Computers, Molecular↗

SegMantX: A Novel Tool for Detecting DNA Duplications Uncovers Prevalent Duplications in Plasmids.

Segmental duplications play an important role in genome evolution via their contribution to copy-number variation, gene-family diversification, and the emergence of novel functions. The detection of segmental duplications is challenging due to heterogeneous amelioration of sequence similarity among duplicates, which hinders the reconstruction of continuous sequence alignment. Here we introduce SegMantX, a novel approach for the identification of diverged segmental duplications in prokaryote genomes using local alignment chaining. In this approach, local alignments resulting from a preliminary sequence similarity search (e.g. BLASTn) are chained into continuous segments. Evaluating the performance of SegMantX using simulated sequences shows that the tool can detect diverged duplications beyond the sensitivity limits of standard alignment-based methods. Applying SegMantX to 6,784 enterobacterial plasmids, we find that 65% plasmids contain duplicated regions and gene duplications, most of which correspond either to dispersed, noncoding regions or duplicated mobile genetic elements (MGEs; e.g. transposons and insertion sequences). Furthermore, we demonstrate the applicability of SegMantX for the identification of diverged gene transfers between replicons and plasmid hybridization events. Our findings highlight MGEs as drivers of segmental duplications in plasmid evolution, leading to the amplification of their cargo genes, including antibiotic resistance genes. SegMantX provides a powerful framework for reconstructing diverged segmental duplications and other alignment problems.

Plasmids↗

Family business: the multidrug-resistance related protein (MRP) ABC transporter genes in Arabidopsis thaliana.

Despite the completion of the sequencing of the entire genome of Arabidopsis thaliana (L.) Heynh., the exact determination of each single gene and its function remains an open question. This is especially true for multigene families. An approach that combines analysis of genomic structure, expression data and functional genomics to ascertain the role of the members of the multidrug-resistance-related protein ( MRP) gene family, a subfamily of the ATP-binding cassette (ABC) transporters from Arabidopsis is presented. We used cDNA sequencing and alignment-based re-annotation of genomic sequences to define the exact genic structure of all known AtMRP genes. Analysis of promoter regions suggested different induction conditions even for closely related genes. Expression analysis for the entire gene family confirmed these assumptions. Phylogenetic analysis and determination of segmental duplication in the regions of AtMRP genes revealed that the evolution of the extraordinarily high number of ABC transporter genes in plants cannot solely be explained by polyploidisation during the evolution of the Arabidopsis genome. Interestingly MRP genes from Oryza sativa L. (rice; OsMRP) show very similar genomic structures to those from Arabidopsis. Screening of large populations of T-DNA-mutagenised lines of A. thaliana resulted in the isolation of AtMRP insertion mutants. This work opens the way for the defined analysis of a multigene family of important membrane transporters whose broad variety of functions expands their traditional role as cellular detoxifiers.

ATP-Binding Cassette Transporters↗

Identification of the gene encoding the sole physiological fumarate reductase in Shewanella oneidensis MR-1.

Shewanella oneidensis MR-1 is a Gram-negative, nonfermentative rod with a complex electron transport system which facilitates its ability to use a variety of terminal electron acceptors, including fumarate, for anaerobic respiration. CMTn-3, a mutant isolated by transposon (TnphoA) mutagenesis, can no longer use fumarate as an electron acceptor; it lacks fumarate reductase activity as well as a 65-kDa soluble tetraheme flavocytochrome c. The sequence of the TnphoA-flanking genomic DNA of CMTn-3 did not align to those for fumarate reductase or related electron transport genes from other bacteria. Sequence analysis of the MR-1 genomic database demonstrated that an open reading frame encoding a 65-kDa tetraheme cytochrome c with sequence similarity to the fumarate reductase from S. frigidimarina NCIMB400 was found 8 kb away from the TnphoA-flanking genomic DNA of CMTn-3. PCR analysis demonstrated that a large deletion (>or=9.2 kb and <or=11 kb) of genomic DNA occurred in CMTn-3 as a result of TnphoA insertion. This deletion included at least half of the fumarate reductase gene as well as approximately 8 kb of upstream DNA. Complementation of CMTn-3 with the fumarate reductase gene plus 0.5-kb of upstream DNA restored growth on fumarate. These studies explicitly define the sole physiological fumarate reductase gene from the several possibilities suggested by the genomic sequence of MR-1. Surprisingly, the fumarate reductase gene plus 0.77-kb upstream DNA from S. frigidimarina NCIMB400 did not complement CMTn-3.

DNA Transposable Elements↗

Expression profiling of the Leishmania life cycle: cDNA arrays identify developmentally regulated genes present but not annotated in the genome.

As genomic sequencing of Leishmania nears completion, functional analyses that provide a global genetic perspective on biological processes are important. Despite polycistronic transcription, RNA transcript abundance can be measured using microarrays. To provide a resource to evaluate cDNA arrays, we undertook 5' expressed sequence tag analysis of 2183 full-length randomly selected cDNAs from Leishmania major promastigote (days 3, 7, 10 of culture in vitro), and lesion-derived amastigote libraries. PCR-amplified inserts from 1830 of these cDNA representing 1001 unique genes were spotted onto microarrays, and compared internally with PCR-amplified open reading frames (ORFs) from 904 genes representing 842 unique genes annotated in the L. major genome. Microarrays were screened with RNA from procyclic, metacyclic and amastigote populations of L. major. Redundant clones on the array gave highly reproducible results, providing confidence in identification of stage-specific gene expression. Four hundred and thirty unique (i.e. non-redundant) stage-specific genes were identified. A higher percentage of stage-specific gene expression was observed in amastigotes ( approximately 35%) compared to metacyclics ( approximately 12%) for both cDNAs and ORFs, but cDNAs provided a richer source of regulated genes than currently annotated ORFs from the Leishmania genome. In mapping cDNAs onto the Leishmania genome, we noted that approximately 42% aligned to regions not recognised as genes using current predictive annotation tools. These genes are highly represented in our stage-specific genes, and therefore represent important drug targets and vaccine candidates. Careful annotation of cDNAs onto the Leishmania genome will be important before producing the next generation of oligonucleotide arrays based on annotated genes of the genomic sequencing project.

Animals↗