Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Cloning and expression in phospholipid containing cultures of the gene encoding the specific phosphatidylglycerol/phosphatidylinositol transfer protein from Aspergillus oryzae: evidence that the pg/pi-tp is tandemly arranged with the putative 3-ketoacyl-CoA thiolase gene.

The phosphatidylglycerol/phosphatidylinositol transfer protein (PG/PI-TP) is a new and original phospholipid transfer protein (PLTP) isolated from the Deuteromycete, Aspergillus oryzae. We have isolated a genomic clone of the A. oryzae pg/pi-tp using a probe derived from the corresponding cDNA and sequenced the complete gene. The DNA sequence analysis revealed that pg/pi-tp gene is composed of three exons encoding a 18,823 Da protein of 175 amino acids as previously described and of two introns as deduced by cDNA and genomic sequence alignment. The isolated pg/pi-tp gene do not show similarity with other PLTP genes or the deduced PG/PI-TP protein with proteins already known. Comparison of the encoded PG/PI-TP with other deduced proteins from recent genomic or cDNA sequence from databases revealed that the PG/PI-TP was close to two encoded proteins deduced from the cDNA database of Aspergillus nidulans (54% identity and 68% similarity) and the second from Neurospora crassa (53% identity and 76% similarity). Therefore, we suggested that both proteins might belong to the PLTP family. Southern blot analysis of A. oryzae genomic DNA show that the PG/PI-TP was encoded by a single gene. Expression of pg/pi-tp was performed in phospholipid containing cultures with increasing carbon source concentrations in order to study the regulation of the PLTPs in the filamentous fungus cell. This was done to know if a high density culture could yield a high amount of biomass with high phospholipid transfer activity. Results showed that phospholipids as compared to glucose in standard cultures stimulated mycelial growth and global phospholipid transfer activity, but not the pg/pi-tp transcript accumulation. However, high concentration of both carbon sources yielded an inhibition of the expression of the pg/pi-tp gene and of the global phospholipid transfer activity. In conclusion, both carbon sources are not suitable to increase the PLTP production in high density cultures for biotechnological applications. Finally, using the gene walking sequencing method it is demonstrated that the pg/pi-tp is tandemly arranged on opposite DNA strands in a tail-to-tail orientation with a putative gene encoding the 3-ketoacyl-CoA thiolase (EC 2.3.1.16). Unlike the pg/pi-tp gene, this thiolase gene show a putative 'beta-oxidation box' and encodes a putative 44,150 Da protein of 321 amino acids composed of a putative N-terminal PTS2 (Peroxisomal Targeting Signal) consensus sequence for the peroxisome targeting. Comparison of the amino acid sequence of the A. oryzae thiolase to that of the Yarrowia lipolytica showed a 50% identity and a 69% similarity.

Acetyl-CoA C-Acyltransferase↗

Generation of an integrated transcription map of the BRCA2 region on chromosome 13q12-q13.

An integrated approach involving physical mapping, identification of transcribed sequences, and computational analysis of genomic sequence was used to generate a detailed transcription map of the 1. 0-Mb region containing the breast cancer susceptibility locus BRCA2 on chromosome 13q12-q13. This region is included in the genetic interval bounded by D13S1444 and D13S310. Retrieved sequences from exon amplification or hybrid selection procedures were grouped into physical intervals and subsequently grouped into transcription units by clone overlap. Overlap was established by direct hybridization, cDNA library screening, PCR cDNA linking (island hopping), and/or sequence alignment. Extensive genomic sequencing was performed in an effort to understand transcription unit organization. In total, approximately 500 kb of genomic sequence was completed. The transcription units were further characterized by hybridization to RNA from a series of human tissues. Evidence for seven genes, two putative pseudogenes, and nine additional putative transcription units was obtained. One of the transcription units was recently identified as BRCA2 but all others are novel genes of unknown function as only limited alignment to sequences in public databases was observed. One large gene with a transcript size of 10.7 kb showed significant similarity to a gene predicted by the Caenorhabditis elegans genome and the Saccharomyces cerevisiae genome sequencing efforts, while another contained a motif sequence similar to the human 2',3' cyclic nucleotide 3' phosphodiesterase gene. Several retrieved transcribed sequences were not aligned into transcription units because no corresponding cDNAs were obtained when screening libraries or because of a lack of definitive evidence for splicing signals or putative coding sequence based on computational analysis. However, the presence of additional genes in the BRCA2 interval is suggested as groups of putative exons and hybrid selected clones that were transcribed in consistent orientations could be localized to common physical intervals.

BRCA2 Protein↗

Gene structure prediction from consensus spliced alignment of multiple ESTs matching the same genomic locus.

MOTIVATION: Accurate gene structure annotation is a challenging computational problem in genomics. The best results are achieved with spliced alignment of full-length cDNAs or multiple expressed sequence tags (ESTs) with sufficient overlap to cover the entire gene. For most species, cDNA and EST collections are far from comprehensive. We sought to overcome this bottleneck by exploring the possibility of using combined EST resources from fairly diverged species that still share a common gene space. Previous spliced alignment tools were found inadequate for this task because they rely on very high sequence similarity between the ESTs and the genomic DNA. RESULTS: We have developed a computer program, GeneSeqer, which is capable of aligning thousands of ESTs with a long genomic sequence in a reasonable amount of time. The algorithm is uniquely designed to tolerate a high percentage of mismatches and insertions or deletions in the EST relative to the genomic template. This feature allows use of non-cognate ESTs for gene structure prediction, including ESTs derived from duplicated genes and homologous genes from related species. The increased gene prediction sensitivity results in part from novel splice site prediction models that are also available as a stand-alone splice site prediction tool. We assessed GeneSeqer performance relative to a standard Arabidopsis thaliana gene set and demonstrate its utility for plant genome annotation. In particular, we propose that this method provides a timely tool for the annotation of the rice genome, using abundant ESTs from other cereals and plants. AVAILABILITY: The source code is available for download at http://bioinformatics.iastate.edu/bioinformatics2go/gs/download.html. Web servers for Arabidopsis and other plant species are accessible at http://www.plantgdb.org/cgi-bin/AtGeneSeqer.cgi and http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi, respectively. For non-plant species, use http://bioinformatics.iastate.edu/cgi-bin/gs.cgi. The splice site prediction tool (SplicePredictor) is distributed with the GeneSeqer code. A SplicePredictor web server is available at http://bioinformatics.iastate.edu/cgi-bin/sp.cgi

Algorithms↗

Genome-wide detection and analysis of homologous recombination among sequenced strains of Escherichia coli.

BACKGROUND: Comparisons of complete bacterial genomes reveal evidence of lateral transfer of DNA across otherwise clonally diverging lineages. Some lateral transfer events result in acquisition of novel genomic segments and are easily detected through genome comparison. Other more subtle lateral transfers involve homologous recombination events that result in substitution of alleles within conserved genomic regions. This type of event is observed infrequently among distantly related organisms. It is reported to be more common within species, but the frequency has been difficult to quantify since the sequences under comparison tend to have relatively few polymorphic sites. RESULTS: Here we report a genome-wide assessment of homologous recombination among a collection of six complete Escherichia coli and Shigella flexneri genome sequences. We construct a whole-genome multiple alignment and identify clusters of polymorphic sites that exhibit atypical patterns of nucleotide substitution using a random walk-based method. The analysis reveals one large segment (approximately 100 kb) and 186 smaller clusters of single base pair differences that suggest lateral exchange between lineages. These clusters include portions of 10% of the 3,100 genes conserved in six genomes. Statistical analysis of the functional roles of these genes reveals that several classes of genes are over-represented, including those involved in recombination, transport and motility. CONCLUSION: We demonstrate that intraspecific recombination in E. coli is much more common than previously appreciated and may show a bias for certain types of genes. The described method provides high-specificity, conservative inference of past recombination events.

Alleles↗

Exon-intron organization of TRGC genes in sheep.

A series of genomic clones derived from a sheep library were used to determine the germline configuration and the exon-intron organization of TRGC2, TRGC3, and TRGC4 genes. Based on the outcomes of molecular analysis, we compared and aligned the genomic sequences with the known complete cDNA sequences of sheep and deduced the exon-intron organization of TRGC genes in this ruminant animal, EX1, corresponding to the disulfide-linked constant domain, and EX3, corresponding to the transmembrane and cytoplasmatic domains, are similar in length in all genes. Conversely, the hinge-encoding EX2A, EX2B, and EX2C exons differ in number and length between genes, and EX2A contains the TTKPP motif irrespective of whether it occurs in single or triplicate form. The molecular data also indicate that at least one additional gene is present in sheep. Phylogenetic analysis grouped the ruminant TRGC genes in two clusters that could have emerged from two ancestral forms that underwent a series of duplications giving rise to the new sequences that were selected and then fixed in the ruminant lineages. A correlation between the cluster distribution in the phylogenetic tree of TRGC genes and their expression during fetal development is discussed.

Amino Acid Sequence↗

Evolutionary inventions and continuity of CORE-SINEs in mammals.

We characterized short interspersed elements (SINEs), of the CORE-suprafamily in egg-laying (monotremes), pouched (marsupials) and placental mammals. Five families of these repeats distinguished by the presence of distinct LINE-related 3'-segments shared tRNA-like promoter and the central core region. The putative active elements were reconstructed from the alignment of genomic repeats representing molecular fossils of sequences that amplified in the past and since then underwent multiple mutations. Their mode of proliferation by retroposition was indicated by the presence of: (1) internal RNA PolIII promoter; (2) simple sequence repeated tail; (3) direct repeats; and (4) subfamilies recording the evolution of elements. The copy number of CORE-SINEs in placental genomes was estimated at about 300,000; they were highly divergent and apparently ceased to amplify before radiation of these lineages. On the other hand, among almost half a million fossil elements present in marsupials and monotremes, the youngest subfamilies could still be retropositionally active. CORE-SINEs terminate in sequence repeats of a few nucleotides similar to their 3'-segment LINE-homologues, CR1, L2 and Bov-B. These three LINE elements fall into clades distinct from that of L1 elements which, similar to their co-amplifying SINEs, end in a poly(A) tail. We propose a model in which new CORE-families, with distinct 3'-segments, are created at the RNA level due to template switching between LINE and CORE-RNA during reverse transcription. The proposed mechanism suggests that such an adaptation to the changing amplification machinery facilitated the survival and prosperity of CORE-elements over long evolutionary periods in different lineages.

Animals↗

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals↗

A novel method for multiple alignment of sequences with repeated and shuffled elements.

We describe ABA (A-Bruijn alignment), a new method for multiple alignment of biological sequences. The major difference between ABA and existing multiple alignment methods is that ABA represents an alignment as a directed graph, possibly containing cycles. This representation provides more flexibility than does a traditional alignment matrix or the recently introduced partial order alignment (POA) graph by allowing a larger class of evolutionary relationships between the aligned sequences. Our graph representation is particularly well-suited to the alignment of protein sequences with shuffled and/or repeated domain structure, and allows one to construct multiple alignments of proteins containing (1) domains that are not present in all proteins, (2) domains that are present in different orders in different proteins, and (3) domains that are present in multiple copies in some proteins. In addition, ABA is useful in the alignment of genomic sequences that contain duplications and inversions. We provide several examples illustrating the applications of ABA.

Algorithms↗

Primate molecular divergence dates.

With genomic data, alignments can be assembled that greatly increase the number of informative sites for analysis of molecular divergence dates. Here, we present an estimate of the molecular divergence dates for all of the major primate groups. These date estimates are based on a Bayesian analysis of approximately 59.8 kbp of genomic data from 13 primates and 6 mammalian outgroups, using a range of paleontologically supported calibration estimates. Results support a Cretaceous last common ancestor of extant primates (approximately 77 mya), an Eocene divergence between platyrrhine and catarrhine primates (approximately 43 mya), an Oligocene origin of apes and Old World monkeys (approximately 31 mya), and an early Miocene (approximately 18 mya) divergence of Asian and African great apes. These dates are examined in the context of other molecular clock studies.

Animals↗

Utilizing evolutionary conservation to detect deleterious mutations and improve genomic prediction in cassava.

INTRODUCTION: Cassava (Manihot esculenta) is an annual root crop which provides the major source of calories for over half a billion people around the world. Since its domestication ~10,000 years ago, cassava has been largely clonally propagated through stem cuttings. Minimal sexual recombination has led to an accumulation of deleterious mutations made evident by heavy inbreeding depression. METHODS: To locate and characterize these deleterious mutations, and to measure selection pressure across the cassava genome, we aligned 52 related Euphorbiaceae and other related species representing millions of years of evolution. With single base-pair resolution of genetic conservation, we used protein structure models, amino acid impact, and evolutionary conservation across the Euphorbiaceae to estimate evolutionary constraint. With known deleterious mutations, we aimed to improve genomic evaluations of plant performance through genomic prediction. We first tested this hypothesis through simulation utilizing multi-kernel GBLUP to predict simulated phenotypes across separate populations of cassava. RESULTS: Simulations showed a sizable increase of prediction accuracy when incorporating functional variants in the model when the trait was determined by<100 quantitative trait loci (QTL). Utilizing deleterious mutations and functional weights informed through evolutionary conservation, we saw improvements in genomic prediction accuracy that were dependent on trait and prediction. CONCLUSION: We showed the potential for using evolutionary information to track functional variation across the genome, in order to improve whole genome trait prediction. We anticipate that continued work to improve genotype accuracy and deleterious mutation assessment will lead to improved genomic assessments of cassava clones.

cassava (Manihot esculenta)↗

Identifying multiple alignment regions satisfying simple formulas and patterns.

MOTIVATION: When studying multiple alignments of genomic sequences one frequently aims to locate and count regions which satisfy a set of constraints. These regions may be putatively functional, but researchers may also be interested in quantifying the frequency of occurrences of certain patterns. RESULTS: We have developed a program that applies simple formulas and pattern specifications to multiple alignments, reporting the positions and counts of conforming regions. As an example, we have navigated a 15-species alignment of the CAV2-CAV1 region and outlined some findings regarding PPARgamma binding sites. AVAILABILITY: Our software and the accompanying documentation can be obtained at no charge by contacting the authors. It can also be accessed at http://ranger.uta.edu/~nick/compgen

Algorithms↗

An algorithm for linear metabolic pathway alignment.

Metabolic pathway alignment represents one of the most powerful tools for comparative analysis of metabolism. It involves recognition of metabolites common to a set of functionally-related metabolic pathways, interpretation of biological evolution processes and determination of alternative metabolic pathways. Moreover, it is of assistance in function prediction and metabolism modeling. Although research on genomic sequence alignment is extensive, the problem of aligning metabolic pathways has received less attention. We are motivated to develop an algorithm of metabolic pathway alignment to reveal the similarities between metabolic pathways. A new definition of the metabolic pathway is introduced. The algorithm has been implemented into the PathAligner system; its web-based interface is available at http://bibiserv.techfak.uni-bielefeld.de/pathaligner/.

Algorithms↗

Quality assessment of maize assembled genomic islands (MAGIs) and large-scale experimental verification of predicted genes.

Recent sequencing efforts have targeted the gene-rich regions of the maize (Zea mays L.) genome. We report the release of an improved assembly of maize assembled genomic islands (MAGIs). The 114,173 resulting contigs have been subjected to computational and physical quality assessments. Comparisons to the sequences of maize bacterial artificial chromosomes suggest that at least 97% (160 of 165) of MAGIs are correctly assembled. Because the rates at which junction-testing PCR primers for genomic survey sequences (90-92%) amplify genomic DNA are not significantly different from those of control primers ( approximately 91%), we conclude that a very high percentage of genic MAGIs accurately reflect the structure of the maize genome. EST alignments, ab initio gene prediction, and sequence similarity searches of the MAGIs are available at the Iowa State University MAGI web site. This assembly contains 46,688 ab initio predicted genes. The expression of almost half (628 of 1,369) of a sample of the predicted genes that lack expression evidence was validated by RT-PCR. Our analyses suggest that the maize genome contains between approximately 33,000 and approximately 54,000 expressed genes. Approximately 5% (32 of 628) of the maize transcripts discovered do not have detectable paralogs among maize ESTs or detectable homologs from other species in the GenBank NR nucleotide/protein database. Analyses therefore suggest that this assembly of the maize genome contains approximately 350 previously uncharacterized expressed genes. We hypothesize that these "orphans" evolved quickly during maize evolution and/or domestication.

Chromosomes, Artificial, Bacterial↗

A high-resolution 6.0-megabase transcript map of the type 2 diabetes susceptibility region on human chromosome 20.

Recent linkage studies and association analyses indicate the presence of at least one type 2 diabetes susceptibility gene in human chromosome region 20q12-q13.1. We have constructed a high-resolution 6.0-megabase (Mb) transcript map of this interval using two parallel, complementary strategies to construct the map. We assembled a series of bacterial artificial chromosome (BAC) contigs from 56 overlapping BAC clones, using STS/marker screening of 42 genes, 43 ESTs, 38 STSs, 22 polymorphic, and 3 BAC end sequence markers. We performed map assembly with GraphMap, a software program that uses a greedy path searching algorithm, supplemented with local heuristics. We anchored the resulting BAC contigs and oriented them within a yeast artificial chromosome (YAC) scaffold by observing the retention patterns of shared markers in a panel of 21 YAC clones. Concurrently, we assembled a sequence-based map from genomic sequence data released by the Human Genome Project, using a seed-and-walk approach. The map currently provides near-continuous coverage between SGC32867 and WI-17676 ( approximately 6.0 Mb). EST database searches and genomic sequence alignments of ESTs, mRNAs, and UniGene clusters enabled the annotation of the sequence interval with experimentally confirmed and putative transcripts. We have begun to systematically evaluate candidate genes and novel ESTs within the transcript map framework. So far, however, we have found no statistically significant evidence of functional allelic variants associated with type 2 diabetes. The combination of the BAC transcript map, YAC-to-BAC scaffold, and reference Human Genome Project sequence provides a powerful integrated resource for future genomic analysis of this region.

Base Composition↗

Molecular and culture-based analyses of aerobic carbon monoxide oxidizer diversity.

Isolates belonging to six genera not previously known to oxidize CO were obtained from enrichments with aquatic and terrestrial plants. DNA from these and other isolates was used in PCR assays of the gene for the large subunit of carbon monoxide dehydrogenase (coxL). CoxL and putative coxL fragments were amplified from known CO oxidizers (e.g., Oligotropha carboxidovorans and Bradyrhizobium japonicum), from novel CO-oxidizing isolates (e.g., Aminobacter sp. strain COX, Burkholderia sp. strain LUP, Mesorhizobium sp. strain NMB1, Stappia strains M4 and M8, Stenotrophomonas sp. strain LUP, and Xanthobacter sp. strain COX), and from several well-known isolates for which the capacity to oxidize CO is reported here for the first time (e.g., Burkholderia fungorum LB400, Mesorhizobium loti, Stappia stellulata, and Stappia aggregata). PCR products from several taxa, e.g., O. carboxidovorans, B. japonicum, and B. fungorum, yielded sequences with a high degree (>99.6%) of identity to those in GenBank or genome databases. Aligned sequences formed two phylogenetically distinct groups. Group OMP contained sequences from previously known CO oxidizers, including O. carboxidovorans and Pseudomonas thermocarboxydovorans, plus a number of closely related sequences. Group BMS was dominated by putative coxL sequences from genera in the Rhizobiaceae and other alpha-PROTEOBACTERIA: PCR analyses revealed that many CO oxidizers contained two coxL sequences, one from each group. CO oxidation by M. loti, for which whole-genome sequencing has revealed a single BMS-group putative coxL gene, strongly supports the notion that BMS sequences represent functional CO dehydrogenase proteins that are related to but distinct from previously characterized aerobic CO dehydrogenases.

Aerobiosis↗

BLAT--the BLAST-like alignment tool.

Analyzing vertebrate genomes requires rapid mRNA/DNA and cross-species protein alignments. A new tool, BLAT, is more accurate and 500 times faster than popular existing tools for mRNA/DNA alignments and 50 times faster for protein alignments at sensitivity settings typically used when comparing vertebrate sequences. BLAT's speed stems from an index of all nonoverlapping K-mers in the genome. This index fits inside the RAM of inexpensive computers, and need only be computed once for each genome assembly. BLAT has several major stages. It uses the index to find regions in the genome likely to be homologous to the query sequence. It performs an alignment between homologous regions. It stitches together these aligned regions (often exons) into larger alignments (typically genes). Finally, BLAT revisits small internal exons possibly missed at the first stage and adjusts large gap boundaries that have canonical splice sites where feasible. This paper describes how BLAT was optimized. Effects on speed and sensitivity are explored for various K-mer sizes, mismatch schemes, and number of required index matches. BLAT is compared with other alignment programs on various test sets and then used in several genome-wide applications. http://genome.ucsc.edu hosts a web-based BLAT server for the human genome.

Animals↗

The first non-LTR retrotransposon characterised in the cephalochordate amphioxus, BfCR1, shows similarities to CR1-like elements.

BfCR1 is the first non-long terminal repeat retrotransposon to be characterised in the amphioxus genome. Sequence alignment of the predicted translation product reveals that BfCR1 belongs to the CR1-like retroposon class, a family widely distributed in vertebrate and invertebrate lineages. Structural analysis shows conservation of the specific motifs of the ORF2-CR1 elements: the N-terminal endonuclease, the reverse transcriptase and the C-terminal domains. The BfCR1 element possesses an atypical 3' terminus consisting of the tandem repeat (AAG)6. We gathered evidence supporting the mobility of this element and report an estimated 15 copies of BfCR1, mostly truncated, per haploid genome, a remarkably low number when compared to that of vertebrates. Phylogenetic analysis, including the amphioxus element, seems to indicate that (i) CR1-like retroposons cluster in a monophyletic group and (ii) the CR1-like family was already present in the chordate ancestor. Our data provide further support for the horizontal transmission of CR1-like elements during early vertebrate evolution.

Amino Acid Sequence↗