Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Genomic DNA sequence of the cystic fibrosis transmembrane conductance regulator (CFTR) gene.

The gene responsible for cystic fibrosis, the most common severe autosomal recessive disorder, is located on the long arm of human chromosome 7, region q31-q32. The gene has recently been identified and shown to be approximately 250 kb in size. To understand the structure and to provide the basis for a systematic analysis of the disease-causing mutations in the gene, genomic DNA clones spanning different regions of the previously reported cDNA were isolated and used to determine the coding regions and sequences of intron/exon boundaries. A total of 22,708 bp of sequence, accounting for approximately 10% of the entire gene, was obtained. Alignment of the genomic DNA sequence with the cDNA sequence showed perfect colinearity between the two and a total of 27 exons, each flanked by consensus splice signals. A number of repetitive elements, including the Alu and Kpn families and simple repeats, such as (GT)17, (GATT)7, and (TA)14, were detected in close vicinity of some of the intron/exon boundaries. At least three of the simple repeats were found to be polymorphic in the population. Although an internal amino acid sequence homology could be detected between the two halves of the predicted polypeptide, especially in the regions of the two putative nucleotide-binding folds (NBF1 and NBF2), the lack of alignment of the nucleotide sequence as well as the different positions of the exon/intron boundaries does not seem to support the hypothesis of a recent gene duplication event. To facilitate detection of mutations by direct sequence analysis of genomic DNA, 28 sets of oligonucleotide primers were designed and tested for their ability to amplify individual exons and the immediately flanking sequences in the introns.

Amino Acid Sequence↗

Bioinformatic analyses of bacterial HPr kinase/phosphorylase homologues.

HPr kinase/phosphorylases (HprKs) regulate catabolite repression and sugar transport in Gram-positive bacteria by phosphorylating the small phosphotransferase system (PTS) protein HPr on a serine residue. We identified homologues of HprK in currently sequenced genomes and multiply aligned their sequences in order to perform phylogenetic and motif analyses. Seventy-eight homologues from bacteria and one from an archaeon comprise nine phylogenetic clusters. Some homologues come from bacteria whose genomes contain multiple highly divergent paralogues that cluster loosely together. Many of these proteins are truncated or show little or no identifiable similarity outside of the Walker A nucleotide binding domain. HprK homologues were identified in Gram-negative bacteria that appear to lack PTS permeases, suggesting modes of action and substrates that differ from those characterized in Gram-positive bacteria.

Amino Acid Sequence↗

Construction and characterization of a genomic BAC library for the Mus m. musculus mouse subspecies (PWD/Ph inbred strain).

BACKGROUND: The genome of classical laboratory strains of mice is an artificial mosaic of genomes originated from several mouse subspecies with predominant representation (>90%) of the Mus m. domesticus component. Mice of another subspecies, East European/Asian Mus m. musculus, can interbreed with the classical laboratory strains to generate hybrids with unprecedented phenotypic and genotypic variations. To study these variations in depth we prepared the first genomic large insert BAC library from an inbred strain derived purely from the Mus m. musculus-subspecies. The library will be used to seek and characterize genomic sequences controlling specific monogenic and polygenic complex traits, including modifiers of dominant and recessive mutations. RESULTS: A representative mouse genomic BAC library was derived from a female mouse of the PWD/Ph inbred strain of Mus m. musculus subspecies. The library consists of 144,768 primary clones from which 97% contain an insert of 120 kb average size. The library represents an equivalent of 6.7 x mouse haploid genome, as estimated from the total number of clones carrying genomic DNA inserts and from the average insert size. The clones were arrayed in duplicates onto eight high-density membranes that were screened with seven single-copy gene probes. The individual probes identified four to eleven positive clones, corresponding to 6.9-fold coverage of the mouse genome. Eighty-seven BAC-ends of PWD/Ph clones were sequenced, edited, and aligned with mouse C57BL/6J (B6) genome. Seventy-three BAC-ends displayed unique hits on B6 genome and their alignment revealed 0.92 single nucleotide polymorphisms (SNPs) per 100 bp. Insertions and deletions represented 0.3% of the BAC end sequences. CONCLUSION: Analysis of the novel genomic library for the PWD/Ph inbred strain demonstrated coverage of almost seven mouse genome equivalents and a capability to recover clones for specific regions of PWD/Ph genome. The single nucleotide polymorphism between the strains PWD/Ph and C57BL/6J was 0.92/100 bp, a value significantly higher than between classical laboratory strains. The library will serve as a resource for dissecting the phenotypic and genotypic variations between mice of the Mus m. musculus subspecies and classical laboratory mouse strains.

Animals↗

The UCSC genome browser database: update 2007.

The University of California, Santa Cruz Genome Browser Database contains, as of September 2006, sequence and annotation data for the genomes of 13 vertebrate and 19 invertebrate species. The Genome Browser displays a wide variety of annotations at all scales from the single nucleotide level up to a full chromosome and includes assembly data, genes and gene predictions, mRNA and EST alignments, and comparative genomics, regulation, expression and variation data. The database is optimized for fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. In the past year, 22 new assemblies and several new sets of human variation annotation have been released. New features include VisiGene, a fully integrated in situ hybridization image browser; phyloGif, for drawing evolutionary tree diagrams; a redesigned Custom Track feature; an expanded SNP annotation track; and many new display options. The Genome Browser, other tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Animals↗

Noncoding RNA gene detection using comparative sequence analysis.

BACKGROUND: Noncoding RNA genes produce transcripts that exert their function without ever producing proteins. Noncoding RNA gene sequences do not have strong statistical signals, unlike protein coding genes. A reliable general purpose computational genefinder for noncoding RNA genes has been elusive. RESULTS: We describe a comparative sequence analysis algorithm for detecting novel structural RNA genes. The key idea is to test the pattern of substitutions observed in a pairwise alignment of two homologous sequences. A conserved coding region tends to show a pattern of synonymous substitutions, whereas a conserved structural RNA tends to show a pattern of compensatory mutations consistent with some base-paired secondary structure. We formalize this intuition using three probabilistic "pair-grammars": a pair stochastic context free grammar modeling alignments constrained by structural RNA evolution, a pair hidden Markov model modeling alignments constrained by coding sequence evolution, and a pair hidden Markov model modeling a null hypothesis of position-independent evolution. Given an input pairwise sequence alignment (e.g. from a BLASTN comparison of two related genomes) we classify the alignment into the coding, RNA, or null class according to the posterior probability of each class. CONCLUSIONS: We have implemented this approach as a program, QRNA, which we consider to be a prototype structural noncoding RNA genefinder. Tests suggest that this approach detects noncoding RNA genes with a fair degree of reliability.

Algorithms↗

Comparative genomics approaches to study organism similarities and differences.

Comparative genomics is a large-scale, holistic approach that compares two or more genomes to discover the similarities and differences between the genomes and to study the biology of the individual genomes. Comparative studies can be performed at different levels of the genomes to obtain multiple perspectives about the organisms. We discuss in detail the type of analyses that offer significant biological insights in the comparisons of (1) genome structure including overall genome statistics, repeats, genome rearrangement at both DNA and gene level, synteny, and breakpoints; (2) coding regions including gene content, protein content, orthologs, and paralogs; and (3) noncoding regions including the prediction of regulatory elements. We also briefly review the currently available computational tools in comparative genomics such as algorithms for genome-scale sequence alignment, gene identification, and nonhomology-based function prediction.

Animals↗

Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.

With a draft genome-sequence assembly for the chimpanzee available, it is now possible to perform genome-wide analyses to identify, at a submicroscopic level, structural rearrangements that have occurred between chimpanzees and humans. The goal of this study was to investigate chromosomal regions that are inverted between the chimpanzee and human genomes. Using the net alignments for the builds of the human and chimpanzee genome assemblies, we identified a total of 1,576 putative regions of inverted orientation, covering more than 154 mega-bases of DNA. The DNA segments are distributed throughout the genome and range from 23 base pairs to 62 mega-bases in length. For the 66 inversions more than 25 kilobases (kb) in length, 75% were flanked on one or both sides by (often unrelated) segmental duplications. Using PCR and fluorescence in situ hybridization we experimentally validated 23 of 27 (85%) semi-randomly chosen regions; the largest novel inversion confirmed was 4.3 mega-bases at human Chromosome 7p14. Gorilla was used as an out-group to assign ancestral status to the variants. All experimentally validated inversion regions were then assayed against a panel of human samples and three of the 23 (13%) regions were found to be polymorphic in the human genome. These polymorphic inversions include 730 kb (at 7p22), 13 kb (at 7q11), and 1 kb (at 16q24) fragments with a 5%, 30%, and 48% minor allele frequency, respectively. Our results suggest that inversions are an important source of variation in primate genome evolution. The finding of at least three novel inversion polymorphisms in humans indicates this type of structural variation may be a more common feature of our genome than previously realized.

Animals↗

BACFinder: genomic localisation of large insert genomic clones based on restriction fingerprinting.

We have developed software that allows the prediction of the genomic location of a bacterial artificial chromosome (BAC) clone, or other large genomic clone, based on a simple restriction digest of the BAC. The mapping is performed by comparing the experimentally derived restriction digest of the BAC DNA with a virtual restriction digest of the whole genome sequence. Our trials indicate that this program identified the genomic regions represented by BAC clones with a degree of accuracy comparable to that of end-sequencing, but at considerably less cost. Although the program has been developed principally for use with Arabidopsis BACs, it should align large insert genomic clones to any fully sequenced genome.

Arabidopsis↗

Detection of House Dust Mite-derived DNA in Human Lung Tumors by Whole-Genome Sequencing.

Lung cancer in never-smokers (LCINS) accounts for an increasing proportion of lung cancer cases, yet its risk factors remain poorly understood. House dust mites (HDM) are common aeroallergens that induce airway inflammation, but their potential contribution to lung cancer is unknown. We analyzed unmapped whole-genome sequencing reads from 783 lung cancers from the Sherlock-Lung (n = 621 never-smokers) and EAGLE (n = 162 smokers) cohorts, including 328 matched adjacent normal lung tissues. After removal of human sequences, reads were aligned to reference genomes from the two major HDM species and confirmed by BLAST. Samples with top BLAST matches were classified as HDM-detected. Associations between HDM detection and genomic, microbiome, and bulk RNA-seq-derived immune features were evaluated. HDM-derived DNA was detected at low abundance in a subset of tumors and adjacent normal tissues, with higher detection frequencies in tumors than matched normal tissues and in smokers than never-smokers. In LCINS tumors, HDM detection was not associated with tumor mutational burden or recurrent driver alterations but was associated with modest differences in immune cell composition and a limited but reproducible bacterial co-detection pattern. These findings provide a foundation for investigating aeroallergen-derived DNA signatures and their potential relationship to the lung tumor microenvironment.

Environmental exposure↗

Molecular cloning and partial characterization of a parrot papillomavirus.

The genome of a papillomavirus isolated from a cutaneous lesion on the head of an African grey parrot (Psittacus erithacus timneh) was cloned into pBR322, and a restriction map was prepared. Several short portions of the DNA were sequenced allowing the genome to be aligned with HPV-la. In Southern blot hybridizations, under conditions of low and medium stringency (Tm-40 and -33 degrees), but not at a higher stringency, this viral DNA annealed weakly with only 1 of 17 mammalian papillomavirus genomes tested. Furthermore, PePV DNA hybridized with the DNA from the European chaffinch only at low stringency, indicating that it represents a unique avian papillomavirus.

Animals↗

Comparative genomics reveals unusually long motifs in mammalian genomes.

MOTIVATION: The recent discovery of the first small modulatory RNA (smRNA) presents the challenge of finding other molecules of similar length and conservation level. Unlike short interfering RNA (siRNA) and micro-RNA (miRNA), effective computational and experimental screening methods are not currently known for this species of RNA molecule, and the discovery of the one known example was partly fortuitous because it happened to be complementary to a well-studied DNA binding motif (the Neuron Restrictive Silencer Element). RESULTS: The existing comparative genomics approaches (e.g., phylogenetic footprinting) rely on alignments of orthologous regions across multiple genomes. This approach, while extremely valuable, is not suitable for finding motifs with highly diverged "non-alignable" flanking regions. Here we show that several unusually long and well conserved motifs can be discovered de novo through a comparative genomics approach that does not require an alignment of orthologous upstream regions. These motifs, including Neuron Restrictive Silencer Element, were missed in recent comparative genomics studies that rely on phylogenetic footprinting. While the functions of these motifs remain unknown, we argue that some may represent biologically important sites. AVAILABILITY: Our comparative genomics software, a web-accessible database of our results and a compilation of experimentally validated binding sites for NRSE can be found at http://www.cse.ucsd.edu/groups/bioinformatics.

Algorithms↗

Genome-wide detection of alternative splicing in expressed sequences using partial order multiple sequence alignment graphs.

We present a method for high-throughput alternative splicing detection in expressed sequence data. This method effectively copes with many of the problems inherent in making inferences about splicing and alternative splicing on the basis of EST sequences, which in addition to being fragmentary and full of sequencing errors, may also be chimeric, misoriented, or contaminated with genomic sequence. Our method, which relies both on the Partial Order Alignment (POA) program for constructing multiple sequence alignments, and its Heaviest Bundling function for generating consensus sequences, accounts for the real complexity of expressed sequence data by building and analyzing a single multiple sequence alignment containing all of the expressed sequences in a particular cluster aligned to genomic sequence. We illustrate application of this method to human UniGene Cluster Hs.1162, which contains expressed sequences from the human HLA-DMB gene. We have used this method to generate databases, published elsewhere, of splices and alternative splicing relationships for the human, mouse and rat genomes. We present statistics from these calculations, as well as the CPU time for running our method on expressed sequence clusters of varying size, to verify that it truly scales to complete genomes.

Alternative Splicing↗

Genomic organization and DNA sequences of two human phenol sulfotransferase genes (STP1 and STP2) on the short arm of chromosome 16.

A family of human phenol sulfotransferase genes has been suggested by the cloning of numerous cDNA isolates from different tissues. We have previously cloned and sequenced the STM gene encoding the monoamine neurotransmitter-preferring sulfotransferase, M-PST, and a portion of the STP1 gene encoding the phenol-preferring isozyme, P-PST1 (BBRC 205, 1325-1332; Genomics 18, 440-443). Both genes were mapped to a small region on the short arm of chromosome 16 (BBRC 205, 482-489). Here we report on the sequencing and genomic organization of the STP1 and STP2 genes from a single cosmid clone obtained from chromosome 16p12.1-p11.2. STP1 and STP2 are 95.9% identical at the amino acid sequence level, whereas the STM gene is only 92.9% and 90.5% identical to STP1 and STP2, respectively. Alignment of the genomic sequences indicated that all three genes have 7 coding exons and conserved intron-exon boundaries. These results facilitated the assignment of previously published cDNA isolates as "alleles" of the individual STM, STP1, and STP2 loci on 16p, and provide to us a greater understanding of the complexity and roles of the phenol sulfotransferase gene family in the metabolism of endogenous and xenobiotic agents.

Alleles↗

The institute for genomic research Osa1 rice genome annotation database.

We have developed a rice (Oryza sativa) genome annotation database (Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis (Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 non-transposable element-related genes, 18,545 (42.4%) were annotated with a putative function, 5,777 (13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 (44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms (5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser. All data can be obtained through the project Web pages at http://rice.tigr.org.

Computational Biology↗

CoverM: read alignment statistics for metagenomics.

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Metabolomics↗

Alignment of receptor nomenclature with the human genome: classification of 5-HT1B and 5-HT1D receptor subtypes.

The continuing rapid progress towards a complete database of structural information on the human genome creates a challenge of ensuring that current schemes for classifying and naming receptors and ion channels effectively integrate this information with functional data to provide unambiguous principles for classification. In this article, Paul Hartig and colleagues review the recent deliberations of the Serotonin Club Nomenclature Committee and outline a number of its recommendations aimed at encouraging consistency in current and future receptor nomenclature. Based on these principles, the present classification of 5-HT1B and 5-HT1D receptors is reconsidered, and a revised nomenclature for 5-HT1B, 5-HT1D alpha and 5-HT1D beta receptor subtypes is suggested.

Animals↗

Divergent structures of Caenorhabditis elegans cytochrome P450 genes suggest the frequent loss and gain of introns during the evolution of nematodes.

The Caenorhabditis elegans genome contains more than 60 cytochrome P450 (CYP) genes. The exon-intron organizations of all of the available and potentially active C. elegans CYP genes were inferred by a newly developed program for predicting protein-coding exons based on the alignment of a genomic DNA sequence and a protein profile. From the predicted amino acid sequences, all of the C. elegans CYP genes except one were classified into three groups, which were closely related to the mammalian drug-metabolizing P450 gene families CYP2, CYP3, and CYP4. The gene structures were strikingly divergent within each group; 20, 10, and 5 unique gene organizations were identified among 40, 18, and 5 genes in the CYP2-, CYP3-, and CYP4-related groups, respectively. The degrees of divergence in gene organization were strongly correlated with those in the amino acid sequences of encoding proteins, and the minimum rate of change in an intron insertion site was estimated to be about 90 times less frequent than amino acid substitutions. Parsimonious analyses suggested that frequent loss and gain of introns has occurred during the evolution of CYP genes in each group after the divergence of nematodes, arthropods, and deuterostomia. Few, if any, incidents of intron sliding were evident, and a model that did not allow intron insertions was highly inconsistent with the observations. All of these findings are explained better by the intron-late view than by the intron-early view.

Amino Acid Sequence↗

Whitefly (Bemisia tabaci) genome project: analysis of sequenced clones from egg, instar, and adult (viruliferous and non-viruliferous) cDNA libraries.

BACKGROUND: The past three decades have witnessed a dramatic increase in interest in the whitefly Bemisia tabaci, owing to its nature as a taxonomically cryptic species, the damage it causes to a large number of herbaceous plants because of its specialized feeding in the phloem, and to its ability to serve as a vector of plant viruses. Among the most important plant viruses to be transmitted by B. tabaci are those in the genus Begomovirus (family, Geminiviridae). Surprisingly, little is known about the genome of this whitefly. The haploid genome size for male B. tabaci has been estimated to be approximately one billion bp by flow cytometry analysis, about five times the size of the fruitfly Drosophila melanogaster. The genes involved in whitefly development, in host range plasticity, and in begomovirus vector specificity and competency, are unknown. RESULTS: To address this general shortage of genomic sequence information, we have constructed three cDNA libraries from non-viruliferous whiteflies (eggs, immature instars, and adults) and two from adult insects that fed on tomato plants infected by two geminiviruses: Tomato yellow leaf curl virus (TYLCV) and Tomato mottle virus (ToMoV). In total, the sequence of 18,976 clones was determined. After quality control, and removal of 5,542 clones of mitochondrial origin 9,110 sequences remained which included 3,843 singletons and 1,017 contigs. Comparisons with public databases indicated that the libraries contained genes involved in cellular and developmental processes. In addition, approximately 1,000 bases aligned with the genome of the B. tabaci endosymbiotic bacterium Candidatus Portiera aleyrodidarum, originating primarily from the egg and instar libraries. Apart from the mitochondrial sequences, the longest and most abundant sequence encodes vitellogenin, which originated from whitefly adult libraries, indicating that much of the gene expression in this insect is directed toward the production of eggs. CONCLUSION: This is the first functional genomics project involving a hemipteran (Homopteran) insect from the subtropics/tropics. The B. tabaci sequence database now provides an important tool to initiate identification of whitefly genes involved in development, behaviour, and B. tabaci-mediated begomovirus transmission.

Amino Acid Sequence↗