Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

GeneSeqer@PlantGDB: Gene structure prediction in plant genomes.

The GeneSeqer@PlantGDB Web server (http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi) provides a gene structure prediction tool tailored for applications to plant genomic sequences. Predictions are based on spliced alignment with source-native ESTs and full-length cDNAs or non-native probes derived from putative homologous genes. The tool is illustrated with applications to refinement of current gene structure annotation and de novo annotation of draft genomic sequences. The service should facilitate expert annotation as a community effort by providing convenient access to all public plant sequences via the PlantGDB database, a simple four-step protocol for spliced alignment and visually appealing displays of the predicted gene structures in addition to detailed sequence alignments.

Arabidopsis↗

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

SUMMARY: Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors." These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac. We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC. Alevin-fry-atac is highly scalable and, when using 32 threads, is 2.8 times faster than Chromap (the second fastest approach) while using only 33% of the memory required by Chromap. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual color-enhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger. AVAILABILITY AND IMPLEMENTATION: Alevin-fry-atac is written in Rust and C++17, and is freely-available under a BSD 3-clause license. It is integrated into piscem (https://github.com/COMBINE-lab/piscem) and alevin-fry (https://github.com/COMBINE-lab/alevin-fry), and is also supported directly as part of simpleaf (https://github.com/COMBINE-lab/simpleaf).

Single-Cell Analysis↗

Software tools for analyzing pairwise alignments of long sequences.

Pairwise comparison of long stretches of genomic DNA sequence can identify regions conserved across species, which often indicate functional significance. However, the novel insights frequently must be windowed from a flood of information; for instance, running an alignment program on two 50-kilobase sequences might yield over a hundred pages of alignments. Direct inspection of such a volume of printed output is infeasible, or at best highly undesirable, and computer tools are needed to summarize the information, to assist in its analysis, and to report the findings. This paper describes two such software tools. One tool prepares publication-quality pictorial representations of alignments, while another facilitates interactive browsing of pairwise alignment data. Their effectiveness is illustrated by comparing the beta-like globin gene clusters between humans and rabbits. A second example compares the chloroplast genomes of tobacco and liverwort.

Animals↗

Computational approaches for the analysis of gene neighbourhoods in prokaryotic genomes.

Gene order in prokaryotes is conserved to a much lesser extent than protein sequences. Only some operons, primarily those that encode physically interacting proteins, are conserved in all or most of the bacterial and archaeal genomes. Nevertheless, even the limited conservation of operon organisation that is observed provides valuable evolutionary and functional clues through multiple genome comparisons. With the rapid growth in the number and diversity of sequenced prokaryotic genomes, functional inferences for uncharacterized genes located in the same conserved gene neighborhood with well-studied genes are becoming increasingly important. In this review, we discuss various computational approaches for identification of conserved gene strings and construction of local alignments of gene orders in prokaryotic genomes.

Algorithms↗

The genomes of bovine papillomaviruses types 3 and 4 are colinear.

The 7.2 kb genomic DNA of bovine papillomavirus type 3 (BPV-3) was molecularly cloned using its unique EcoRI site, and the 7.3 kb genome of BPV-4 was cloned using its single BamHI site. The viral genomes were compared by liquid hybridization, Southern blot hybridization and heteroduplex mapping. Low stringency hybridization conditions revealed that the genomes are colinear but the sequences are extensively mismatched. The relative alignment of the restriction endonuclease maps of the two viral genomes has been determined. It was found that the genomes as linearized for cloning are out of phase by 1.7 kb, so that the single EcoRI site of BPV-3 appears to coincide with the BPV-4 EcoRI site at 0.22 map units. It is concluded that the genomes of BPV-3 and BPV-4, both of which cause true epithelial warts, share the same physical organization but exhibit sequence divergence.

Base Sequence↗

The Helicobacter pylori genome: from sequence analysis to structural and functional predictions.

Fold assignments for proteins from the Helicobacter pylori genome are carried out using BASIC, a profile-profile alignment algorithm recently tested on the Mycoplasma genitalium and Escherichia coli genomes. The fold assignments are followed by automated function evaluation, based on the multilevel description of functional sites in proteins. Over 40% of the proteins encoded in the H. pylori genome can be recognized as belonging to a protein family with known structure. Previous estimates suggested that only 10-15% of genome proteins could be characterized this way. This dramatic increase in the number of recognized homologies between H. pylori proteins and structurally characterized protein families is partly due to the rapid increase of the database of known protein structures, but mostly it is due to the significant improvement in prediction algorithms. Knowledge of a protein fold adds a new dimension to our understanding of its function and, similarly, structure prediction can also add to understanding, verification, and/or prediction of function for uncharacterized proteins. Several examples analyzed in more detail in this article illustrate insights that can be achieved from structure and detailed function prediction.

Algorithms↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

Molecular Identification and Genotyping of Blastocystis Spp. In Children with Clinical Symptoms in Southeast Iran Using PCR-Sequencing Method.

Blastocystis spp. is a zoonotic anaerobic parasite that has been identified in the large intestine of humans and many vertebrates. It is predominantly encountered in individuals with frequent contact with animals. The present study aims to identify the prevalence of Blastocystis spp. and its common genotypes in children with clinical symptoms of diarrhea in the city of Zahedan, located in the southeast of Iran. A cross-sectional descriptive study was conducted on 60 children under ten years of age with gastrointestinal symptoms, especially diarrhea. Following the collection of samples, stool samples were subjected to direct stool testing for the initial diagnosis. Following this, a microscopic diagnosis was made, after which DNA was extracted and a Polymerase Chain Reaction (PCR) test with a small subunit ribosomal RNA (SSU rRNA) gene target was performed. The PCR products were then purified and sequenced. The resulting nucleotide sequences were then subjected to a thorough review using Chromas biotechnology software version 2.4 and CLC genomic work bench software 11. The alignment of the nucleotide sequences was subsequently facilitated by utilizing the BLAST database, and these sequences were then compared with the reference genotypes of Blastocystis spp. that are stored within the gene bank. The genotyping of the sequences was conducted using CLC genomic work bench software 11, and a phylogenetic tree was constructed using MEGA7 software with the Neighbor-Joining statistical method, which applied the Kimura 2-parameter method. Out of the 60 cases that were examined, 5 children (8.33%) were found to be positive by direct microscopic and PCR tests, where a 500 (479) bp fragment in the SSU-rRNA target was detected. Subsequent genetic analysis identified four distinct subtypes, including subtypes 1, 2, 3, and 5. The percentage of nucleotide identity with the sequences in the gene bank was found to be between 93 and 100%. Given the presence of subtypes 3 and 5 in the study and the evidence of their zoonotic nature, it can be concluded that examining parasite dynamics and epidemiological principles can be effective in the control strategy.

Blastocystis↗

A whole genome shotgun gene fusion method for isolation of translation initiation sites in Escherichia coli: identification of Haemophilus influenzae translation initiation sites in E. coli.

We have developed a new method for isolating translation initiation sites based on the expression of Haemophilus influenzae Rd gene fusions with the Escherichia coli galactokinase (galK) gene. We cloned random DNA fragments of H. influenzae Rd DNA into a plasmid vector containing the galK coding sequence from which the translation initiation site (the ribosome binding site and translation initiation codon) had been removed. A subset of the cloned DNA fragments contained translation initiation sites that, when fused to the galK gene, produced active galactokinase and complemented the host galK mutation. Molecules expressing galactokinase activity were isolated and characterized by DNA sequence analysis, and the sequences were aligned with the recently completed whole genomic sequence of H. influenzae Rd. Translation initiation sites for known, hypothetical, and new genes were identified. Translation initiation sites internal to the coding sequences of a number of genes were identified, suggesting that internal translation initiation sites are common, especially in large genes. This shotgun method provides functional information on translation initiation sites and helps to define gene coding sequences.

Amino Acid Sequence↗

Construction of a radiation hybrid map of chicken chromosome 2 and alignment to the chicken draft sequence.

BACKGROUND: The ChickRH6 whole chicken genome radiation hybrid (RH) panel recently produced has already been used to build radiation hybrid maps for several chromosomes, generating comparative maps with the human and mouse genomes and suggesting improvements to the chicken draft sequence assembly. Here we present the construction of a RH map of chicken chromosome 2. Markers from the genetic map were used for alignment to the existing GGA2 (Gallus gallus chromosome 2) linkage group and EST were used to provide valuable comparative mapping information. Finally, all markers from the RH map were localised on the chicken draft sequence assembly to check for eventual discordances. RESULTS: Eighty eight microsatellite markers, 10 genes and 219 EST were selected from the genetic map or on the basis of available comparative mapping information. Out of these 317 markers, 270 gave reliable amplifications on the radiation hybrid panel and 198 were effectively assigned to GGA2. The final RH map is 2794 cR6000 long and is composed of 86 framework markers distributed in 5 groups. Conservation of synteny was found between GGA2 and eight human chromosomes, with segments of conserved gene order of varying lengths. CONCLUSION: We obtained a radiation hybrid map of chicken chromosome 2. Comparison to the human genome indicated that most of the 8 groups of conserved synteny studied underwent internal rearrangements. The alignment of our RH map to the first draft of the chicken genome sequence assembly revealed a good agreement between both sets of data, indicative of a low error rate.

Animals↗

cDNA sequencing and analysis of POV1 (PB39): a novel gene up-regulated in prostate cancer.

We recently identified a novel gene (PB39) (HGMW-approved symbol POV1) whose expression is up-regulated in human prostate cancer using tissue microdissection-based differential display analysis. In the present study we report the full-length sequencing of PB39 cDNA, genomic localization of the PB39 gene, and genomic sequence of the mouse homologue. The full-length human cDNA is 2317 nucleotides in length and contains an open reading frame of 559 amino acids which does not show homology with any reported human genes. The N-terminus contains charged amino acids and a helical loop pattern suggestive of an srp leader sequence for a secreted protein. Fluorescence in situ hybridization using PB39 cDNA as probe mapped the gene to chromosome 11p11.1-p11.2. Comparison of PB39 cDNA sequence with murine sequence available in the public database identified a region of previously sequenced mouse genomic DNA showing 67% amino acid sequence homology with human PB39. Based on alignment and comparison to the human cDNA the mouse genomic sequence suggests there are at least 14 exons in the mouse gene spread over approximately 100 kb of genomic sequence. Further analysis of PB39 expression in human tissues shows the presence of a unique splice variant mRNA that appears to be primarily associated with fetal tissues and tumors. Interestingly, the unique splice variant appears in prostatic intraepithelial neoplasia, a microscopic precursor lesion of prostate cancer. The current data support the hypothesis that PB39 plays a role in the development of human prostate cancer and will be useful in the analysis of the gene product in further human and murine studies.

Amino Acid Sequence↗

Structure alignment via Delaunay tetrahedralization.

A novel protein structure alignment technique has been developed reducing much of the secondary and tertiary structure to a sequential representation greatly accelerating many structural computations, including alignment. Constructed from incidence relations in the Delaunay tetrahedralization, alignments of the sequential representation describe structural similarities that cannot be expressed with rigid-body superposition and complement existing techniques minimizing root-mean-squared distance through superposition. Restricting to the largest substructure superimposable by a single rigid-body transformation determines an alignment suitable for root-mean-squared distance comparisons and visualization. Restricted alignments of a test set of histones and histone-like proteins determined superpositions nearly identical to those produced by the established structure alignment routines of DaliLite and ProSup. Alignment of three, increasingly complex proteins: ferredoxin, cytidine deaminase, and carbamoyl phosphate synthetase, to themselves, demonstrated previously identified regions of self-similarity. All-against-all similarity index comparisons performed on a test set of 45 class I and class II aminoacyl-tRNA synthetases closely reproduced the results of established distance matrix methods while requiring 1/16 the time. Principal component analysis of pairwise tetrahedral decomposition similarity of 2300 molecular dynamics snapshots of tryptophanyl-tRNA synthetase revealed discrete microstates within the trajectory consistent with experimental results. The method produces results with sufficient efficiency for large-scale multiple structure alignment and is well suited to genomic and evolutionary investigations where no geometric model of similarity is known a priori.

Algorithms↗

The Genomic Threading Database: a comprehensive resource for structural annotations of the genomes from key organisms.

Currently, the Genomic Threading Database (GTD) contains structural assignments for the proteins encoded within the genomes of nine eukaryotes and 101 prokaryotes. Structural annotations are carried out using a modified version of GenTHREADER, a reliable fold recognition method. The Gen THREADER annotation jobs are distributed across multiple clusters of processors using grid technology and the predictions are deposited in a relational database accessible via a web interface at http://bioinf.cs.ucl.ac.uk/GTD. Using this system, up to 84% of proteins encoded within a genome can be confidently assigned to known folds with 72% of the residues aligned. On average in the GTD, 64% of proteins encoded within a genome are confidently assigned to known folds and 58% of the residues are aligned to structures.

Animals↗

Computational identification of noncoding RNAs in E. coli by comparative genomics.

Some genes produce noncoding transcripts that function directly as structural, regulatory, or even catalytic RNAs [1, 2]. Unlike protein-coding genes, which can be detected as open reading frames with distinctive statistical biases, noncoding RNA (ncRNA) gene sequences have no obvious inherent statistical biases [3]. Thus, genome sequence analyses reveal novel protein-coding genes, but any novel ncRNA genes remain invisible. Here, we describe a computational comparative genomic screen for ncRNA genes. The key idea is to distinguish conserved RNA secondary structures from a background of other conserved sequences using probabilistic models of expected mutational patterns in pairwise sequence alignments. We report the first whole-genome screen for ncRNA genes done with this method, in which we applied it to the "intergenic" spacers of Escherichia coli using comparative sequence data from four related bacteria. Starting from >23,000 conserved interspecies pairwise alignments, the screen predicted 275 candidate structural RNA loci. A sample of 49 candidate loci was assayed experimentally. At least 11 loci expressed small, apparently noncoding RNA transcripts of unknown function. Our computational approach may be used to discover structural ncRNA genes in any genome for which appropriate comparative genome sequence data are available.

Animals↗

Deriving ribosomal binding site (RBS) statistical models from unannotated DNA sequences and the use of the RBS model for N-terminal prediction.

Accurate prediction of the position of translation initiation (N-terminal prediction) is a difficult problem. N-terminal prediction from DNA sequence alone is ambiguous is several candidate start sites are close to each other. Protein similarity search is usually unable to indicate the true start of a gene as it would require a strong protein sequence similarity at the N-terminal portion of a protein where conservative regions are rarely situated. With the aid of the GeneMark program for gene identification, we extract DNA sequence fragments presumably containing ribosome binding sites (RBS) from unannotated complete genomic sequences. These DNA segments are aligned to generate the RBS model using the Gibbs' sampling method. N-terminal prediction is then performed by using the RBS model in conjunction with the GeneMark start codon prediction to aid in determining the true N-terminal site.

Base Sequence↗

BLAST2GENE: a comprehensive conversion of BLAST output into independent genes and gene fragments.

SUMMARY: BLAST2GENE is a program that allows a detailed analysis of genomic regions containing completely or partially duplicated genes. From a BLAST (or BL2SEQ) comparison of a protein or nucleotide query sequence with any genomic region of interest, BLAST2GENE processes all high scoring pairwise alignments (HSPs) and provides the disposition of all independent copies along the genomic fragment. The results are provided in text and PostScript formats to allow an automatic and visual evaluation of the respective region. AVAILABILITY: The program is available upon request from the authors. A web server of BLAST2GENE is maintained at http://www.bork.embl.de/blast2gene

Algorithms↗

A whole-genome shotgun optical map of Yersinia pestis strain KIM.

Yersinia pestis is the causative agent of the bubonic, septicemic, and pneumonic plagues (also known as black death) and has been responsible for recurrent devastating pandemics throughout history. To further understand this virulent bacterium and to accelerate an ongoing sequencing project, two whole-genome restriction maps (XhoI and PvuII) of Y. pestis strain KIM were constructed using shotgun optical mapping. This approach constructs ordered restriction maps from randomly sheared individual DNA molecules directly extracted from cells. The two maps served different purposes; the XhoI map facilitated sequence assembly by providing a scaffold for high-resolution alignment, while the PvuII map verified genome sequence assembly. Our results show that such maps facilitated the closure of sequence gaps and, most importantly, provided a purely independent means for sequence validation. Given the recent advancements to the optical mapping system, increased resolution and throughput are enabling such maps to guide sequence assembly at a very early stage of a microbial sequencing project.

Genome, Bacterial↗

The structure and gene repertoire of an ancient red algal plastid genome.

Photosynthetic eukaryotes can, according to features of their chloroplasts, be divided into two major groups: the red and the green lineage of plastid evolution. To extend the knowledge about the evolution of the red lineage we have sequenced and analyzed the chloroplast genome (cp-genome) of Cyanidium caldarium RK1, a unicellular red alga (AF022186). The analysis revealed that this genome shows several unusual structural features, such as a hypothetical hairpin structure in a gene-free region and absence of large repeat units. We provide evidence that this structural organization of the cp-genome of C. caldarium may be that of the most ancient cp-genome so far described. We also compared the cp-genome of C. caldarium to the other known cp-genomes of the red lineage. The cp-genome of C. caldarium cannot be readily aligned with that of Porphyra purpurea, a multicellular red alga, or Guillardia theta due to a displacement of a region of the cp-genome. The phylogenetic tree reveals that the secondary endosymbiosis, through which G. theta evolved, took place after the separation of the ancestors of C. caldarium and P. purpurea. We found several genes unique to the cp-genome of C. caldarium. Five of them seem to be involved in the building of bacterial cell envelopes and may be responsible for the thermotolerance of the chloroplast of this alga. Two additional genes may play a role in stabilizing the photosynthetic machinery against salt stress and detoxification of the chloroplast. Thus, these genes may be unique to the cp-genome of C. caldarium and may be required for the endurance of the extreme living conditions of this alga.

Adaptation, Physiological↗