Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Use of massively parallel signature sequencing to study genes expressed during the plant defense response.

Massively parallel signature sequencing is a sequencing-based method that provides quantitative gene expression data for nearly all transcripts in a particular ribonucleic acid sample. Although the sequencing technology is practiced as a service by a California-based company, we have developed methods for the handling and analysis of these data. This chapter describes the steps involved in obtaining data from massively parallel signature sequencing, aligning the signatures to genomic sequence, identifying novel transcripts, and performing quantitative analyses of genes expressed under conditions such as disease treatments.

Arabidopsis↗

Fluorescent detection of differentially expressed cDNA using SYBR gold nucleic acid gel stain.

We describe herein a modified differential gene display (DGD) technique that can be rapidly and simply performed and that eliminates the need for radioactivity by fluorescent visualization of complementary deoxyribonucleic acid (cDNA) bands with SYBR gold nucleic acid gel stain. To streamline the DGD procedure, a number of modifications were employed. Ribonucleic acid isolated from differentially treated populations of human trabecular bone-derived mesenchymal progenitor cells was reverse-transcribed into cDNA using oligo-dT primer, and subsequent amplification of differentially expressed cDNAs was done using arbitrary 25-mer primers and oligo-dT9 30-mer primers. Moderate-sized nondenaturing 6% polyacrylamide gels (30 x 20 cm) of 1.5-mm thickness were used for easier handling and increased sample loading capacity. Gels were subjected to electrophoresis overnight, stained with SYBR gold, and visualized and photographed using a commercially available gel imager. DNA bands ranging in size from 100 to 400 bp were visualized directly on an ultraviolet transilluminator, excised from the gel, and reamplified. The cDNA amplicons were subcloned, sequenced, and gene sequences were identified by a Basic Local Alignment Search Tool of genomic databases. Overall, this rapid and functional method proved quite effective for identification of novel genes that may be of interest in studies of cartilage and bone differentiation.

DNA, Complementary↗

Gene HI1472 of Haemophilus influenzae Rd is a novel gene involved in DNA repair.

A chimeric plasmid, pJPuvr4, consists of a 16.7 kbp Haemophilus influenzae Rd chromosomal DNA insert at the EcoRI site of vector pJ1-8. This plasmid complements the UV and gamma ray sensitivity of the mutant strain MBH4. This plasmid carries the wild type allele of gene uvr4 which was localised to a 3.8 kbp DraI fragment, with an internal EcoRI site. Partial sequencing of the gene and its alignment with the published genome sequence of H. influenzae Rd revealed uvr4 to be HI1472. HI1472 is a putatively identified open reading frame (ORF), which has been assigned no function so far. The partial sequence did show nt database match with 3D exon of N cadherin gene of homosepians and moaA gene of H. influenzae. Cadherins are involved in cell adhesion, cell to cell contact and morphogenesis in homosepians and moaA gene codes for molybdenum biosynthesis subunitA. This report implicates HI1472 of Haemophilus influenzae Rd in DNA repair. Nucleotide sequence obtained for the gene uvr4 was compared with the published sequence of gene HI1472. A wild type strain variation was observed at the 592nd nucleotide position corresponding to a change from aspartic acid to threonine.

Base Sequence↗

An infrastructure for comparative genomics to functionally characterize genes and proteins.

Current genome projects are resulting in a flood of sequence data. The interpretation of these sequences is lagging, and optimized data analysis strategies need to be developed. Much can be learned from comparing different genomes, as genomes of distant organisms may still encode proteins with high sequence similarity. The order of genes (co linearity) in genomes may also be conserved to some extend. We have employed both these observations to create a multi-functional, computational analysis system (genomeSCOUT) which allows for rapid identification and functional characterization of genes and proteins through genome comparison. With a number of independent algorithms, information about different levels of protein homology (concerning e.g. paralogs, orthologs and clusters of orthologous groups, COGs) and gene order is collected and stored in several value added databases. These databases are then used for interactive comparison of genomes and subsequent analysis. The application is based on the well established data integration system SRS. This ensures (1) fast handling of large genomic data sets, (2) straightforward access to a multitude of biological databases, (3) unique linking functions between these databases, (4) highly efficient collection of information on genes and proteins, and 5. fully integrated and user friendly graphical representations of search results. This application can be used for projects as diverse as the correct annotation of genomes, the optimization of (micro) organisms for industrial production, or the identification of drug targets.

Computational Biology↗

[An introduction of several programs used in genomic analysis].

Genomics is a novel subject that has been developed accompanying with the progress of human genome project. Genomics deals with the chemistry component, structure organization and evolution of genome at global level. As genomics associated with huge data, bioinformatics plays an important role in these processes of data production, data management and data mining. At present, many reliable programs have been used in genomic research successfully, which are usually accessible and downloaded freely. We address here the principles of some programs used wildly in genomics such as sequence alignment, sequence assembly, repeat identification and gene prediction, which are exemplified with typical programs respectively.

English Abstract↗

Haplotype assembly from aligned weighted SNP fragments.

Given an assembled genome of a diploid organism the haplotype assembly problem can be formulated as retrieval of a pair of haplotypes from a set of aligned weighted SNP fragments. Known computational formulations (models) of this problem are minimum letter flips (MLF) and the weighted minimum letter flips (WMLF; Greenberg et al. (INFORMS J. Comput. 2004, 14, 211-213)). In this paper we show that the general WMLF model is NP-hard even for the gapless case. However the algorithmic solutions for selected variants of WMFL can exist and we propose a heuristic algorithm based on a dynamic clustering technique. We also introduce a new formulation of the haplotype assembly problem that we call COMPLETE WMLF (CWMLF). This model and algorithms for its implementation take into account a simultaneous presence of multiple kinds of data errors. Extensive computational experiments indicate that the algorithmic implementations of the CWMLF model achieve higher accuracy of haplotype reconstruction than the WMLF-based algorithms, which in turn appear to be more accurate than those based on MLF.

Algorithms↗

Conserved fragments of transposable elements in intergenic regions: evidence for widespread recruitment of MIR- and L2-derived sequences within the mouse and human genomes.

We analysed the distribution of transposable elements (TEs) in 100 aligned pairs of orthologous intergenic regions from the mouse and human genomes. Within these regions, conserved segments of high similarity between the two species alternate with segments of low similarity. Identifiable TEs comprise 40-60% of segments of low similarity. Within such segments, a particular copy of a TE found in one species has no orthologue in the other. Overall, TEs comprise only approximately 20 % of conserved segments. However, TEs from two families, MIR and L2, are rather common within conserved segments. Statistical analysis of the distributions of TEs suggests that a majority of the MIR and L2 elements present in murine intergenic regions have human orthologues. These elements must have been present in the common ancestor of human and mouse and have remained under substantial negative selection that prevented their divergence beyond recognition. If so, recruitment of MIR- and L2-derived sequences to perform a function that increases host fitness is rather common, with at least two such events per host gene. The central part of the MIR consensus sequence is over-represented in conserved segments given its background frequency in the genome, suggesting that it is under the strongest selective constraint.

Animals↗

Consensus folding of aligned sequences as a new measure for the detection of functional RNAs by comparative genomics.

Facing the ever-growing list of newly discovered classes of functional RNAs, it can be expected that further types of functional RNAs are still hidden in recently completed genomes. The computational identification of such RNA genes is, therefore, of major importance. While most known functional RNAs have characteristic secondary structures, their free energies are generally not statistically significant enough to distinguish RNA genes from the genomic background. Additional information is required. Considering the wide availability of new genomic data of closely related species, comparative studies seem to be the most promising approach. Here, we show that prediction of consensus structures of aligned sequences can be a significant measure to detect functional RNAs. We report a new method to test multiple sequence alignments for the existence of an unusually structured and conserved fold. We show for alignments of six types of well-known functional RNA that an energy score consisting of free energy and a covariation term significantly improves sensitivity compared to single sequence predictions. We further test our method on a number of non-coding RNAs from Caenorhabditis elegans/Caenorhabditis briggsae and seven Saccharomyces species. Most RNAs can be detected with high significance. We provide a Perl implementation that can be used readily to score single alignments and discuss how the methods described here can be extended to allow for efficient genome-wide screens.

Algorithms↗

Sequence alignment by cross-correlation.

Many recent advances in biology and medicine have resulted from DNA sequence alignment algorithms and technology. Traditional approaches for the matching of DNA sequences are based either on global alignment schemes or heuristic schemes that seek to approximate global alignment algorithms while providing higher computational efficiency. This report describes an approach using the mathematical operation of cross-correlation to compare sequences. It can be implemented using the fast fourier transform for computational efficiency. The algorithm is summarized and sample applications are given. These include gene sequence alignment in long stretches of genomic DNA, finding sequence similarity in distantly related organisms, demonstrating sequence similarity in the presence of massive (approximately 90%) random point mutations, comparing sequences related by internal rearrangements (tandem repeats) within a gene, and investigating fusion proteins. Application to RNA and protein sequence alignment is also discussed. The method is efficient, sensitive, and robust, being able to find sequence similarities where other alignment algorithms may perform poorly.

Algorithms↗

A novel approach to structural alignment using realistic structural and environmental information.

In the era of structural genomics, it is necessary to generate accurate structural alignments in order to build good templates for homology modeling. Although a great number of structural alignment algorithms have been developed, most of them ignore intermolecular interactions during the alignment procedure. Therefore, structures in different oligomeric states are barely distinguishable, and it is very challenging to find correct alignment in coil regions. Here we present a novel approach to structural alignment using a clique finding algorithm and environmental information (SAUCE). In this approach, we build the alignment based on not only structural coordinate information but also realistic environmental information extracted from biological unit files provided by the Protein Data Bank (PDB). At first, we eliminate all environmentally unfavorable pairings of residues. Then we identify alignments in core regions via a maximal clique finding algorithm. Two extreme value distribution (EVD) form statistics have been developed to evaluate core region alignments. With an optional extension step, global alignment can be derived based on environment-based dynamic programming linking. We show that our method is able to differentiate three-dimensional structures in different oligomeric states, and is able to find flexible alignments between multidomain structures without predetermined hinge regions. The overall performance is also evaluated on a large scale by comparisons to current structural classification databases as well as to other alignment methods.

Algorithms↗

Enhanced genome annotation using structural profiles in the program 3D-PSSM.

A method (three-dimensional position-specific scoring matrix, 3D-PSSM) to recognise remote protein sequence homologues is described. The method combines the power of multiple sequence profiles with knowledge of protein structure to provide enhanced recognition and thus functional assignment of newly sequenced genomes. The method uses structural alignments of homologous proteins of similar three-dimensional structure in the structural classification of proteins (SCOP) database to obtain a structural equivalence of residues. These equivalences are used to extend multiply aligned sequences obtained by standard sequence searches. The resulting large superfamily-based multiple alignment is converted into a PSSM. Combined with secondary structure matching and solvation potentials, 3D-PSSM can recognise structural and functional relationships beyond state-of-the-art sequence methods. In a cross-validated benchmark on 136 homologous relationships unambiguously undetectable by position-specific iterated basic local alignment search tool (PSI-Blast), 3D-PSSM can confidently assign 18 %. The method was applied to the remaining unassigned regions of the Mycoplasma genitalium genome and an additional 13 regions were assigned with 95 % confidence. 3D-PSSM is available to the community as a web server: http://www.bmm.icnet.uk/servers/3dpssm

Algorithms↗

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide↗

Identification of the alpha-aminoadipic semialdehyde dehydrogenase-phosphopantetheinyl transferase gene, the human ortholog of the yeast LYS5 gene.

In mammals, L-lysine is first catabolized to alpha-aminoadipate semialdehyde by the bifunctional enzyme alpha-aminoadipate semialdehyde synthase (AASS), followed by a conversion to alpha-aminoadipate by alpha-aminoadipate semialdehyde dehydrogenase. In Saccharomyces cerevisiae, which synthesize rather than degrade lysine, the latter activity requires two distinct genes. LYS2 encodes the alpha-aminoadipate reductase activity, while LYS5 encodes a phosphopantetheinyl transferase activity that is required to activate Lys2p. We have identified a full-length human cDNA homologous to the yeast LYS5 gene. The cDNA contains an open-reading frame of 930 bp predicted to encode 309 amino acids, and the human protein is 26% identical and 44% similar to its yeast counterpart. In Northern blot analysis the cDNA hybridizes to a single transcript of approximately 3 kb in all tissues except testis, where there is an additional transcript of 1.5 kb. Expression is highest in brain followed by heart and skeletal muscle, and to a lesser extent in liver. We further identified three human genomic BAC clones containing the human gene. Fluorescence in situ hybridization (FISH) analysis using the BAC clones mapped the gene to chromosome 11q22 while alignment of the cDNA and genomic sequences allowed partial identification of the intron-exon boundaries. Finally, using one-step homologous recombination in S. cerevisiae we generated a lys5 knockout strain. Complementation studies in the yeast knockout demonstrate that the human homolog encodes alpha-aminoadipate dehydrogenase phosphopantetheinyl transferase activity. We hypothesize that defects in this gene may result in pipecolic acidemia.

Aldehyde Oxidoreductases↗

Singletrack: an algorithm for improving memory consumption and performance of gap-affine sequence alignment.

MOTIVATION: Advances in DNA sequencing have outpaced advances in computation, making sequence alignment a major bottleneck in genome data analyses. Classical dynamic programming (DP) algorithms are particularly memory-intensive, especially when computing gap-affine and dual gap-affine alignments. Existing strategies to reduce memory consumption often sacrifice speed or alignment accuracy. RESULTS: We present Singletrack, an efficient algorithm for backtrace gap-affine and dual gap-affine alignments that requires storing a single DP matrix while preserving optimal alignment results. Compared to classical DP algorithms, Singletrack removes the need to store additional matrices (i.e. 2 for gap-affine and 4 for dual gap-affine), significantly reducing memory consumption and, in turn, reducing pressure on the memory hierarchy and improving overall performance. Most importantly, Singletrack is a general backtrace method compatible with state-of-the-art DP-based algorithms and heuristics, such as the Suzuki-Kasahara (SK) and the Wavefront Alignment (WFA) algorithms. We demonstrate that Singletrack reduces memory consumption for both SK and WFA algorithms, lowering SK usage by 2× and 4× and WFA usage by 3× and 5× for gap-affine and dual gap-affine alignments, respectively. Moreover, replacing KSW2's memory-reduction technique with Singletrack accelerates its SK implementation by up to 1.4× at the cost of doubling memory consumption, while Singletrack increases the performance of the WFA implementation in WFA2-lib by 1.2-2.1×. Compared to the efficient linear-memory BiWFA algorithm, the Singletrack-accelerated version of WFA trades a practical increase in memory usage for up to 5.2× higher performance. AVAILABILITY AND IMPLEMENTATION: The Singletrack implementations presented in this work are available on Zenodo (DOI: 10.5281/zenodo.18770585) and GitHub (https://github.com/LorienLV/singletrack).

Algorithms↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

Plant genome resources at the national center for biotechnology information.

The National Center for Biotechnology Information (NCBI) integrates data from more than 20 biological databases through a flexible search and retrieval system called Entrez. A core Entrez database, Entrez Nucleotide, includes GenBank and is tightly linked to the NCBI Taxonomy database, the Entrez Protein database, and the scientific literature in PubMed. A suite of more specialized databases for genomes, genes, gene families, gene expression, gene variation, and protein domains dovetails with the core databases to make Entrez a powerful system for genomic research. Linked to the full range of Entrez databases is the NCBI Map Viewer, which displays aligned genetic, physical, and sequence maps for eukaryotic genomes including those of many plants. A specialized plant query page allow maps from all plant genomes covered by the Map Viewer to be searched in tandem to produce a display of aligned maps from several species. PlantBLAST searches against the sequences shown in the Map Viewer allow BLAST alignments to be viewed within a genomic context. In addition, precomputed sequence similarities, such as those for proteins offered by BLAST Link, enable fluid navigation from unannotated to annotated sequences, quickening the pace of discovery. NCBI Web pages for plants, such as Plant Genome Central, complete the system by providing centralized access to NCBI's genomic resources as well as links to organism-specific Web pages beyond NCBI.

Biotechnology↗