Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

The chloroplast genome of Nymphaea alba: whole-genome analyses and the problem of identifying the most basal angiosperm.

Angiosperms (flowering plants) dominate contemporary terrestrial flora with roughly 250,000 species, but their origin and early evolution are still poorly understood. In recent years, molecular evidence has accumulated suggesting a dicotyledonous origin of monocots. Phylogenetic reconstructions have suggested that several dicotyledonous groups that include taxa such as Amborella, Austrobaileya, and Nymphaea branch off as the most basal among angiosperms. This has led to the concept of monocots, "eudicots," "basal dicots," and "ANITA" groupings. Here, we present the sequence and phylogenetic analyses of the chloroplast DNA of Nymphaea alba. Phylogenetic analyses of our 14-species data set, consisting of 29,991 aligned nucleotide positions per chloroplast genome, revealed consistent support for Nymphaea being a divergent member of a monophyletic dicot assemblage. Three distinct angiosperm lineages were supported in the majority of our phylogenetic analyses-eudicots, Magnoliopsida, and monocots. However, the monocot lineage leading to the grasses was the deepest branching. Although analyses of only one individual gene alignment (out of 61) is consistent with some recently proposed hypotheses for the paraphyly of dicots, we also report observations that nine genes do not support paraphyly of dicots. Instead, they support the basal monocot-dicot split. Consistent with this finding, we also report observations suggesting that the monocot lineage leading to the grasses has the strongest phylogenetic affinity to gymnosperms. Our findings have general implications for studies of substitution model specification and analyses of concatenated genome data.

Base Sequence↗

A thermostable single-strand DNase from Methanococcus jannaschii related to the RecJ recombination and repair exonuclease from Escherichia coli.

The RecJ protein of Escherichia coli plays an important role in a number of DNA repair and recombination pathways. RecJ catalyzes processive degradation of single-stranded DNA in a 5'-to-3' direction. Sequences highly related to those encoding RecJ can be found in most of the eubacterial genomes sequenced to date. From alignment of these sequences, seven conserved motifs are apparent. At least five of these motifs are shared among a large family of proteins in eubacteria, eukaryotes, and archaea, including the PPX1 polyphosphatase of yeast and Drosophila Prune. Archaeal genomes are particularly rich in such sequences, but it has not been clear whether any of the encoded proteins play a functional role similar to that of RecJ exonuclease. We have investigated three such proteins from Methanococcus jannaschii with the strongest overall sequence similarity to E. coli RecJ. Two of the genes, MJ0977 and MJ0831, partially complement a recJ mutant phenotype in E. coli. The expression of MJ0977 in E. coli resulted in high levels of a thermostable single-stranded DNase activity with properties similar to those of RecJ exonuclease. Despite overall weak sequence similarity between the MJ0977 product and RecJ, these nucleases are likely to have similar biological functions.

Acid Anhydride Hydrolases↗

Primary structural comparison of RNA-dependent polymerases from plant, animal and bacterial viruses.

Possible alignments for portions of the genomic codons in eight different plant and animal viruses are presented: tobacco mosaic, brome mosaic, alfalfa mosaic, sindbis, foot-and-mouth disease, polio, encephalomyocarditis, and cowpea mosaic viruses. Since in one of the viruses (polio) the aligned sequence has been identified as an RNA-dependent polymerase, this would imply the identification of the polymerases in the other viruses. A conserved fourteen-residue segment consisting of an Asp-Asp sequence flanked by hydrophobic residues has also been found in retroviral reverse transcriptases, a bacteriophage, influenza virus, cauliflower mosaic virus and hepatitis B virus, suggesting this span as a possible active site or nucleic acid recognition region for the polymerases. Evolutionary implications are discussed.

Animals↗

A cSNP map and database for human chromosome 21.

Single nucleotide polymorphisms (SNPs) are likely to contribute to the study of complex genetic diseases. The genomic sequence of human chromosome 21q was recently completed with 225 annotated genes, thus permitting efficient identification and precise mapping of potential cSNPs by bioinformatics approaches. Here we present a human chromosome 21 (HC21) cSNP database and the first chromosome-specific cSNP map. Potential cSNPs were generated using three approaches: (1) Alignment of the complete HC21 genomic sequence to cognate ESTs and mRNAs. Candidate cSNPs were automatically extracted using a novel program for context-dependent SNP identification that efficiently discriminates between true variation, poor quality sequencing, and paralogous gene alignments. (2) Multiple alignment of all known HC21 genes to all other human database entries. (3) Gene-targeted cSNP discovery. To date we have identified 377 cSNPs averaging ~1 SNP per 1.5 kb of transcribed sequence, covering 65% of known genes in the chromosome. Validation of our bioinformatics approach was demonstrated by a confirmation rate of 78% for the predicted cSNPs, and in total 32% of the cSNPs in our database have been confirmed. The database is publicly available at http://csnp.unige.ch or http://csnp.isb-sib.ch. These SNPs provide a tool to study the contribution of HC21 loci to complex diseases such as bipolar affective disorder and allele-specific contributions to Down syndrome phenotypes.

Base Composition↗

DAGchainer: a tool for mining segmental genome duplications and synteny.

SUMMARY: Given the positions of protein-coding genes along genomic sequence and probability values for protein alignments between genes, DAGchainer identifies chains of gene pairs sharing conserved order between genomic regions, by identifying paths through a directed acyclic graph (DAG). These chains of collinear gene pairs can represent segmentally duplicated regions and genes within a single genome or syntenic regions between related genomes. Automated mining of the Arabidopsis genome for segmental duplications illustrates the use of DAGchainer.

Algorithms↗

A generalized affine gap model significantly improves protein sequence alignment accuracy.

Sequence alignment underpins common tasks in molecular biology, including genome annotation, molecular phylogenetics, and homology modeling. Fundamental to sequence alignment is the placement of gaps, which represent character insertions or deletions. We assessed the ability of a generalized affine gap cost model to reliably detect remote protein homology and to produce high-quality alignments. Generalized affine gap alignment with optimal gap parameters performed as well as the traditional affine gap model in remote homology detection. Evaluation of alignment quality showed that the generalized affine model aligns fewer residue pairs than the traditional affine model but achieves significantly higher per-residue accuracy. We conclude that generalized affine gap costs should be used when alignment accuracy carries more importance than aligned sequence length.

Algorithms↗

Genomic organization and the 5' upstream sequences associated with the specific spatio-temporal expression of HrEpiC, an epidermis-specific gene of the ascidian Halocynthia roretzi.

An epidermis-specific gene HrEpiC of the ascidian Halocynthia roretzi is activated in all presumptive blastomeres by the 64-cell stage. To explore the molecular mechanisms underlying the regulation of the timing of activation of HrEpiC, we studied the genomic organization and the 5' upstream sequences of HrEpiC associated with specific spatio-temporal expression of the gene. The restriction site mapping and sequencing of genomic clones showed that the H. roretzi genome contained two copies of HrEpiC gene, HrEpiC1 and HrEpiC2, aligned tandemly in about 8 kb of the genome. Analysis of various deletion constructs with the 5' flanking sequences of HrEpiC1 revealed that 103 bp of the 5' flanking region was sufficient for the minimal epidermis-specific expression of HrEpiC1 and that the region between -281 bp and -198 bp of the 5' flanking region was associated with the amplification of the minimal expression of the reporter gene in the epidermis. This module between -281 bp and -198 bp was also shown to be associated with the timing of the activation of HrEpiC1 by the 64-cell stage. We discussed how the spatio-temporal expression pattern of HrEpiC1 is regulated by the two modules.

Amino Acid Sequence↗

Statistical evaluation and comparison of a pairwise alignment algorithm that a priori assigns the number of gaps rather than employing gap penalties.

MOTIVATION: Although pairwise sequence alignment is essential in comparative genomic sequence analysis, it has proven difficult to precisely determine the gap penalties for a given pair of sequences. A common practice is to employ default penalty values. However, there are a number of problems associated with using gap penalties. First, alignment results can vary depending on the gap penalties, making it difficult to explore appropriate parameters. Second, the statistical significance of an alignment score is typically based on a theoretical model of non-gapped alignments, which may be misleading. Finally, there is no way to control the number of gaps for a given pair of sequences, even if the number of gaps is known in advance. RESULTS: In this paper, we develop and evaluate the performance of an alignment technique that allows the researcher to assign a priori set of the number of allowable gaps, rather than using gap penalties. We compare this approach with the Smith-Waterman and Needleman-Wunsch techniques on a set of structurally aligned protein sequences. We demonstrate that this approach outperforms the other techniques, especially for short sequences (56-133 residues) with low similarity (<25%). Further, by employing a statistical measure, we show that it can be used to assess the quality of the alignment in relation to the true alignment with the associated optimal number of gaps. AVAILABILITY: The implementation of the described methods SANK_AL is available at http://cbbc.murdoch.edu.au/ CONTACT: matthew@cbbc.murdoch.edu.au.

Algorithms↗

Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/), a scientific database of the molecular biology and genetics of the yeast Saccharomyces cerevisiae, has recently developed several new resources that allow the comparison and integration of information on a genome-wide scale, enabling the user not only to find detailed information about individual genes, but also to make connections across groups of genes with common features and across different species. The Fungal Alignment Viewer displays alignments of sequences from multiple fungal genomes, while the Sequence Similarity Query tool displays PSI-BLAST alignments of each S.cerevisiae protein with similar proteins from any species whose sequences are contained in the non-redundant (nr) protein data set at NCBI. The Yeast Biochemical Pathways tool integrates groups of genes by their common roles in metabolism and displays the metabolic pathways in a graphical form. Finally, the Find Chromosomal Features search interface provides a versatile tool for querying multiple types of information in SGD.

Amino Acid Sequence↗

Interspecific chromosome-wide transcription profiles reveal the existence of mammalian-specific and species-specific chromosome domains.

A long-range exploration of expression levels through wide chromosome territories was carried out in three species (pig, cattle, and chicken) by aligning EST counts against the human genome. This strategy made it possible to produce expression profiles that were very similar between pig and cattle and that were significantly correlated with chicken levels of expression. In parallel with these alignments, we developed a statistical approach enabling us to screen genomic regions for both underexpression and overexpression at the chromosome level within a given species, as well as interspecifically. The observed correlations are indicative of the existence of interspecifically conserved domains of gene expression, not only for housekeeping genes (which are highly expressed), but also for regions where genes are significantly underexpressed. Furthermore, our strategy made it possible to point out regions that are differentially regulated between species. These expression data were crossed with available comparative mapping information for pigs and cattle, suggesting that coregulated regions are syntenic in various mammals.

Animals↗

A map of the common chimpanzee genome.

The completion of the chimpanzee genome will greatly help us determine which genetic changes are unique to humanity. Chimpanzees are our closest living relative, and a recent study has made considerable progress towards decoding the genome of our sister taxon.1 Over 75,000 common chimpanzee (Pan troglodytes) bacterial artificial chromosome end sequences were aligned and mapped to the human genome. This study shows the remarkable genetic similarity (98.77%) between humans and chimpanzees, while highlighting intriguing areas of potential difference. If we wish to understand the genetic basis of humankind, the completion of the chimpanzee genome deserves high priority.

Animals↗

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

Evolutionary sequence analysis of complete eukaryote genomes.

BACKGROUND: Gene duplication and gene loss during the evolution of eukaryotes have hindered attempts to estimate phylogenies and divergence times of species. Although current methods that identify clusters of orthologous genes in complete genomes have helped to investigate gene function and gene content, they have not been optimized for evolutionary sequence analyses requiring strict orthology and complete gene matrices. Here we adopt a relatively simple and fast genome comparison approach designed to assemble orthologs for evolutionary analysis. Our approach identifies single-copy genes representing only species divergences (panorthologs) in order to minimize potential errors caused by gene duplication. We apply this approach to complete sets of proteins from published eukaryote genomes specifically for phylogeny and time estimation. RESULTS: Despite the conservative criterion used, 753 panorthologs (proteins) were identified for evolutionary analysis with four genomes, resulting in a single alignment of 287,000 amino acids. With this data set, we estimate that the divergence between deuterostomes and arthropods took place in the Precambrian, approximately 400 million years before the first appearance of animals in the fossil record. Additional analyses were performed with seven, 12, and 15 eukaryote genomes resulting in similar divergence time estimates and phylogenies. CONCLUSION: Our results with available eukaryote genomes agree with previous results using conventional methods of sequence data assembly from genomes. They show that large sequence data sets can be generated relatively quickly and efficiently for evolutionary analyses of complete genomes.

Animals↗

Development and evaluation of a one-pot RPA-Cas12a assay based on a primer-driven reverse screening strategy for preliminary screening of megalocytivirus-related viruses.

A primer-driven reverse-screening strategy was used to identify an RPA-Cas12a target suitable for the rapid preliminary screening of megalocytivirus-related viruses. The ISKNV reference genome NC_003494.1 was used as the initial template, and candidate amplification units were designed according to RPA primer-design requirements, primer physicochemical properties, and the availability of Cas12a protospacer-adjacent motif (PAM) sites and crRNA target sequences. Following preliminary amplification assessment, the retained candidate primers were aligned individually against 75 complete genome sequences of megalocytivirus-related viruses. Of these, 67 sequences met the predefined criteria for target-region integrity, primer-binding-site compatibility, and Cas12a recognition. Retrospective mapping to the reference genome located the candidate amplification region within ORF057L. Based on the resulting candidate detection unit, a one-pot RPA-Cas12a assay incorporating a commercially available lyophilized RPA amplification module was developed. Optimization showed that 400&#x202f;nM reporter and 80&#x202f;nM crRNA-1 provided relatively stable fluorescence output. A cut-off value of 1281.6 relative fluorescence units (RFU) was established as the mean plus three standard deviations of the endpoint fluorescence values obtained from 20 qPCR-negative samples. In analytical sensitivity testing, the assay generated fluorescence signals above the negative control at low plasmid copy numbers. However, because only a limited number of replicates were tested at these low template concentrations, these findings were not used to define a formal limit of detection. ISKNV, RSIV, and TRBIV samples tested positive, whereas the MRV sample produced an endpoint fluorescence value below the cut-off. Repeatability analysis of the same sample in six independent reactions yielded a coefficient of variation of 8.03%. Among the 39 samples examined, no discordant qualitative results were observed between the RPA-Cas12a assay and qPCR. These findings support the use of the ORF057L-targeted one-pot RPA-Cas12a assay as a rapid preliminary screening tool for megalocytivirus-related viruses. Nevertheless, its formal limit of detection, inter-batch stability, cross-reactivity with additional non-target pathogens, and clinical diagnostic performance require further evaluation.

Lyophilized RPA↗

Identification and assessment of known and novel human papillomaviruses by polymerase chain reaction amplification, restriction fragment length polymorphisms, nucleotide sequence, and phylogenetic algorithms.

The identification and taxonomy of papillomaviruses has become increasingly complex, as approximately 70 human papillomavirus (HPV) types have been described and novel HPV genomes continue to be identified. Methods and corresponding DNA sequence data bases were designed for the reliable identification of mucosal HPV genomes from clinical specimens. HPVs are identified by the amplification of a fragment of the L1 region by consensus primer polymerase chain reaction (PCR) and subsequent hybridization or restriction fragment length polymorphism analysis. L1 PCR fragments may be further characterized by nucleotide sequencing. Conservation of 30 (of 151) predicted amino acids identifies HPV genomic fragments, and nucleotide sequence alignments allow calculation of their phylogenetic relatedness. Sequence differences > 10% from any known HPV type suggest a novel HPV type. Phylogenetic relationships with known HPV types may permit predictions of biology. With these criteria, 10 PCR fragments were identified that would qualify as new genital HPV types after complete genomic isolation.

Amino Acid Sequence↗

Polymorphisms in the non-coding region of the human mitochondrial genome in unrelated plateletapheresis donors.

Human mitochondrial DNA polymorphisms are unique targets to discriminate nucleated cells and platelets between donor and recipient in the setting of transplantation or transfusion. We have previously used this approach to discriminate allogeneic platelets from autologous platelets after transfusion. In the present study, we used DNA sequencing to investigate polymorphisms present in two of the hypervariable segments (HVR1 and HVR2) found within the non-coding region of the mitochondrial genome among 100 plateletapheresis donors. Alignments were made with the Cambridge Reference Sequence (CRS) for human mitochondrial DNA (mtDNA). Combining the sequencing information of HVR1 and HVR2 we could demonstrate that, of the 100 investigated mtDNA samples, none was identical to the CRS. We found a total of 2-17 polymorphisms per donor in the investigated regions, most of them were basepair substitutions (563) and insertions (151). No deletions were found. Sixty-six of the 110 detected polymorphisms were detected in more than one sample. Seven polymorphisms are newly described and have not been published in the Mitomap database. Our results demonstrate that polymerase chain reaction analysis of the many polymorphisms found in the hypervariable region of mitochondrial DNA represents a more informative target than previously described mitochondrial polymorphisms for discriminating donor-recipient cells after transfusion or transplantation.

Blood Platelets↗

GenDiS: Genomic Distribution of protein structural domain Superfamilies.

Several proteins that have substantially diverged during evolution retain similar three-dimensional structures and biological function inspite of poor sequence identity. The database on Genomic Distribution of protein structural domain Superfamilies (GenDiS) provides record for the distribution of 4001 protein domains organized as 1194 structural superfamilies across 18,997 genomes at various levels of hierarchy in taxonomy. GenDiS database provides a survey of protein domains enlisted in sequence databases employing a 3-fold sequence search approach. Lineage-specific literature is obtained from the taxonomy database for individual protein members to provide a platform for performing genomic and phyletic studies across organisms. The database documents residual properties and provides alignments for the various superfamily members in genomes, offering insights into the rational design of experiments and for the better understanding of a superfamily. GenDiS database can be accessed at http://www.ncbs.res.in/~faculty/mini/gendis/home.html.

Databases, Protein↗

Analysis of Chlamydomonas reinhardtii genome structure using large-scale sequencing of regions on linkage groups I and III.

Chlamydomonas reinhardtii is a unicellular green alga that has been used as a model organism for the study of flagella and basal bodies as well as photosynthesis. This report analyzes finished genomic DNA sequence for 0.5% of the nuclear genome. We have used three gene prediction programs as well as EST and protein homology data to estimate the total number of genes in Chlamydomonas to be between 12,000 and 16,400. Chlamydomonas appears to have many more genes than any other unicellular organism sequenced to date. Twenty-seven percent of the predicted genes have significant identity to both ESTs and to known proteins in other organisms, 32% of the predicted genes have significant identity to ESTs alone, and 14% have significant similarity to known proteins in other organisms. For gene prediction in Chlamydomonas, GreenGenie appeared to have the highest sensitivity and specificity at the exon level, scoring 71% and 82%. respectively. Two new alternative splicing events were predicted by aligning Chlamydomonas ESTs to the genomic sequence. Finally recombination differs between the two sequenced contigs. The 350-Kb of the Linkage group III contig is devoid of recombination, while the Linkage group I contig is 30 map units long over 33-kb.

Amino Acid Sequence↗