Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

[Characterization of intertype specific epitopes on adenoviruses hexon].

BACKGROUND: To characterize the intertype epitopes on human adenovirus (HAdV) hexon. METHODS: Based on computerized analysis on adenoviruses sequence of genomic alignment, antigenicity prediction and 3-D structure characteristics of hexon subunit, several peptides of hexon of adenoviruses were chosen to be synthesized or recombinant proteins of the hexon were expressed in E. coli by use of PGEX-5X. To identify the existence of intertype epitopes, the antisera raised with synthetic peptides or purified recombinant proteins were analyzed with Western blot and immunofluorescent assay. RESULTS: The results of Western blot indicated that both peptide and recombinant antibodies showed specific reactivities with hexons of HADv-3, 4, 7 individually. Meanwhile, typical stain of immunofluorescence was found on HeLa cells infected with these HAdV by incubation with peptide as well as recombinant antibodies. Also, antibodies raised against peptide recognized the recombinant hexon protein in which a corresponding region of peptides was covered. CONCLUSIONS: Most of the predicted intertype epitopes of HAdV hexon wer e exclusively found in synthetic peptides and recombinant proteins. These intertype epitopes showed to be continuous and sequential which could be employed for development of antibodies of diagnostic use.

Adenoviruses, Human↗

Subunit VIa of yeast cytochrome c oxidase is not necessary for assembly of the enzyme complex but modulates the enzyme activity. Isolation and characterization of the nuclear-coded gene.

COX13, the nuclear gene for cytochrome c oxidase subunit VIa of Saccharomyces cerevisiae, has been isolated in two steps. First, the partial amino acid sequence information of the subunit was used to design two degenerate oligodeoxynucleotide primers to amplify part of the gene in a polymerase chain reaction. Next, the amplified product was used to screen a yeast genomic library in order to obtain the entire gene and its flanking sequences. COX13 is present as a single copy gene per haploid genome. Alignment of the N-terminal sequence of mature, subunit VIa with the amino acid sequence deduced from the DNA sequence indicates that subunit VIa is synthesized as a precursor comprised of a leader sequence of 9 amino acid residues and a mature polypeptide of 120 amino acid residues. The mature polypeptide shares 34% identical amino acid residues with the human subunit isoform VIa-L. Sequence analysis of the 3'-flanking region of COX13 revealed that the gene is located 599 base pairs downstream of CDC55, a gene which has been mapped to the left arm of chromosome VII. Null mutants of COX13, generated by gene replacement, showed a slightly reduced growth rate on nonfermentable carbon sources. Heme spectra and analysis of immunopurified cytochrome c oxidase from a null strain demonstrated that the enzyme is fully assembled without subunit VIa. At low ionic strength, cytochrome c oxidase missing subunit VIa was more active, whereas at high ionic strength, it was less active than the enzyme complex in which subunit VIa was present. In addition, distinct effects of ATP on the activity of the null and wild type enzyme were found. The results suggest that ATP interacts specifically with subunit VIa and thereby modulates the cytochrome c oxidase activity.

Adenosine Triphosphate↗

Sites other than nucleotide 234 determine cardiovirulence in natural isolates of coxsackievirus B3.

The genetic site(s) that naturally determine the cardiovirulence phenotype of coxsackievirus B3 (CVB3) have yet to be mapped. Using two closely related CVB3 strains that differed in terms of cardiovirulence phenotype in mice, we previously reported the difference in phenotype mapped to a single site, nucleotide 234 (nt234) in the 5' non-translated region (NTR) of the CVB3 genome. When nt234 was C, the virus was attenuated and when U, the virus was cardiovirulent. To determine whether this finding was applicable to other strains of CVB3, we examined 13 different naturally occurring CVB3 strains isolated in different years in the United States. We determined that only two isolates induced severe inflammatory heart muscle disease in C3H/HeJ male mice. Using PCR products as sequencing templates, we determined the 5' NTR sequence from each viral genome. Alignment of these sequences and other published CVB3 5' NTR sequences suggests as many as four separate lineages, with commonly used laboratory strains clustering closely in one branch. An examination of the sequences showed that regardless of cardiovirulence phenotype, nt234 was invariably uridine. Thus, the previously reported cytidine at nt234 is most likely the result of a rare mutation and is not a naturally occurring variation and other sites must account for the variance in virulence seen in natural isolates of CVB3.

Animals↗

Analysis of gene expression in Arabidopsis thaliana by array hybridization with genomic DNA fragments aligned along chromosomal regions.

The availability of the entire genomic sequence of the higher plant Arabidopsis thaliana prompted an analysis of chromosomal regions for gene expression with the use of high-density DNA array filters spotted with genomic DNA fragments (genomic DNA array analysis). One TAC and nine P1 clones, each of which contains a genomic DNA insert of approximately 80 kb and was used for sequencing of chromosome 5, were arbitrarily selected for analysis. The total size of the genomic regions corresponding to these clones is 819 kb. A total of 339 DNA fragments (average size, 2.9 kb) that cover contiguously the 10 chromosomal regions was selected and spotted onto nylon filters. The filters were then subjected to hybridization with (33)P-labelled cDNA molecules that had been synthesized from polyadenylated RNA isolated from 3-week-old plants. Quantitative and reproducible measurement of hybridization signals allowed analysis of the transcription of all genes in the targeted regions that were expressed at a level above the limit of detection. The data revealed that the analysed chromosomal regions are rich in active genes, and that they also provided a basis for the identification of novel transcripts whose sequences are not represented in the expressed sequence tag (EST) database.

Arabidopsis↗

Differences between pair-wise and multi-sequence alignment methods affect vertebrate genome comparisons.

Producing complete and accurate alignments of multiple genomic sequences is complex and prone to errors, especially with sequences generated from highly diverged species. In this article, we show that multi-sequence (as opposed to pair-wise) alignment methods are substantially better at aligning (or 'capturing') all of the available orthologous sequence from phylogenetically diverse vertebrates (i.e. those separated by relatively long branch lengths). Maximum gains are obtained only when sequences from many species are aligned. Such multi-sequence alignments contain significant amounts of exonic and highly conserved non-exonic sequences that are not captured in pair-wise alignments, thus illustrating the importance of the alignment method used for performing comparative genome analyses.

Animals↗

Fast and sensitive alignment of large genomic sequences.

Comparative analysis of syntenic genome sequences can be used to identify functional sites such as exons and regulatory elements. Here, the first step is to align two or several evolutionary related sequences and, in recent years, a number of computer programs have been developed for alignment of large genomic sequences. Some of these programs are extremely fast but often time-efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast heuristic identifies a chain of strong sequence similarities that serve as anchor points. In a second step, regions between these anchor points are aligned using a slower but more sensitive method. We present CHAOS, a novel algorithm for rapid identification of chains of local sequence similarities among large genomic sequences. Similarities identified by CHAOS are used as anchor points to improve the running time of the DIALIGN alignment program. Systematic test runs show that this method can reduce the running time of DIALIGN by more than 93% while affecting the quality of the resulting alignments by only 1%. The source code for CHAOS is available at http://www.stanford.edu/~brudno/chaos/ An integrated program package containing CHAOS and DIALIGN is available at http://bibiserv.techfak.uni-bielefeld.de/dialign/

Algorithms↗

Distinguishing regulatory DNA from neutral sites.

We explore several computational approaches to analyzing interspecies genomic sequence alignments, aiming to distinguish regulatory regions from neutrally evolving DNA. Human-mouse genomic alignments were collected for three sets of human regions: (1) experimentally defined gene regulatory regions, (2) well-characterized exons (coding sequences, as a positive control), and (3) interspersed repeats thought to have inserted before the human-mouse split (a good model for neutrally evolving DNA). Models that potentially could distinguish functional noncoding sequences from neutral DNA were evaluated on these three data sets, as well as bulk genome alignments. Our analyses show that discrimination based on frequencies of individual nucleotide pairs or gaps (i.e., of possible alignment columns) is only partially successful. In contrast, scoring procedures that include the alignment context, based on frequencies of short runs of alignment columns, dramatically improve separation between regulatory and neutral features. Such scoring functions should aid in the identification of putative regulatory regions throughout the human genome.

Animals↗

CGAT: a comparative genome analysis tool for visualizing alignments in the analysis of complex evolutionary changes between closely related genomes.

BACKGROUND: The recent accumulation of closely related genomic sequences provides a valuable resource for the elucidation of the evolutionary histories of various organisms. However, although numerous alignment calculation and visualization tools have been developed to date, the analysis of complex genomic changes, such as large insertions, deletions, inversions, translocations and duplications, still presents certain difficulties. RESULTS: We have developed a comparative genome analysis tool, named CGAT, which allows detailed comparisons of closely related bacteria-sized genomes mainly through visualizing middle-to-large-scale changes to infer underlying mechanisms. CGAT displays precomputed pairwise genome alignments on both dotplot and alignment viewers with scrolling and zooming functions, and allows users to move along the pre-identified orthologous alignments. Users can place several types of information on this alignment, such as the presence of tandem repeats or interspersed repetitive sequences and changes in G+C contents or codon usage bias, thereby facilitating the interpretation of the observed genomic changes. In addition to displaying precomputed alignments, the viewer can dynamically calculate the alignments between specified regions; this feature is especially useful for examining the alignment boundaries, as these boundaries are often obscure and can vary between programs. Besides the alignment browser functionalities, CGAT also contains an alignment data construction module, which contains various procedures that are commonly used for pre- and post-processing for large-scale alignment calculation, such as the split-and-merge protocol for calculating long alignments, chaining adjacent alignments, and ortholog identification. Indeed, CGAT provides a general framework for the calculation of genome-scale alignments using various existing programs as alignment engines, which allows users to compare the outputs of different alignment programs. Earlier versions of this program have been used successfully in our research to infer the evolutionary history of apparently complex genome changes between closely related eubacteria and archaea. CONCLUSION: CGAT is a practical tool for analyzing complex genomic changes between closely related genomes using existing alignment programs and other sequence analysis tools combined with extensive manual inspection.

Algorithms↗

Genome comparison without alignment using shortest unique substrings.

BACKGROUND: Sequence comparison by alignment is a fundamental tool of molecular biology. In this paper we show how a number of sequence comparison tasks, including the detection of unique genomic regions, can be accomplished efficiently without an alignment step. Our procedure for nucleotide sequence comparison is based on shortest unique substrings. These are substrings which occur only once within the sequence or set of sequences analysed and which cannot be further reduced in length without losing the property of uniqueness. Such substrings can be detected using generalized suffix trees. RESULTS: We find that the shortest unique substrings in Caenorhabditis elegans, human and mouse are no longer than 11 bp in the autosomes of these organisms. In mouse and human these unique substrings are significantly clustered in upstream regions of known genes. Moreover, the probability of finding such short unique substrings in the genomes of human or mouse by chance is extremely small. We derive an analytical expression for the null distribution of shortest unique substrings, given the GC-content of the query sequences. Furthermore, we apply our method to rapidly detect unique genomic regions in the genome of Staphylococcus aureus strain MSSA476 compared to four other staphylococcal genomes. CONCLUSION: We combine a method to rapidly search for shortest unique substrings in DNA sequences and a derivation of their null distribution. We show that unique regions in an arbitrary sample of genomes can be efficiently detected with this method. The corresponding programs shustring (SHortest Unique subSTRING) and shulen are written in C and available at http://adenine.biz.fh-weihenstephan.de/shustring/.

Algorithms↗

Multiple alignment of complete sequences (MACS) in the post-genomic era.

Multiple alignment, since its introduction in the early seventies, has become a cornerstone of modern molecular biology. It has traditionally been used to deduce structure / function by homology, to detect conserved motifs and in phylogenetic studies. There has recently been some renewed interest in the development of multiple alignment techniques, with current opinion moving away from a single all-encompassing algorithm to iterative and / or co-operative strategies. The exploitation of multiple alignments in genome annotation projects represents a qualitative leap in the functional analysis process, opening the way to the study of the co-evolution of validated sets of proteins and to reliable phylogenomic analysis. However, the alignment of the highly complex proteins detected by today's advanced database search methods is a daunting task. In addition, with the explosion of the sequence databases and with the establishment of numerous specialized biological databases, multiple alignment programs must evolve if they are to successfully rise to the new challenges of the post-genomic era. The way forward is clearly an integrated system bringing together sequence data, knowledge-based systems and prediction methods with their inherent unreliability. The incorporation of such heterogeneous, often non-consistent, data will require major changes to the fundamental alignment algorithms used to date. Such an integrated multiple alignment system will provide an ideal workbench for the validation, propagation and presentation of this information in a format that is concise, clear and intuitive.

Amino Acid Sequence↗

Multi-alignment of orthologous genome regions in five species provides new insights into the evolutionary make-up of mammalian genomes.

Evidence has shown that bacterial genomes have undergone random shuffling of genomic elements consisting of one to two genes. In order to delineate such genome-shuffling events in mammals, we constructed a high-resolution map of Sus scrofa chromosome 3 (SSC3) with a total of 116 genes/markers. Alignment of this pig map to orthologous regions in human, dog, mouse and rat led to the identification of 31 provisional conserved ancestral blocks (CABs) in these five species. Among them, only 3 CABs (<10%) had one gene, indicating that one-gene shuffling is not frequent in mammals. The sizes of CABs vary significantly within a species, but each may be relatively consistent in different species with a scale to species-genome evolution. The type and frequency of rearrangement events that takes place, either intra- or interchromosomal, depends on the evolutionary regions and species under comparison. Characterization of 36 tentative breakpoint regions flanking these 31 CABs indicated that they occupied approximately 43 Mb in length and featured genome deserts, gene duplications, and birth/death of species-specific genes in humans. Identification of CABs provides an alternative for further determination of the evolutionary make-up of mammalian genomes.

Animals↗

Alignments of mitochondrial genome arrangements: applications to metazoan phylogeny.

Mitochondrial genomes provide a valuable dataset for phylogenetic studies, in particular of metazoan phylogeny because of the extensive taxon sample that is available. Beyond the traditional sequence-based analysis it is possible to extract phylogenetic information from the gene order. Here we present a novel approach utilizing these data based on cyclic list alignments of the gene orders. A progressive alignment approach is used to combine pairwise list alignments into a multiple alignment of gene orders. Parsimony methods are used to reconstruct phylogenetic trees, ancestral gene orders, and consensus patterns in a straightforward approach. We apply this method to study the phylogeny of protostomes based exclusively on mitochondrial genome arrangements. We, furthermore, demonstrate that our approach is also applicable to the much larger genomes of chloroplasts.

Algorithms↗

Alignments anchored on genomic landmarks can aid in the identification of regulatory elements.

MOTIVATION: The transcription start site (TSS) has been located for an increasing number of genes across several organisms. Statistical tests have shown that some cis-acting regulatory elements have positional preferences with respect to the TSS, but few strategies have emerged for locating elements by their positional preferences. This paper elaborates such a strategy. First, we align promoter regions without gaps, anchoring the alignment on each promoter's TSS. Second, we apply a novel word-specific mask. Third, we apply a clustering test related to gapless BLAST statistics. The test examines whether any specific word is placed unusually consistently with respect to the TSS. Finally, our program A-GLAM, an extension of the GLAM program, uses significant word positions as new 'anchors' to realign the sequences. A Gibbs sampling algorithm then locates putative cis-acting regulatory elements. Usually, Gibbs sampling requires a preliminary masking step, to avoid convergence onto a dominant but uninteresting signal from a DNA repeat. However, since the positional anchors focus A-GLAM on the motif of interest, masking DNA repeats during Gibbs sampling becomes unnecessary. RESULTS: In a set of human DNA sequences with experimentally characterized TSSs, the placement of 791 octonucleotide words was unusually consistent (multiple test corrected P < 0.05). Alignments anchored on these words sometimes located statistically significant motifs inaccessible to GLAM or AlignACE. AVAILABILITY: The A-GLAM program and a list of statistically significant words are available at ftp://ftp.ncbi.nih.gov/pub/spouge/papers/archive/AGLAM/.

Amino Acid Motifs↗

Fast and sensitive algorithm for aligning ESTs to human genome.

There is a pressing need to align growing set of expressed sequence tags (ESTs) to newly sequenced human genome. The problem is, however, complicated by the exon/intron structure of eucaryotic genes, misread nucleotides in ESTs, and millions of repeptive sequences in genomic sequences. Indeed, to solve this, algorithms that use dynamic programming have been proposed, but in reality, these algorithms require an enormous amount of processing time. In an effort to improve the computational efficiency of these classical DP algorithms, we develop software that fully utilizes the lookup-table for allowing the efficient detection of the start- and endpoints of an EST within a given DNA sequence, and subsequently, the prompt identification of exons and introns. In addition, high sensitivity and accuracy must be achieved by calculating locations of all spliced sites correctly for more ESTs while retaining high computational efficiency. This goal is hard to accomplish in practice, owing to misread nucleotides in ESTs and repeptive sequences in the genome, but we present a couple of heuristics effective in settling this issue. Experimental results have confirmed that our technique improves the overall computation time by orders of magnitude compared with common tools such as sim4 and BLAT, and attains high sensitivity and accuracy against datasets of clean and documented genes at the same time.

Algorithms↗

Protein structural alignments and functional genomics.

Structural genomics-the systematic solution of structures of the proteins of an organism-will increasingly often produce molecules of unknown function with no close relative of known function. Prediction of protein function from structure has thereby become a challenging problem of computational molecular biology. The strong conservation of active site conformations in homologous proteins suggests a method for identifying them. This depends on the relationship between size and goodness-of-fit of aligned substructures in homologous proteins. For all pairs of proteins studied, the root-mean-square deviation (RMSD) as a function of the number of residues aligned varies exponentially for large common substructures and linearly for small common substructures. The exponent of the dependence at large common substructures is well correlated with the RMSD of the core as originally calculated by Chothia and Lesk (EMBO J 1986;5:823-826), affording the possibility of reconciling different structural alignment procedures. In the region of small common substructures, reduced aligned subsets define active sites and can be used to suggest the locations of active sites in homologous proteins.

Bacillus subtilis↗

Manipulating multiple sequence alignments via MaM and WebMaM.

MaM is a software tool that processes and manipulates multiple alignments of genomic sequence. MaM computes the exact location of common repeat elements, exons and unique regions within aligned genomics sequences using a variety of user identified programs, databases and/or tables. The program can extract subalignments, corresponding to these various regions of DNA to be analyzed independently or in conjunction with other elements of genomic DNA. Graphical displays further allow an assessment of sequence variation throughout these different regions of the aligned sequence, providing separate displays for their repeat, non-repeat and coding portions of genomic DNA. The program should facilitate the phylogenetic analysis and processing of different portions of genomic sequence as part of large-scale sequencing efforts. MaM source code is freely available for non-commercial use at http://compbio.cs.sfu.ca/MAM.htm; and the web interface WebMaM is hosted at http://atgc.lirmm.fr/mam.

Exons↗

Combining transcriptome data with genomic and cDNA sequence alignments to make confident functional assignments for Aspergillus nidulans genes.

Whole genome sequencing of several filamentous ascomycetes is complete or in progress; these species, such as Aspergillus nidulans, are relatives of Saccharomyces cerevisiae. However, their genomes are much larger and their gene structure more complex, with genes often containing multiple introns. Automated annotation programs can quickly identify open reading frames for hypothetical genes, many of which will be conserved across large evolutionary distances, but further information is required to confirm functional assignments. We describe a comparative and functional genomics approach using sequence alignments and gene expression data to predict the function of Aspergillus nidulans genes. By highlighting examples of discrepancies between the automated genome annotation and cDNA or EST sequencing, we demonstrate that the greater complexity of gene structure in filamentous fungi demands independent data on gene expression and the gene sequence be used to make confident functional assignments.

Aspergillus nidulans↗

The mouse tumor necrosis factor receptor 2 gene: genomic structure and characterization of the two transcripts.

The mouse TNFR2 gene has been cloned, sequenced, and characterized as a gene spanning >44 kb of the genome. By alignment of five genomic clones we have established that TNFR2 consists of 10 exons and 9 introns with exons ranging in size from 35 bp to 2.6 kb and introns ranging from 322 bp to >16 kb. All splice acceptor and donor sites conform to the canonical AG/GT rule. The translation initiation and termination sites are located in exon 1 and 10, respectively. Although TNFR2 lacks a canonical TATA box, the gene is transcribed from a unique start site located 70 bp upstream of the ATG initiation codon that conforms to the consensus Inr motif. Several cis-elements for transcription factors were identified in the 5' flanking region, including NF-1, Sp-1, AP2, gamma-IRE, and NF-kappaBeta motifs. Functional analysis indicates that the region -705/-412 contains a negative cis-acting element and that the minimal promoter contains motifs that confer LPS inducibility. Two mouse TNFR2 mRNAs of 3.2 and 4.1 kb are detected by Northern blot analysis, but until now their origin has not been explained. No evidence of alternative splicing of the coding exons was found. However, hybridization studies and amplification of cDNA ends suggest the use of a noncanonical polyadenylation signal in the untranslated region of exon 10. A comparative analysis of the 3' untranslated regions of the human and mouse TNFR2 genes shows highly divergent 3' ends. The possibility of an ancestral mouse TNFR2 mRNA similar to the short transcript is discussed.

Amino Acid Sequence↗