The colinear alignment of the genomes of papovaviruses JC, BK, and SV40.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The gene (Vn) encoding the mouse vitronectin was isolated and its nucleotide sequence determined. The gene covers approximately 3 kb of genomic DNA. Alignment of the genomic sequence with that of the cDNA revealed that Vn consists of eight exons, interrupted by seven introns ranging in size from 78 to 723 bp.
Connexin 45 is a gap junction protein that is prominent in early embryos and is widely expressed in many mature cell types. To elucidate its gene structure, expression, and regulation, we isolated mouse Cx45 genomic clones. Alignment of the genomic DNA and cDNA sequences revealed the presence of three exons and two introns. The first two exons contained only 5' untranslated sequences, while exon 3 contained the remaining 5' UTR, the entire coding region, and the 3' UTR. An RT-PCR with exon-specific primers was utilized to examine exon usage in F9 mouse embryonal carcinoma cells and adult mouse tissues. In all samples, PCR products amplified using exon 2/exon 3 or exon 3/exon 3 primer pairs were much more abundant than products produced using exon 1/exon 2 or exon 1/exon 3 primer pairs, suggesting that Cx45 mRNAs containing exon 1 were relatively rare compared with mRNAs containing the other exons. Rapid amplification of cDNA ends (5'-RACE) was performed using antisense primers from within exon 3 and template RNA prepared from F9 cells or from adult mouse kidney. We obtained multiple RACE products from both templates, including products that contained all three exons and were spliced identically to the cDNA. However, clones were also isolated (from kidney) that began within the region previously identified as intron 1 and continued upstream with a sequence identical to the cDNA, including splicing to exon 3. These results show that mouse Cx45 has a gene structure that differs from that of previously studied connexins and allows the production of heterogeneous Cx45 mRNAs with differing 5' UTRs. These differences might contribute to regulation of Cx45 protein levels by modulating mRNA stability or translational efficiency.
Cartilage-hair hypoplasia (CHH) is a pleiotropic disease caused by recessive mutations in the RMRP gene that result in a wide spectrum of manifestations including short stature, sparse hair, metaphyseal dysplasia, anemia, immune deficiency, and increased incidence of cancer. Molecular diagnosis of CHH has implications for management, prognosis, follow-up, and genetic counseling of affected patients and their families. We report 20 novel mutations in 36 patients with CHH and describe the associated phenotypic spectrum. Given the high mutational heterogeneity (62 mutations reported to date), the high frequency of variations in the region (eight single nucleotide polymorphisms in and around RMRP), and the fact that RMRP is not translated into protein, prediction of mutation pathogenicity is difficult. We addressed this issue by a comparative genomic approach and aligned the genomic sequences of RMRP gene in the entire class of mammals. We found that putative pathogenic mutations are located in highly conserved nucleotides, whereas polymorphisms are located in non-conserved positions. We conclude that the abundance of variations in this small gene is remarkable and at odds with its high conservation through species; it is unclear whether these variations are caused by a high local mutation rate, a failure of repair mechanisms, or a relaxed selective pressure. The marked diversity of mutations in RMRP and the low homozygosity rate in our patient population indicate that CHH is more common than previously estimated, but may go unrecognized because of its variable clinical presentation. Thus, RMRP molecular testing may be indicated in individuals with isolated metaphyseal dysplasia, anemia, or immune dysregulation.
As part of an initiative to develop Brachypodium distachyon as a genomic "bridge" species between rice and the temperate cereals and grasses, a BAC library has been constructed for the two diploid (2n = 2x = 10) genotypes, ABR1 and ABR5. The library consists of 9100 clones, with an approximate average insert size of 88 kb, representing 2.22 genome equivalents. To validate the usefulness of this species for comparative genomics and gene discovery in its larger genome relatives, the library was screened by PCR using primers designed on previously mapped rice and Poaceae sequences. Screening indicated a degree of synteny between these species and B. distachyon, which was confirmed by fluorescent in situ hybridization of the marker-selected BACs (BAC landing) to the 10 chromosome arms of the karyotype, with most of the BACs hybridizing as single loci on known chromosomes. Contiguous BACs colocalized on individual chromosomes, thereby confirming the conservation of genome synteny and proving that B. distachyon has utility as a temperate grass model species alternative to rice.
We aligned and analyzed 100 pairs of complete, orthologous intergenic regions from the human and mouse genomes (average length approximately 12 000 nucleotides). The alignments alternate between highly similar segments and dissimilar segments, indicating a wide variation of selective constraint. The average number of selectively constrained nucleotides within a mammalian intergenic region is at least 2000. This is threefold higher than within a nematode intergenic region and at least twofold higher than the number of selectively constrained nucleotides coding for an average protein. Because mammals possess only two- to threefold more proteins than Caenorhabditis elegans, the higher complexity of mammals might be primarily because of the functioning of intergenic DNA.
Analysis of multiple sequence alignments can generate important, testable hypotheses about the phylogenetic history and cellular function of genomic sequences. We describe the MultiPipMaker server, which aligns multiple, long genomic DNA sequences quickly and with good sensitivity (available at http://bio.cse.psu.edu/ since May 2001). Alignments are computed between a contiguous reference sequence and one or more secondary sequences, which can be finished or draft sequence. The outputs include a stacked set of percent identity plots, called a MultiPip, comparing the reference sequence with subsequent sequences, and a nucleotide-level multiple alignment. New tools are provided to search MultiPipMaker output for conserved matches to a user-specified pattern and for conserved matches to position weight matrices that describe transcription factor binding sites (singly and in clusters). We illustrate the use of MultiPipMaker to identify candidate regulatory regions in WNT2 and then demonstrate by transfection assays that they are functional. Analysis of the alignments also confirms the phylogenetic inference that horses are more closely related to cats than to cows.
We have hybridized all 28 chromosome-specific painting probes from the domestic sheep (Ovis aries, 2n = 54) onto metaphase chromosomes of the Indian muntjac deer (Muntiacus muntjak vaginalis, 2n = 6,7) and identified 35 conserved chromosomal segments. Results from this study show that most of the sheep acrocentric chromosomes hybridized to single regions in the Indian muntjac genome. This conserved hybridization pattern supports the concept that the large Indian muntjac chromosomes were derived from multiple tandem fusions from an ancestral deer species. Using previously reported fluorescence in situ hybridization data in which human chromosomes were hybridized onto the Indian muntjac genome, we were able to align chromosomal segments of the sheep and human genomes. Using this three-species genome alignment approach, we have identified a minimum of 42 conserved chromosomal segments between sheep and human genomes including 7 new regions not previously reported.
This study examines genomic duplications, deletions, and rearrangements that have happened at scales ranging from a single base to complete chromosomes by comparing the mouse and human genomes. From whole-genome sequence alignments, 344 large (>100-kb) blocks of conserved synteny are evident, but these are further fragmented by smaller-scale evolutionary events. Excluding transposon insertions, on average in each megabase of genomic alignment we observe two inversions, 17 duplications (five tandem or nearly tandem), seven transpositions, and 200 deletions of 100 bases or more. This includes 160 inversions and 75 duplications or transpositions of length >100 kb. The frequencies of these smaller events are not substantially higher in finished portions in the assembly. Many of the smaller transpositions are processed pseudogenes; we define a "syntenic" subset of the alignments that excludes these and other small-scale transpositions. These alignments provide evidence that approximately 2% of the genes in the human/mouse common ancestor have been deleted or partially deleted in the mouse. There also appears to be slightly less nontransposon-induced genome duplication in the mouse than in the human lineage. Although some of the events we detect are possibly due to misassemblies or missing data in the current genome sequence or to the limitations of our methods, most are likely to represent genuine evolutionary events. To make these observations, we developed new alignment techniques that can handle large gaps in a robust fashion and discriminate between orthologous and paralogous alignments.
The multiple species de novo gene prediction problem can be stated as follows: given an alignment of genomic sequences from two or more organisms, predict the location and structure of all protein-coding genes in one or more of the sequences. Here, we present a new system, N-SCAN (a.k.a. TWINSCAN 3.0), for addressing this problem. N-SCAN can model the phylogenetic relationships between the aligned genome sequences, context dependent substitution rates, and insertions and deletions. An implementation of N-SCAN was created and used to generate predictions for the entire human genome and the genome of the fruit fly Drosophila melanogaster. Analyses of the predictions reveal that N-SCAN's accuracy in both human and fly exceeds that of all previously published whole-genome de novo gene predictors.
MOTIVATION: Sequence alignment techniques have been developed into extremely powerful tools for identifying the folding families and function of proteins in newly sequenced genomes. For a sufficiently low sequence identity it is necessary to incorporate additional structural information to positively detect homologous proteins. We have carried out an extensive analysis of the effectiveness of incorporating secondary structure information directly into the alignments for fold recognition and identification of distant protein homologs. A secondary structure similarity matrix based on a database of three-dimensionally aligned proteins was first constructed. An iterative application of dynamic programming was used which incorporates linear combinations of amino acid and secondary structure sequence similarity scores. Initially, only primary sequence information is used. Subsequently contributions from secondary structure are phased in and new homologous proteins are positively identified if their scores are consistent with the predetermined error rate. RESULTS: We used the SCOP40 database, where only PDB sequences that have 40% homology or less are included, to calibrate homology detection by the combined amino acid and secondary structure sequence alignments. Combining predicted secondary structure with sequence information results in a 8-15% increase in homology detection within SCOP40 relative to the pairwise alignments using only amino acid sequence data at an error rate of 0.01 errors per query; a 35% increase is observed when the actual secondary structure sequences are used. Incorporating predicted secondary structure information in the analysis of six small genomes yields an improvement in the homology detection of approximately 20% over SSEARCH pairwise alignments, but no improvement in the total number of homologs detected over PSI-BLAST, at an error rate of 0.01 errors per query. However, because the pairwise alignments based on combinations of amino acid and secondary structure similarity are different from those produced by PSI-BLAST and the error rates can be calibrated, it is possible to combine the results of both searches. An additional 25% relative improvement in the number of genes identified at an error rate of 0.01 is observed when the data is pooled in this way. Similarly for the SCOP40 dataset, PSI-BLAST detected 15% of all possible homologs, whereas the pooled results increased the total number of homologs detected to 19%. These results are compared with recent reports of homology detection using sequence profiling methods. AVAILABILITY: Secondary structure alignment homepage at http://lutece.rutgers.edu/ssas CONTACT: anders@rutchem.rutgers.edu; ronlevy@lutece.rutgers.edu SUPPLEMENTARY INFORMATION: Genome sequence/structure alignment results at http://lutece.rutgers.edu/ss_fold_predictions.
MOTIVATION: Genome sequencing projects require the periodic application of analysis tools that can classify and multiply align related protein sequence domains. Full automation of this task requires an efficient integration of similarity and alignment techniques. RESULTS: We have developed a fully automated process that classifies entire protein sequence databases, resulting in alignment of the homologous sequences. The successive steps of the procedure are based on compositional and local sequence similarity searches followed by multiple sequence alignments. Global similarities are detected from the pairwise comparison of amino acid and dipeptide compositions of each protein. After the elimination of all but one sequence from each detected cluster of closely related proteins, the remaining sequences are compiled in a suffix tree which is self-compared to detect local sequence similarities. Sets of proteins which share similar sequence segments are then weighted according to their closeness and multiply aligned using a fast hierarchical dynamic programming algorithm. Computational strategies were devised to minimize computer processing time and memory space requirements. The accuracy of the sequence classifications has been evaluated for 12 462 primary structures distributed over 341 known families. The percentage of sequences with missed or incorrect family assignments was 6.8% on the test set. This low error level is only twice that of the manually constructed PROSITE database ( 3.4% ) and is substantially better than that found for the automatically built PRODOM database ( 34.9% ). AVAILABILITY: The resulting database, called DOMO, is available through database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr
The alignment of full-length human cDNA sequences to the finished sequence of the human genome provides a unique opportunity to study the distribution of genes throughout the genome. By analyzing the distances between 23,752 genes, we identified a class of divergently transcribed gene pairs, representing more than 10% of the genes in the genome, whose transcription start sites are separated by less than 1000 base pairs. Although this bidirectional arrangement has been previously described in humans and other species, the prevalence of bidirectional gene pairs in the human genome is striking, and the mechanisms of regulation of all but a few bidirectional genes are unknown. Our work shows that the transcripts of many bidirectional pairs are coexpressed, but some are antiregulated. Further, we show that many of the promoter segments between two bidirectional genes initiate transcription in both directions and contain shared elements that regulate both genes. We also show that the bidirectional arrangement is often conserved among mouse orthologs. These findings demonstrate that a bidirectional arrangement provides a unique mechanism of regulation for a significant number of mammalian genes.
Recent sequencing of the human and other mammalian genomes has brought about the necessity to align them, to identify and characterize their commonalities and differences. Programs that align whole genomes generally use a seed-and-extend technique, starting from exact or near-exact matches and selecting a reliable subset of these, called anchors, and then filling in the remaining portions between the anchors using a combination of local and global alignment algorithms, but their choices for the parameters so far have been primarily heuristic. We present a statistical framework and practical methods for selecting a set of matches that is both sensitive and specific and can constitute a reliable set of anchors for a one-to-one mapping of two genomes from which a whole-genome alignment can be built. Starting from exact matches, we introduce a novel per-base repeat annotation, the Z-score, from which noise and repeat filtering conditions are explored. Dynamic programming-based chaining algorithms are also evaluated as context-based filters. We apply the methods described here to the comparison of two progressive assemblies of the human genome, NCBI build 28 and build 34 (www.genome.ucsc.edu), and show that a significant portion of the two genomes can be found in selected exact matches, with very limited amount of sequence duplication.
BACKGROUND: Currently available methods to predict splice sites are mainly based on the independent and progressive alignment of transcript data (mostly ESTs) to the genomic sequence. Apart from often being computationally expensive, this approach is vulnerable to several problems--hence the need to develop novel strategies. RESULTS: We propose a method, based on a novel multiple genome-EST alignment algorithm, for the detection of splice sites. To avoid limitations of splice sites prediction (mainly, over-predictions) due to independent single EST alignments to the genomic sequence our approach performs a multiple alignment of transcript data to the genomic sequence based on the combined analysis of all available data. We recast the problem of predicting constitutive and alternative splicing as an optimization problem, where the optimal multiple transcript alignment minimizes the number of exons and hence of splice site observations. We have implemented a splice site predictor based on this algorithm in the software tool ASPIC (Alternative Splicing PredICtion). It is distinguished from other methods based on BLAST-like tools by the incorporation of entirely new ad hoc procedures for accurate and computationally efficient transcript alignment and adopts dynamic programming for the refinement of intron boundaries. ASPIC also provides the minimal set of non-mergeable transcript isoforms compatible with the detected splicing events. The ASPIC web resource is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility. CONCLUSION: Extensive bench marking shows that ASPIC outperforms other existing methods in the detection of novel splicing isoforms and in the minimization of over-predictions. ASPIC also requires a lower computation time for processing a single gene and an EST cluster. The ASPIC web resource is available at http://aspic.algo.disco.unimib.it/aspic-devel/.
Splicing is a biological phenomenon that removes the non-coding sequence from the transcripts to produce a mature transcript suitable for translation. To study this phenomenon, information on the intron-exon arrangement of a gene is essential, usually obtained by aligning mRNA/EST sequences to their cognate genomic sequences. MGAlign is a novel, rapid, memory efficient and practical method for aligning mRNA/EST and genome sequences. We present here a freely available web service, MGAlignIt (http://origin.bic.nus.edu.sg/mgalign/mgalignit), based on MGAlign. Besides the alignment itself, this web service allows users to effectively visualize the alignment in a graphical manner and to perform limited analysis on the alignment output. The server also permits the alignment to be saved in several forms, both graphical and text, suitable for further processing and analysis by other programs.
The Mouse Genome Analysis Consortium aligned the human and mouse genome sequences for a variety of purposes, using alignment programs that suited the various needs. For investigating issues regarding genome evolution, a particularly sensitive method was needed to permit alignment of a large proportion of the neutrally evolving regions. We selected a program called BLASTZ, an independent implementation of the Gapped BLAST algorithm specifically designed for aligning two long genomic sequences. BLASTZ was subsequently modified, both to attain efficiency adequate for aligning entire mammalian genomes and to increase its sensitivity. This work describes BLASTZ, its modifications, the hardware environment on which we run it, and several empirical studies to validate its results.
FELINES (Finding and Examining Lots of Intron 'N' Exon Sequences) is a utility written to automate construction and analysis of high quality intron and exon sequence databases produced from EST (expressed sequence tag) to genomic sequence alignments. We demonstrated the various programs of the FELINES utility by creating intron and exon sequence databases for the fungal organism Schizosaccharomyces pombe from alignments of EST to genomic sequences. In addition, we analyzed our constructed S.pombe sequence databases and the well-established Saccharomyces cerevisiae intron database from Manuel Ares' Laboratory for conserved sequence motifs. FELINES was shown to be useful for characterizing branchsites, polypyrimidine tracts and 5' and 3' splice sites in the intron databases and exonic splicing enhancers (ESEs) in S.pombe exons. FELINES is available at http://www.genome.ou.edu/informatics.html.