Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A new approach for gene annotation using unambiguous sequence joining.

The problem addressed by this paper is accurate and automatic gene annotation following precise identification/ annotation of exon and intron boundaries of biologically verified nucleotide sequences using the alignment of human genomic DNA to curated mRNA transcripts. We provide a detailed description of a new cDNA/DNA homology gene annotation algorithm that combines the results of BLASTN searches and spliced alignments. Compared to other programs currently in use, annotation quality is significantly increased through the unambiguous junction of genomic DNA sequences. We also address gene annotation with both non-canonic splice sites and short exons. The approach has been tested on the Genie learning subset as well as full-scale human RefSeq, and has demonstrated performance as high as 97%.

Algorithms↗

Organization of human FcRII and FcRII-like (beta FcRII) genes: structural homology to HLA class I and class II genes.

Genomic EMBL 3 DNA clones representing part of the human Fc gamma receptor II and the beta Fc gamma receptor II genes were characterized. One of them contains the first five exons including the 5' flanking region of the beta FcRII gene. The signal peptide, the extracellular domains, the putative membrane spanning region, and the first amino acids of the cytoplasmic region encoded by these five exons are spread over approximately 11 kb. Another genomic DNA clone comprises four exons encoding the second extracellular domain and the transmembrane and cytoplasmic regions of the FcRII. Alignment of the genomic DNA clones reveals that these FcR genes are identically organized. Comparison of the corresponding regions of these clones shows that not only the exons are strikingly homologous but also the splice junctions and parts of the intervening sequences are conserved. Furthermore, the genomic organization of FcRII and HLA-class I resemble each other.

Antigens, Differentiation↗

Nucleotide sequence and organization of the human S-protein gene: repeating peptide motifs in the "pexin" family and a model for their evolution.

The S-protein/vitronectin gene was isolated from a human genomic DNA library, and its sequence of about 5.3 kilobases including the adjacent 5' and 3' flanking regions was established. Alignment of the genomic DNA nucleotide sequence and the cDNA sequence indicated that the gene consisted of eight exons and seven introns. The intron positions in the S-protein gene and their phase type were compared to those in the hemopexin gene which shares amino acid sequence homologies with transin and the S-protein. Three introns have been found at equivalent positions; two other introns are very close to these positions and are interpreted as cases of intron sliding. Introns 3-7 occur at a conserved glycine residue within repeating peptide segments, whereas introns 1 and 2 are at the boundaries of the Somatomedin B domain of S-protein. The analysis of the exon structure in relation to repeating peptide motifs within the S-protein strongly suggests that it contains only seven repeats, one less than the hemopexin molecule. A very similar repeat pattern like that in hemopexin is shown to be present also in two other related proteins, transin and interstitial collagenase. An evolutionary model for the generation of the repeat pattern in the S-protein and the other members of this novel "pexin" gene family is proposed, and the sequence modifications for some of the repeats during divergent evolution are discussed in relation to known unique functional properties of hemopexin and S-protein.

Amino Acid Sequence↗

eShadow: a tool for comparing closely related sequences.

Primate sequence comparisons are difficult to interpret due to the high degree of sequence similarity shared between such closely related species. Recently, a novel method, phylogenetic shadowing, has been pioneered for predicting functional elements in the human genome through the analysis of multiple primate sequence alignments. We have expanded this theoretical approach to create a computational tool, eShadow, for the identification of elements under selective pressure in multiple sequence alignments of closely related genomes, such as in comparisons of human-to-primate or mouse-to-rat DNA. This tool integrates two different statistical methods and allows for the dynamic visualization of the resulting conservation profile. eShadow also includes a versatile optimization module capable of training the underlying Hidden Markov Model to differentially predict functional sequences. This module grants the tool high flexibility in the analysis of multiple sequence alignments and in comparing sequences with different divergence rates. Here, we describe the eShadow comparative tool and its potential uses for analyzing both multiple nucleotide and protein alignments to predict putative functional elements.

Animals↗

Genetic variation among isolates of White spot syndrome virus.

White spot syndrome virus (WSSV), member of a new virus family called Nimaviridae, is a major scourge in worldwide shrimp cultivation. Geographical isolates of WSSV identified so far are very similar in morphology and proteome, and show little difference in restriction fragment length polymorphism (RFLP) pattern. We have mapped the genomic differences between three completely sequenced WSSV isolates, originating from Thailand (WSSV-TH), China (WSSV-CN) and Taiwan (WSSV-TW). Alignment of the genomic sequences of these geographical isolates revealed an overall nucleotide identity of 99.32%. The major difference among the three isolates is a deletion of approximately 13 kb (WSSV-TH) and 1 kb (WSSV-CN), present in the same genomic region, relative to WSSV-TW. A second difference involves a genetically variable region of about 750 bp. All other variations >2 bp between the three isolates are located in repeat regions along the genome. Except for the homologous regions ( hr1, hr3, hr8 and hr9), these variable repeat regions are almost exclusively located in ORFs, of which the genomic repeat regions in ORF75, ORF94 and ORF125 can be used for PCR based classification of WSSV isolates in epidemiological studies. Furthermore, the comparison identified highly invariable genomic loci, which may be used for reliable monitoring of WSSV infections and for shrimp health certification.

Base Sequence↗

Chaos game representation for comparison of whole genomes.

BACKGROUND: Chaos game representation of genome sequences has been used for visual representation of genome sequence patterns as well as alignment-free comparisons of sequences based on oligonucleotide frequencies. However the potential of this representation for making alignment-based comparisons of whole genome sequences has not been exploited. RESULTS: We present here a fast algorithm for identifying all local alignments between two long DNA sequences using the sequence information contained in CGR points. The local alignments can be depicted graphically in a dot-matrix plot or in text form, and the significant similarities and differences between the two sequences can be identified. We demonstrate the method through comparison of whole genomes of several microbial species. Given two closely related genomes we generate information on mismatches, insertions, deletions and shuffles that differentiate the two genomes. CONCLUSION: Addition of the possibility of large scale sequence alignment to the repertoire of alignment-free sequence analysis applications of chaos game representation, positions CGR as a powerful sequence analysis tool.

Algorithms↗

[Comparative analysis of primary structure of nucleic acids and proteins].

The review considers the original works on the primary structure of biopolymers, which were carried out from 1983 to 2003. Most works were supported by the Russian program Human Genome and earlier similar Russian programs. Little-known publications of 1983-1993 and recent unpublished results are described in detail. In the field of genome comparisons, these concern the OWEN hierarchic algorithm aligning syntenic regions of two genome sequences. The resulting global alignment is obtained as an ordered chain of local similarities. Alignment of sequences sized about 10(6) nucleotides takes several minutes. The concept of local similarity conflicts is generalized to multiple comparisons. New algorithms aligning protein sequences are described and compared with the Smith-Waterman algorithm, which is now most accurate. The ANCHOR hierarchic algorithm generates alignments of much the same accuracy and is twice as rapid as the Smith-Waterman one. The STRSWer algorithm takes an account of the secondary structures of proteins under study. With the secondary structures predicted using the PSI-PRED software for pairs of proteins having 10-30% similarity, the average accuracy of alignments generated by STRSWer is 15% higher than that achieved with the Smith-Waterman algorithm.

Algorithms↗

[Study of the 3'noncoding region of Chinese hepatitis C virus genome].

OBJECTIVE: To analyze the 3' noncoding region (3' NCR) of HCV genome from Chinese hepatitis C patients so as to facilitate further study of mechanism of HCV gene replication. METHODS: Two different strategies were employed to amplify the full-length of the 3' noncoding region of HCV genome from sera of HCV infected patients in Shanghai area: one was to amplify the full-length fragment directly by nested PCR and the other amplify two overlapping fragments. The PCR products were further analyzed by sequencing and nucleotide alignments. A HCV genome 3'NCR based RT-PCR was developed and its specificity and sensitivity for HCV RNA detection in sera was compared with the established 5'NCR based RT-PCR. RESULTS: Sequence analysis showed that Chinese HCV genomic 3' NCR consists of three parts: the 5' region, poly (U-UC) tract and the 98-base region. Sequence alignments revealed that, while the 98-base regions were completely conserved in different isolates and were identical to the reported sequences, the poly (U-UC) region shared highly diversities. A high degree of concordance(95%) between the 3'NCR and 5'NCR RT-PCR for detection of HCV RNA in sera was found. CONCLUSION: The high conservation at the 3' NCR(98 bases) of HCV genome among different isolates indicated that this region may be critical for HCV gene replication The 3'NCR based RT-PCR may be a useful addition to available systems to diagnosis HCV infection.

3' Untranslated Regions↗

Linking porcine microsatellite markers to known genome regions by identifying their human orthologs.

Microsatellites, or tandem simple sequence repeats (SSRs), have become one of the most popular molecular markers in genome mapping because of their abundance across genomes and because of their high levels of polymorphism. However, information on which genes surround or flank them has remained very limited for most SSRs, especially in livestock species. In this study, an in silico comparative mapping approach was developed to link porcine SSRs to known genome regions by identifying their human orthologs. From a total of 1321 porcine microsatellites used in this study, 228 were found to have blocks in alignment with human genomic sequences. These 228 SSRs span about 1459 cM of the porcine genome, but with uneven distributions, ranging from 2 on SSC12 to 24 on SSC14. Linking these porcine SSRs to the known genome regions in the human genome also revealed 16 new putative synteny groups between these two species. Fifteen SSRs on SSC3 with identified human orthologs were typed on a pig-hamster radiation hybrid (RH) panel and used in a joint analysis with 80 known gene markers previously mapped on SSC3 using the same panel. The analysis revealed that they were all highly linked to either one or both adjacent markers. These results indicated that assigning the porcine SSRs to known genome regions by identifying their human orthologs is a reliable approach. The process will provide a foundation for positional cloning of causative genes for economically important traits.

Animals↗

Detailed alignment of saccharum and sorghum chromosomes: comparative organization of closely related diploid and polyploid genomes.

The complex polyploid genomes of three Saccharum species have been aligned with the compact diploid genome of Sorghum (2n = 2x = 20). A set of 428 DNA probes from different Poaceae (grasses) detected 2460 loci in F1 progeny of the crosses Saccharum officinarum Green German x S. spontaneum IND 81-146, and S. spontaneum PIN 84-1 x S. officinarum Muntok Java. Thirty-one DNA probes detected 226 loci in S. officinarum LA Purple x S. robustum Molokai 5829. Genetic maps of the six Saccharum genotypes, including up to 72 linkage groups, were assembled into "homologous groups" based on parallel arrangements of duplicated loci. About 84% of the loci mapped by 242 common probes were homologous between Saccharum and Sorghum. Only one interchromosomal and two intrachromosomal rearrangements differentiated both S. officinarum and S. spontaneum from Sorghum, but 11 additional cases of chromosome structural polymorphism were found within Saccharum. Diploidization was advanced in S. robustum, incipient in S. officinarum, and absent in S. spontaneum, consistent with biogeographic data suggesting that S. robustum is the ancestor of S. officinarum, but raising new questions about the antiquity of S. spontaneum. The densely mapped Sorghum genome will be a valuable tool in ongoing molecular analysis of the complex Saccharum genome.

DNA, Plant↗

Combing the genome for genomic instability.

Genomic instability is one of the major features of cancer cells. The clinical phenotypes associated with several human diseases have been linked to recurrent DNA rearrangements and dysfunction of DNA replication processes that involve unstable genomic regions. Analysis of these rearrangements, which are frequently submicroscopic and can lead to loss or gain of dosage-sensitive genes or gene disruption, requires the development of sensitive, high-resolution techniques. This will lead to a better understanding of the mechanisms underlying genome instability and a greater awareness of the role of chromosomal rearrangements in disease. A new technology that involves molecular combing, a method that permits straightening and aligning molecules of genomic DNA, should make possible a detailed analysis of genomic events at the level of single DNA molecules. Such a single molecule approach could help to elucidate important properties that are masked in bulk studies.

Cell Transformation, Neoplastic↗

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software↗

Universal primers for real-time amplification of DNA from all known Orthohepadnavirus species.

BACKGROUND: The family of Hepadnaviridae is made up of members infecting birds (genus Avihepadnavirus) or mammals (genus Orthohepadnavirus). Hepatitis B virus (HBV), the hepadnavirus infecting humans, can be divided into the seven genotypes A-G. By definition, genotypes differ by more than 8% at the nucleotide level. However, some genotypes differ by more than 14% from others. OBJECTIVES: The diversity of HBV genotypes necessitates great care in primer design to find primers suitable for routine diagnostic procedures that are highly conserved. Our aim was to find a target sequence on the HBV genome that is highly conserved among all known orthohepadnaviruses, to avoid false-negative polymerase chain reaction (PCR) results due to uncommon variants of HBV. METHODS: Using an alignment of 177 genomes of orthohepadnaviruses from GenBank, we selected a primer pair from a highly conserved region, corresponding to hydrophobic transmembrane domains of the major surface protein of HBV. RESULTS: The primer pair chosen was suitable to amplify genome sequences from HBV and to the genetically most distant woodchuck hepatitis virus in real-time PCR using the LightCycler, Roche. Moreover, the primers were suitable for accurate quantitation of both viral genomes over a range from 100 to 10(10) genomes/ml. CONCLUSION: The described primers are useful for reliable detection and accurate quantitation of all known hepadnaviral genomes and may be used for the search for unknown orthohepadnaviruses.

Animals↗

A phylogenetic survey of recombination frequency in plant RNA viruses.

The severe economic consequences of emerging plant viruses highlights the importance of studies of plant virus evolution. One question of particular relevance is the extent to which the genomes of plant viruses are shaped by recombination. To this end we conducted a phylogenetic survey of recombination frequency in a wide range of positive-sense RNA plant viruses, utilizing 975 capsid gene sequences and 157 complete genome sequences. In total, 12 of the 36 RNA virus species analyzed showed evidence for recombination, comprising 17% of the capsid gene sequence alignments and 44% of the genome sequence alignments. Given the conservative nature of our analysis, we propose that recombination is a relatively common process in some plant RNA viruses, most notably the potyviruses.

Computational Biology↗

Genome-wide assembly and analysis of alternative transcripts in mouse.

To build a mouse gene index with the most comprehensive coverage of alternative transcription/splicing (ATS), we developed an algorithm and a fully automated computational pipeline for transcript assembly from expressed sequences aligned to the genome. We identified 191,946 genomic loci, which included 27,497 protein-coding genes and 11,906 additional gene candidates (e.g., nonprotein-coding, but multiexon). Comparison of the resulting gene index with TIGR, UniGene, DoTS, and ESTGenes databases revealed that it had a greater number of transcripts, a greater average number of exons and introns with proper splicing sites per gene, and longer ORFs. The 27,497 protein-coding genes had 77,138 transcripts, i.e., 2.8 transcripts per gene on average. Close examination of transcripts led to a combinatorial table of 23 types of ATS units, only nine of which were previously described, i.e., 14 types of alternative splicing, seven types of alternative starts, and two types of alternative termination. The 47%, 18%, and 14% of 20,323 multiexon protein-coding genes with proper splice sites had alternative splicings, alternative starts, and alternative terminations, respectively. The gene index with the comprehensive ATS will provide a useful platform for analyzing the nature and mechanism of ATS, as well as for designing the accurate exon-based DNA microarrays. The sequence data from this study have been submitted to GenBank under accession numbers: CK329321-CK334090; CF891695-CF906652; CF906741-CF916750; CK334091-CK347104; CK387035-CK393993; CN660032-CN690720; CN690721-CN725493.

Algorithms↗

Reindeer papillomavirus transforming properties correlate with a highly conserved E5 region.

A papillomavirus was isolated from the epithelial layer of a cutaneous fibropapilloma on a Swedish reindeer (Rangifer tarandus). Reindeer papillomavirus (RPV) is morphologically indistinguishable from other papillomaviruses, but the restriction enzyme cleavage pattern of its genome is different. No sequence homology was detected between RPV DNA and the DNAs of bovine papillomavirus type 1 (BPV-1) and avian papillomavirus when hybridization was performed under stringent conditions. However, the RPV genome hybridized to the genome of the European elk papillomavirus and the deer papillomavirus under stringent conditions. A physical map of the RPV genome was constructed, and selected regions of the genome, covering the open translational reading frame (ORF) E5 and part of the E1 and L1 ORFs, were studied by nucleotide sequence analysis. The results made it possible to align the RPV genome with the genome of BPV-1. The E5 ORF of RPV has the potential to encode a 44-amino-acid, exceptionally hydrophobic polypeptide which is very similar to the E5 polypeptides of BPV-1 and deer and European elk papillomaviruses. RPV is oncogenic for hamsters and transforms C127 mouse cells in vitro. Several virus-specific mRNAs were detected in RPV-transformed C127 cells.

Amino Acid Sequence↗

Comparative genomic mapping of the bovine Fragile Histidine Triad (FHIT) tumour suppressor gene: characterization of a 2 Mb BAC contig covering the locus, complete annotation of the gene, analysis of cDNA and of physiological expression profiles.

BACKGROUND: The Fragile Histidine Triad gene (FHIT) is an oncosuppressor implicated in many human cancers, including vesical tumors. FHIT is frequently hit by deletions caused by fragility at FRA3B, the most active of human common fragile sites, where FHIT lays. Vesical tumors affect also cattle, including animals grazing in the wild on bracken fern; compounds released by the fern are known to induce chromosome fragility and may trigger cancer with the interplay of latent Papilloma virus. RESULTS: The bovine FHIT was characterized by assembling a contig of 78 BACs. Sequence tags were designed on human exons and introns and used directly to select bovine BACs, or compared with sequence data in the bovine genome database or in the trace archive of the bovine genome sequencing project, and adapted before use. FHIT is split in ten exons like in man, with exons 5 to 9 coding for a 149 amino acids protein. VISTA global alignments between bovine genomic contigs retrieved from the bovine genome database and the human FHIT region were performed. Conservation was extremely high over a 2 Mb region spanning the whole FHIT locus, including the size of introns. Thus, the bovine FHIT covers about 1.6 Mb compared to 1.5 Mb in man. Expression was analyzed by RT-PCR and Northern blot, and was found to be ubiquitous. Four cDNA isoforms were isolated and sequenced, that originate from an alternative usage of three variants of exon 4, revealing a size very close to the major human FHIT cDNAs. CONCLUSION: A comparative genomic approach allowed to assemble a contig of 78 BACs and to completely annotate a 1.6 Mb region spanning the bovine FHIT gene. The findings confirmed the very high level of conservation between human and bovine genomes and the importance of comparative mapping to speed the annotation process of the recently sequenced bovine genome. The detailed knowledge of the genomic FHIT region will allow to study the role of FHIT in bovine cancerogenesis, especially of vesical papillomavirus-associated cancers of the urinary bladder, and will be the basis to define the molecular structure of the bovine homologue of FRA3B, the major common fragile site of the human genome.

Acid Anhydride Hydrolases↗

Sugarcane yellow leaf virus: an emerging virus that has evolved by recombination between luteoviral and poleroviral ancestors.

We have derived the genomic nucleotide sequence of an emerging virus, the Sugarcane yellow leaf virus (ScYLV), and shown that it produces one to two subgenomic RNAs. The family Luteoviridae currently includes the Luteovirus, Polerovirus, and Enamovirus genera. With the new ScYLV nucleotide sequence and existing Luteoviridae sequence information, we have utilized new phylogenetic and evolutionary methodologies to identify homologous regions of Luteoviridae genomes, which have statistically significant altered nucleotide substitution ratios and have produced a reconstructed phylogeny of the Luteoviridae. The data indicate that Pea enation mosaic virus-1 (PEMV-1), Soybean dwarf virus (SbDV), and ScYLV exhibit spatial phylogenetic variation (SPV) consistent with recombination events that have occurred between poleroviral and luteoviral ancestors, after the divergence of these two progenitor groups. The reconstructed phylogeny confirms a contention that a continuum in the derived sequence evolution of the Luteoviridae has been established by intrafamilial as well as extrafamilial RNA recombination and expands the database of recombinant Luteoviridae genomes that are currently needed to resolve better defined means for generic discrimination in the Luteoviridae (D'Arcy, C. J. and Mayo, M. 1997. Arch. Virol. 142, 1285-1287). The analyses of the nucleotide substitution ratios from a nucleotide alignment of Luteoviridae genomes substantiates the hypothesis that hot spots for RNA recombination in this virus family are associated with the known sites for the transcription of subgenomic RNAs (Miller et al. 1995. Crit. Rev. Plant Sci. 14, 179-211), and provides new information that might be utilized to better design more effective means to generate transgene-mediated host resistance.

Amino Acid Sequence↗