Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

ASD: a bioinformatics resource on alternative splicing.

Alternative splicing is an important regulatory mechanism of mammalian gene expression. The alternative splicing database (ASD) consortium is systematically collecting and annotating data on alternative splicing. We present the continuation and upgrade of the ASD [T. A. Thanaraj, S. Stamm, F. Clark, J. J. Riethoven, V. Le Texier, J. Muilu (2004) Nucleic Acids Res. 32, D64-D69] that consists of computationally and manually generated data. Its largest parts are AltSplice, a value-added database of computationally delineated alternative splicing events. Its data include alternatively spliced introns/exons, events, isoform splicing patterns and isoform peptide sequences. AltSplice data are generated by examining gene-transcript alignments. The data are annotated for various biological features including splicing signals, expression states, (SNP)-mediated splicing and cross-species conservation. AEdb forms the manually curated component of ASD. It is a literature-based data set containing sequence and properties of alternatively spliced exons, functional enumeration of observed splicing events, characterization of observed splicing regulatory elements, and a collection of experimentally clarified minigene constructs. ASD includes a workbench, which is an analysis tool that enables users to carry out splicing related analysis such as characterization of introns for various splicing signals, identification of splicing regulatory elements on a given RNA sequence, prediction of putative exons and prediction of putative translation start codons. The different ASD modules are integrated and can be accessed through user-friendly interfaces and visualization tools. ASD data has been integrated with Ensembl genome annotation project as a Distributed Annotation System (DAS) resource and can be viewed on Ensembl genome browser. The ASD resource is presented at (http://www.ebi.ac.uk/asd).

Alternative Splicing↗

Comparative sequence analysis of Sordaria macrospora and Neurospora crassa as a means to improve genome annotation.

One of the most challenging parts of large scale sequencing projects is the identification of functional elements encoded in a genome. Recently, studies of genomes of up to six different Saccharomyces species have demonstrated that a comparative analysis of genome sequences from closely related species is a powerful approach to identify open reading frames and other functional regions within genomes [Science 301 (2003) 71, Nature 423 (2003) 241]. Here, we present a comparison of selected sequences from Sordaria macrospora to their corresponding Neurospora crassa orthologous regions. Our analysis indicates that due to the high degree of sequence similarity and conservation of overall genomic organization, S. macrospora sequence information can be used to simplify the annotation of the N. crassa genome.

Base Sequence↗

Recognition of human genes by stochastic parsing.

A gene finding system, GeneDecoder, based on a parsing technique using a stochastic grammar and dictionary of genetic words is introduced. The structure of human genes are expressed by a stochastic grammar and a dictionary, whose components are the genetic words consisting of genetic phonemes, built as hidden Markov models (HMMs). The HMMs represent the nucleotide acid bases, the codons, and the amino acids. The genetic words in the dictionary are described by the sequence of these HMMs and represent exons, introns, intergenic regions, tRNA regions and signals in DNA sequences. The statistics between these regions are expressed by the grammar, which is a stochastic network of the genetic words. Using the same kind of technique of speech recognition by HMMs with a word dictionary and a grammar, the stochastic network of genetic words enables the motif dictionary to be used during the parsing of the DNA sequences. At the same time, stochastic features of donor/acceptor sites, information of the di-codon statistics, and other important features are integrated into stochastic scores during the parsing. As a result, while the system parses DNA sequences and finds the exon/intron structures, the protein motifs are automatically annotated in the regions. It helps to identify the functions of the genes and reduces the cost of homology search for each hypothetical coding regions. This method is different from simply using the information of homology search. This method uses the information of the motif patterns during the parsing process, but searching the motif patterns after/before finding the coding regions cannot directly affect the parsing process itself. Experimental results have shown that this method reasonably finds and annotates the motifs in the exons in the DNA sequence of human.

Amino Acid Sequence↗

The diversity of Rab GTPases in Entamoeba histolytica.

Rab proteins are ubiquitous small GTP-binding proteins that form a highly conserved family and regulate vesicular trafficking. Recent completion of the genome of the enteric protozoan parasite Entamoeba histolytica enabled us to identify an extremely large number (>90) of putative Rab genes. Multiple alignment and phylogenic analysis of amebic, human, and yeast Rab showed that only 22 amebic Rab proteins including EhRab1, EhRab2, EhRab5, EhRab7, EhRab8, EhRab11, and EhRab21 showed significant similarity to Rab from other organisms. The 69 remaining amebic Rab proteins showed only moderate similarity (<40% identity) to Rab proteins from other organisms. Approximately one-third of Rab proteins including Rab7, Rab11, and RabC form 15 subfamilies, which contain up to nine isoforms. Approximately 70% of amebic Rab genes contain single or multiple introns, and this proportion is significantly higher than that of common genes in this organism. Twenty-five Rabs possess an atypical carboxyl terminus such as CXXX, XCXX, XXCX, XXXC, and no cysteine. We propose annotation of amebic Rab genes and discuss biological significance of this extraordinary diversity of EhRab proteins in this organism.

Amino Acid Sequence↗

Comprehensive identification of Drosophila dorsal-ventral patterning genes using a whole-genome tiling array.

Dorsal-ventral (DV) patterning of the Drosophila embryo is initiated by Dorsal, a sequence-specific transcription factor distributed in a broad nuclear gradient in the precellular embryo. Previous studies have identified as many as 70 protein-coding genes and one microRNA (miRNA) gene that are directly or indirectly regulated by this gradient. A gene regulation network, or circuit diagram, including the functional interconnections among 40 Dorsal target genes and 20 associated tissue-specific enhancers, has been determined for the initial stages of gastrulation. Here, we attempt to extend this analysis by identifying additional DV patterning genes using a recently developed whole-genome tiling array. This analysis led to the identification of another 30 protein-coding genes, including the Drosophila homolog of Idax, an inhibitor of Wnt signaling. In addition, remote 5' exons were identified for at least 10 of the approximately 100 protein-coding genes that were missed in earlier annotations. As many as nine intergenic uncharacterized transcription units were identified, including two that contain known microRNAs, miR-1 and -9a. We discuss the potential functions of these recently identified genes and suggest that intronic enhancers are a common feature of the DV gene network.

Animals↗

Genomic divergence between human and chimpanzee estimated from large-scale alignments of genomic sequences.

To study the genomic divergence between human and chimpanzee, large-scale genomic sequence alignments were performed. The genomic sequences of human and chimpanzee were first masked with the RepeatMasker and the repeats were excluded before alignments. The repeats were then reinserted into the alignments of nonrepetitive segments and entire sequences were aligned again. A total of 2.3 million base pairs (Mb) of genomic sequences, including repeats, were aligned and the average nucleotide divergence was estimated to be 1.22%. The Jukes-Cantor (JC) distances (nucleotide divergences) in nonrepetitive (1.44 Mb) and repetitive sequences (0.86 Mb) are 1.14% and 1.34%, respectively, suggesting a slightly higher average rate in repetitive sequences. Annotated coding and noncoding regions of homologous chimpanzee genes were also retrieved from GenBank and compared. The average synonymous and nonsynonymous divergences in 88 coding genes are 1.48% and 0.55%, respectively. The JC distances in intron, 5' flanking, 3' flanking, promoter, and pseudogene regions are 1.47%, 1.41%, 1.68%, 0.75%, and 1.39%, respectively. It is not clear why the genetic distances in most of these regions are somewhat higher than those in genomic sequences. One possible explanation is that some of the genes may be located in regions with higher mutation rates.

Animals↗

Molecular identification of the first insect ecdysis triggering hormone receptors.

The Drosophila Genome Project website (www.flybase.org) contains an annotated gene sequence (CG5911), coding for a G protein-coupled receptor. We cloned the cDNA corresponding to this sequence and found that the gene has not been correctly predicted. The corrected gene CG5911 has five introns and six exons (1-6). Alternative splicing yields two cDNAs called A (containing exons 1-5) and B (containing exons 1-4, 6). We expressed these splicing variants in Chinese hamster ovary cells and found that the corrected CG5911-A and -B cDNAs coded for two different G protein-coupled receptors that could be activated by low concentrations of Drosophila ecdysis triggering hormones-1 and -2. Ecdysis (cuticle shedding) is an important behaviour, allowing growth and metamorphosis in insects and other arthropods. Our paper is the first report on the molecular identification of ecdysis triggering hormone receptors from insects.

Alternative Splicing↗

Virtual Ribosome--a comprehensive DNA translation tool with support for integration of sequence feature annotation.

Virtual Ribosome is a DNA translation tool with two areas of focus. (i) Providing a strong translation tool in its own right, with an integrated ORF finder, full support for the IUPAC degenerate DNA alphabet and all translation tables defined by the NCBI taxonomy group, including the use of alternative start codons. (ii) Integration of sequences feature annotation--in particular, native support for working with files containing intron/exon structure annotation. The software is available for both download and online use at http://www.cbs.dtu.dk/services/VirtualRibosome/.

Animals↗

Dissecting the genetic basis underlying drought tolerance at different development stages in soybean.

INTRODUCTION: Soybean is an indispensable crop supplying protein and oil for humans and animals, and playing an essential role in global food security. Drought represses soybean seed germination, reducing biomass accumulation and even inhibiting yield. METHODS: In order to dissect the genetic components underlying soybean drought tolerance during different development stage, a natural population containing 140 accessions was employed to evaluate seven drought tolerance-related traits under water-welled and drought stress conditions. Subsequently, genome-wide association study (GWAS) was conducted based on 150K single nucleotide polymorphism (SNP) markers of "Zhongdouxin-1". And the drought tolerance coefficient of seven different traits were analyzed with seven GWAS models. RESULTS: A total of 1807 significant SNPs were detected across 20 chromosome, including 569 SNPs for germination stage, and 1242 SNPs for seedling stage. Of 569 SNPs identified in germination stage, 354 SNPs on chromosomes 2, 7, 13, 14, and 17 accounting for 62.21%. Among 1242 SNPs found in seedling stage, 869 SNPs on chromosomes 11, 14, 15, 17 and 18 accounting for 69.97%. Moreover, among 1807 significant SNPs, 163 SNPs exhibited pleiotropic effects, of which 23 were located in exon, 21 in intron, 12 in 5'UTR or 3'UTR and 11 in upstream or downstream. Furthermore, 249 stable SNPs were detected by more than four GWAS models. According to these stable SNPs, RNA expression levels and gene annotations, four causal genes (Glyma.02G080200, Glyma.11G056200, Glyma.12G188900, and Glyma.18G110200) conferring soybean drought tolerance were detected, which participated in ethylene stimulus response, water deprivation response, and proteolysis. DISCUSSION: Collectively, 249 stable SNPs, 163 pleiotropic SNPs and four candidate genes identified in present study provided promising molecular resources and reliable foundation for drought resistance improvement and marker-assisted selective breeding in soybean.

GWAS↗

Target SNP selection in complex disease association studies.

BACKGROUND: The massive amount of SNP data stored at public internet sites provides unprecedented access to human genetic variation. Selecting target SNP for disease-gene association studies is currently done more or less randomly as decision rules for the selection of functional relevant SNPs are not available. RESULTS: We implemented a computational pipeline that retrieves the genomic sequence of target genes, collects information about sequence variation and selects functional motifs containing SNPs. Motifs being considered are gene promoter, exon-intron structure, AU-rich mRNA elements, transcription factor binding motifs, cryptic and enhancer splice sites together with expression in target tissue. As a case study, 396 genes on chromosome 6p21 in the extended HLA region were selected that contributed nearly 20,000 SNPs. By computer annotation ~2,500 SNPs in functional motifs could be identified. Most of these SNPs are disrupting transcription factor binding sites but only those introducing new sites had a significant depressing effect on SNP allele frequency. Other decision rules concern position within motifs, the validity of SNP database entries, the unique occurrence in the genome and conserved sequence context in other mammalian genomes. CONCLUSION: Only 10% of all gene-based SNPs have sequence-predicted functional relevance making them a primary target for genotyping in association studies.

Amino Acid Substitution↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

Full-length messenger RNA sequences greatly improve genome annotation.

BACKGROUND: Annotation of eukaryotic genomes is a complex endeavor that requires the integration of evidence from multiple, often contradictory, sources. With the ever-increasing amount of genome sequence data now available, methods for accurate identification of large numbers of genes have become urgently needed. In an effort to create a set of very high-quality gene models, we used the sequence of 5,000 full-length gene transcripts from Arabidopsis to re-annotate its genome. We have mapped these transcripts to their exact chromosomal locations and, using alignment programs, have created gene models that provide a reference set for this organism. RESULTS: Approximately 35% of the transcripts indicated that previously annotated genes needed modification, and 5% of the transcripts represented newly discovered genes. We also discovered that multiple transcription initiation sites appear to be much more common than previously known, and we report numerous cases of alternative mRNA splicing. We include a comparison of different alignment software and an analysis of how the transcript data improved the previously published annotation. CONCLUSIONS: Our results demonstrate that sequencing of large numbers of full-length transcripts followed by computational mapping greatly improves identification of the complete exon structures of eukaryotic genes. In addition, we are able to find numerous introns in the untranslated regions of the genes.

Alternative Splicing↗

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem↗

Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.

BACKGROUND: This paper describes Pairagon+N-SCAN_EST, a gene annotation pipeline that uses only native alignments. For each expressed sequence it chooses the best genomic alignment. Systems like ENSEMBL and ExoGean rely on trans alignments, in which expressed sequences are aligned to the genomic loci of putative homologs. Trans alignments contain a high proportion of mismatches, gaps, and/or apparently unspliceable introns, compared to alignments of cDNA sequences to their native loci. The Pairagon+N-SCAN_EST pipeline's first stage is Pairagon, a cDNA-to-genome alignment program based on a PairHMM probability model. This model relies on prior knowledge, such as the fact that introns must begin with GT, GC, or AT and end with AG or AC. It produces very precise alignments of high quality cDNA sequences. In the genomic regions between Pairagon's cDNA alignments, the pipeline combines EST alignments with de novo gene prediction by using N-SCAN_EST. N-SCAN_EST is based on a generalized HMM probability model augmented with a phylogenetic conservation model and EST alignments. It can predict complete transcripts by extending or merging EST alignments, but it can also predict genes in regions without EST alignments. Because they are based on probability models, both Pairagon and N-SCAN_EST can be trained automatically for new genomes and data sets. RESULTS: On the ENCODE regions of the human genome, Pairagon+N-SCAN_EST was as accurate as any other system tested in the EGASP assessment, including ENSEMBL and ExoGean. CONCLUSION: With sufficient mRNA/EST evidence, genome annotation without trans alignments can compete successfully with systems like ENSEMBL and ExoGean, which use trans alignments.

Base Sequence↗

Comparative genomic mapping of the bovine Fragile Histidine Triad (FHIT) tumour suppressor gene: characterization of a 2 Mb BAC contig covering the locus, complete annotation of the gene, analysis of cDNA and of physiological expression profiles.

BACKGROUND: The Fragile Histidine Triad gene (FHIT) is an oncosuppressor implicated in many human cancers, including vesical tumors. FHIT is frequently hit by deletions caused by fragility at FRA3B, the most active of human common fragile sites, where FHIT lays. Vesical tumors affect also cattle, including animals grazing in the wild on bracken fern; compounds released by the fern are known to induce chromosome fragility and may trigger cancer with the interplay of latent Papilloma virus. RESULTS: The bovine FHIT was characterized by assembling a contig of 78 BACs. Sequence tags were designed on human exons and introns and used directly to select bovine BACs, or compared with sequence data in the bovine genome database or in the trace archive of the bovine genome sequencing project, and adapted before use. FHIT is split in ten exons like in man, with exons 5 to 9 coding for a 149 amino acids protein. VISTA global alignments between bovine genomic contigs retrieved from the bovine genome database and the human FHIT region were performed. Conservation was extremely high over a 2 Mb region spanning the whole FHIT locus, including the size of introns. Thus, the bovine FHIT covers about 1.6 Mb compared to 1.5 Mb in man. Expression was analyzed by RT-PCR and Northern blot, and was found to be ubiquitous. Four cDNA isoforms were isolated and sequenced, that originate from an alternative usage of three variants of exon 4, revealing a size very close to the major human FHIT cDNAs. CONCLUSION: A comparative genomic approach allowed to assemble a contig of 78 BACs and to completely annotate a 1.6 Mb region spanning the bovine FHIT gene. The findings confirmed the very high level of conservation between human and bovine genomes and the importance of comparative mapping to speed the annotation process of the recently sequenced bovine genome. The detailed knowledge of the genomic FHIT region will allow to study the role of FHIT in bovine cancerogenesis, especially of vesical papillomavirus-associated cancers of the urinary bladder, and will be the basis to define the molecular structure of the bovine homologue of FRA3B, the major common fragile site of the human genome.

Acid Anhydride Hydrolases↗

SNPSplicer: systematic analysis of SNP-dependent splicing in genotyped cDNAs.

Functional annotation of SNPs (as generated by HapMap (http://www.hapmap.org) for instance) is a major challenge. SNPs that lead to single amino acid substitutions, stop codons, or frameshift mutations can be readily interpreted, but these represent only a fraction of known SNPs. Many SNPs are located in sequences of splicing relevance-the canonical splice site consensus sequences, exonic and intronic splice enhancers or silencers (exonic splice enhancer [ESE], intronic splice enhancer [ISE], exonic splicing silencer [ESS], and intronic splicing silencer [ISS]), and others. We propose using sets of matching DNA and complementary DNA (cDNA) as a screening method to investigate the potential splice effects of SNPs in RT-PCR experiments with tissue material from genotyped sources. We have developed a software solution (SNPSplicer; http://www.ikmb.uni-kiel.de/snpsplicer) that aids in the rapid interpretation of such screening experiments. The utility of the approach is illustrated for SNPs affecting the donor splice sites (rs2076530:A>G, rs3816989:G>A) leading to the use of a cryptic splice site and exon skipping, respectively, and an exonic splice enhancer SNP (rs2274987:C/T), leading to inclusion of a new exon. We anticipate that this methodology may help in the functional annotation of SNPs in a more high-throughput fashion.

Alternative Splicing↗

Improved splice site detection in Genie.

We present an improved splice site predictor for the genefinding program Genie. Genie is based on a generalized Hidden Markov Model (GHMM) that describes the grammar of a legal parse of a multi-exon gene in a DNA sequence. In Genie, probabilities are estimated for gene features by using dynamic programming to combine information from multiple content and signal sensors, including sensors that integrate matches to homologous sequences from a database. One of the hardest problems in genefinding is to determine the complete gene structure correctly. The splice site sensors are the key signal sensors that address this problem. We replaced the existing splice site sensors in Genie with two novel neural networks based on dinucleotide frequencies. Using these novel sensors, Genie shows significant improvements in the sensitivity and specificity of gene structure identification. Experimental results in tests using a standard set of annotated genes showed that Genie identified 86% of coding nucleotides correctly with a specificity of 85%, versus 80% and 84% in the older system. In further splice site experiments, we also looked at correlations between splice site scores and intron and exon lengths, as well as at the effect of distance to the nearest splice site on false positive rates.

Animals↗

Prediction of locally optimal splice sites in plant pre-mRNA with applications to gene identification in Arabidopsis thaliana genomic DNA.

Prediction of splice site selection and efficiency from sequence inspection is of fundamental interest (testing the current knowledge of requisite sequence features) and practical importance (genome annotation, design of mutant or transgenic organisms). In plants, the dominant variables affecting splice site selection and efficiency include the degree of matching to the extended splice site consensus and the local gradient of U- and G+C-composition (introns being U-rich and exons G+C-rich). We present a novel method for splice site prediction, which was particularly trained for maize and Arabidopsis thaliana. The method extends our previous algorithm based on logitlinear models by considering three variables simultaneously: intrinsic splice site strength, local optimality and fit with respect to the overall splice pattern prediction. We show that the method considerably improves prediction specificity without compromising the high degree of sensitivity required in gene prediction algorithms. Applications to gene identification are illustrated for Arabidopsis and suggest that successful methods must combine scoring for splice sites, coding potential and similarity with potential homologs in non-trivial ways. A WWW version of the SplicePredictor program is available at http:/gnomic.stanford.edu/volker/SplicePredi ctor.html/

Algorithms↗