Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Two-dimensional gel protein database of Saccharomyces cerevisiae.

With the systematic sequencing of the yeast genome, yeast biology has entered a new era where novel challenges have to be faced. One challenge is the identification of the function of the several hundred novel genes discovered by genome sequencing. Another is to understand how all yeast genes act in concert to ensure and maintain cell organization. Two-dimensional (2-D) gel electrophoresis is the technique of choice to take up these challenges because it provides the opportunity of obtaining an overall view of genome expression. In prospect of these studies we have undertaken the construction of a yeast 2-D gel protein database that contains information on polypeptides of the yeast protein map. In this paper we report the information presently contained in this database. The reported information includes the identification of 250 protein spots and the characterization of polypeptides corresponding to N-terminal acetylated proteins, mitochondrial proteins, glucose-repressed proteins, heat shock induced proteins and proteins encoded by intron-containing genes. In all, 600 spots are annotated. These data can be accessed on the Yeast Protein Map server through the World Wide Web network.

Computer Communication Networks↗

Beyond Exons: Linking Noncoding Heritability and Polygenicity across Complex Human Traits and Disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across exonic, intronic, and intergenic regions for 34 traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons explain a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of a broader set of functional annotations reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less polygenic traits show stronger contributions in promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, pointing to a shift from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects as a key determinant of differences in polygenicity across traits.

Journal Article↗

SSCP-SNP in pearl millet--a new marker system for comparative genetics.

A considerable array of genomic resources are in place in pearl millet, and marker-aided selection is already in use in the public breeding programme at ICRISAT. This paper describes experiments to extend these publicly available resources to a single nucleotide polymorphism (SNP)-based marker system. A new marker system, single-strand conformational polymorphism (SSCP)-SNP, was developed using annotated rice genomic sequences to initially predict the intron-exon borders in millet expressed sequence tags (ESTs) and then to design primers that would amplify across the introns. An adequate supply of millet ESTs was available for us to identify 299 homologues of single-copy rice genes in which the intron positions could be precisely predicted. PCR primers were then designed to amplify approximately 500-bp genomic fragments containing introns. Analysis of these fragments on SSCP gels revealed considerable polymorphism. A detailed DNA sequence analysis of variation at four of the SSCP-SNP loci over a panel of eight inbred genotypes showed complex patterns of variation, with about one SNP or indel (insertion-deletion) every 59 bp in the introns, but considerably fewer in the exons. About two-thirds of the variation was derived from SNPs and one-third from indels. Most haplotypes were detected by SSCP. As a marker system, SSCP-SNP has lower development costs than simple sequence repeats (SSRs), because much of the work is in silico, and similar deployment costs and through-put potential. The rates of polymorphism were lower but useable, with a mean PIC of 0.49 relative to 0.72 for SSRs in our eight inbred genotype panel screen. The major advantage of the system is in comparative applications. Syntenic information can be used to target SSCP-SNP markers to specific chromosomal regions or, conversely, SSCP-SNP markers can be used to unravel detailed syntenic relationships in specific parts of the genome. Finally, a preliminary analysis showed that the millet SSCP-SNP primers amplified in other cereals with a success rate of about 50%. There is also considerable potential to promote SSCP-SNP to a COS (conserved orthologous set) marker system for application across species by more specifically designing primers to precisely match the model genome sequence.

Base Sequence↗

Generation of p53 target database via integration of microarray and global p53 DNA-binding site analysis.

The completion of the human genome sequence and availability of cDNA microarray technology provide new approaches to explore global cellular regulatory mechanisms. Here we present a strategy to identify genes regulated by specific transcription factors in the human genome, and apply it to p53. We first collected promoters or introns of all genes available using two methods: GenBank annotation and a computationally derived transcript map. The "FindPatterns" program is then used to search sequences in regulatory regions that match the p53 DNA-binding consensus sequence, resulting in the p53 Target Database. This database collects human genes that have at least one p53 DNA-binding sequence in their regulatory region. cDNA microarray was also used to identify genes that respond to p53 at a genomic scale. Integration of the microarray data and the p53 Target Database should greatly enrich direct p53 target genes. Taqman analysis and quantitative chromatin immunoprecipitation analysis are used to validate the in silico prediction and microarray data. Enrichment factor analysis is used to demonstrate that in silico prediction greatly enriches for genes that are transcriptionally regulated by p53 and assists us to identify other signaling pathways that are potentially connected to p53. The approaches can be extended to other transcription factors. The methods shown here illustrate a novel approach to the analysis of global gene regulatory networks through the integration of human genomic sequence information and genome-wide gene expression analysis.

Binding Sites↗

Recognition of unknown conserved alternatively spliced exons.

The split structure of most mammalian protein-coding genes allows for the potential to produce multiple different mRNA and protein isoforms from a single gene locus through the process of alternative splicing (AS). We propose a computational approach called UNCOVER based on a pair hidden Markov model to discover conserved coding exonic sequences subject to AS that have so far gone undetected. Applying UNCOVER to orthologous introns of known human and mouse genes predicts skipped exons or retained introns present in both species, while discriminating them from conserved noncoding sequences. The accuracy of the model is evaluated on a curated set of genes with known conserved AS events. The prediction of skipped exons in the approximately 1% of the human genome represented by the ENCODE regions leads to more than 50 new exon candidates. Five novel predicted AS exons were validated by RT-PCR and sequencing analysis of 15 introns with strong UNCOVER predictions and lacking EST evidence. These results imply that a considerable number of conserved exonic sequences and associated isoforms are still completely missing from the current annotation of known genes. UNCOVER also identifies a small number of candidates for conserved intron retention.

Journal Article↗

Analyses of p53 target genes in the human genome by bioinformatic and microarray approaches.

The completion of the human genome sequence (International Human Genome Sequence Consortium (2001) Nature 409, 860-921; Venter, J. C., et al. (2001) Science 291, 1304-1351) allows for new ways to analyze global cellular regulatory mechanisms. Here we present a strategy to identify genes regulated by specific transcription factors in the human genome, and apply it to p53. We first collected promoters or introns of all genes available using two methods: GenBank(TM) annotation and a computationally derived transcript map. 4,852 genes analyzed in this way contained at least one p53 consensus binding sequence. Of 13 genes randomly selected for mRNA analysis, 11 were shown to respond to p53 expression. Five promoters were analyzed by chromatin immunoprecipitation, which revealed that all were bound by p53 in vivo. We then analyzed 33,615 unique human genes on cDNA microarrays, identifying 1,501 genes that respond to p53 expression. A parameter was derived that demonstrates that in silico prediction greatly enriches for genes that are activated and repressed by p53 and assists us to suggest other signaling pathways that may be connected to p53. The methods shown here illustrate a novel approach to analysis of global gene regulatory network through the integration of human genomic sequence information and genome-wide gene expression analysis.

Computational Biology↗

Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22.

In this report, we have achieved a richer view of the transcriptome for Chromosomes 21 and 22 by using high-density oligonucleotide arrays on cytosolic poly(A)(+) RNA. Conservatively, only 31.4% of the observed transcribed nucleotides correspond to well-annotated genes, whereas an additional 4.8% and 14.7% correspond to mRNAs and ESTs, respectively. Approximately 85% of the known exons were detected, and up to 21% of known genes have only a single isoform based on exon-skipping alternative expression. Overall, the expression of the well-characterized exons falls predominately into two categories, uniquely or ubiquitously expressed with an identifiable proportion of antisense transcripts. The remaining observed transcription (49.0%) was outside of any known annotation. These novel transcripts appear to be more cell-line-specific and have lower and less variation in expression than the well-characterized genes. Novel transcripts were further characterized based on their distance to annotations, transcript size, coding capacity, and identification as antisense to intronic sequences. By RT-PCR, 126 novel transcripts were independently verified, resulting in a 65% verification rate. These observations strongly support the argument for a re-evaluation of the total number of human genes and an alternative term for "gene" to encompass these growing, novel classes of RNA transcripts in the human genome.

Cell Line↗

Circular RNA profiling reveals an abundant circLMO7 that regulates myoblasts differentiation and survival by sponging miR-378a-3p.

Circular RNAs (circRNAs) have been identified from various tissues and species, but their regulatory functions during developmental processes are not well understood. We examined circRNA expression profiles of two developmental stages of bovine skeletal muscle (embryonic and adult musculus longissimus) to provide first insights into their potential involvement in bovine myogenesis. We identified 12 981 circRNAs and annotated them to the Bos taurus reference genome, including 530 circular intronic RNAs (ciRNAs). One parental gene could generate multiple circRNA isoforms, with only one or two isoforms being expressed at higher expression levels. Also, several host genes produced different isoforms when comparing development stages. Most circRNA candidates contained two to seven exons, and genomic distances to back-splicing sites were usually less than 50 kb. The length of upstream or downstream flanking introns was usually less than 105 nt (mean≈11 000 nt). Several circRNAs differed in abundance between developmental stages, and real-time quantitative PCR (qPCR) analysis largely confirmed differential expression of the 17 circRNAs included in this analysis. The second part of our study characterized the role of circLMO7-one of the most down-regulated circRNAs when comparing adult to embryonic muscle tissue-in bovine muscle development. Overexpression of circLMO7 inhibited the differentiation of primary bovine myoblasts, and it appears to function as a competing endogenous RNA for miR-378a-3p, whose involvement in bovine muscle development has been characterized beforehand. Congruent with our interpretation, circLMO7 increased the number of myoblasts in the S-phase of the cell cycle and decreased the proportion of cells in the G0/G1 phase. Moreover, it promoted the proliferation of myoblasts and protected them from apoptosis. Our study provides novel insights into the regulatory mechanisms underlying skeletal muscle development and identifies a number of circRNAs whose regulatory potential will need to be explored in the future.

Animals↗

Efficient evidence-based genome annotation with EviAnn.

For many years, machine learning-based ab initio gene finding approaches have been central components of eukaryotic genome annotation pipelines, and they remain so today. The reliance on these approaches was originally sustained by the high cost and low availability of gene expression data, a primary source of evidence for gene annotation along with protein homology. However, innovations in modern sequencing technologies have revolutionized the acquisition of gene expression data, allowing scientists to rely more heavily on this class of evidence. In addition, proteins found in a multitude of well-annotated genomes represent another invaluable resource for gene annotation. Existing annotation packages often underutilize these data sources, which prompted us to develop EviAnn (Evidence-based Annotator), a novel evidence-based eukaryotic gene annotation system. EviAnn takes a strongly data-driven approach, building the exon-intron structure of genes from transcript alignments or protein-sequence homology rather than from purely ab initio gene finding techniques. We show that when provided with the same input data, EviAnn consistently outperforms current state-of-the-art packages including BRAKER3, MAKER2, and FINDER, while utilizing considerably less computer time. Annotation of a mammalian genome can be completed in less than an hour on a single multi-core server. EviAnn is freely available under an open-source license from https://github.com/alekseyzimin/EviAnn_release and from Bioconda as "eviann".

Journal Article↗

Prediction of many new exons and introns in Plasmodium falciparum chromosome 2.

The current prediction of genes in the Plasmodium falciparum genome database relies upon a limited number of specially developed computer algorithms. We have re-annotated the sequence of chromosome 2 of P. falciparum by a computer-assisted manual analysis, which is described here. Of 161 newly predicted introns, we have experimentally confirmed 98. We regard 110 introns from the previously published analyses as probable, we delete 3, change 26 and add 135. We recognise 214 genes in chromosome 2. We have predicted introns in 121 genes. The increased complexity of gene structure on chromosome 2 is likely to be mirrored by the entire genome.

Algorithms↗

The complete set of tRNA species in Nanoarchaeum equitans.

The archaeal parasite Nanoarchaeum equitans was found to generate five tRNA species via a unique process requiring the assembly of seperate 5' and 3' tRNA halves [Randau, L., Munch, R., Hohn, M.J., Jahn, D. and Soll, D. (2005) Nanoarchaeum equitans creates functional tRNAs from separate genes for their 5'- and 3'-halves. Nature 433, 537-541]. Biochemical evidence was missing for one of the computationally-predicted, joined tRNAs designated as tRNA(Trp). Our RT-PCR and sequencing results identify this tRNA as tRNA(Lys) (CUU) joined at the alternative position between bases 30 and 31. We show that the intron-containing tRNA(Trp) was misidentified in the initial Nanoarchaeum equitans genome annotation [E. Waters et al. (2003) The genome of Nanoarchaeum equitans: insights into early archaeal evolution and derived parasitism. Proc. Natl. Acad. Sci. USA 100, 12984-12988]. Along with a previously unidentified joined tRNA(Gln) (UUG), Nanoarchaeum equitans exhibits 44 tRNAs and is enabled to read all 61 sense codons. Features unique to this set of tRNA molecules are discussed.

Base Sequence↗

Pangenome-wide identification and expression analysis of the chalcone synthase (CHS) gene family in five yellowhorn spp.

Chalcone synthase (CHS) is a pivotal enzyme in flavonoid biosynthesis involved in plant development, defense, and secondary metabolism. Xanthoceras sorbifolium (yellowhorn) is a medicinal and ornamental species with high resistance to environmental stresses, but its CHS gene family remains uncharacterized. We performed a pangenome-wide identification of CHS genes across five yellowhorn genomes (Xzs4, Xwf8, Xjg, Xg11, and Xzg2). Across the five yellowhorn genomes, 27 CHS genes were identified and classified into four core pangenes, present in all five genomes, and two dispensable genes, present only in a subset of genomes. Phylogenetic analysis grouped these genes into three major clades, and chromosomal mapping and duplication analyses identified four tandemly duplicated gene pairs under purifying selection. The analyses of conserved structural features, including protein motifs and exon-intron organization, together with promoter cis-regulatory elements and gene ontology annotation, further indicated the potential involvement of CHS genes in flavonoid biosynthesis and stress-responsive mechanisms. Gene expression profiling identified significant upregulation of Xg11_CHS1 and Xg11_CHS3 under cold and drought stress, with tissue-specific expression patterns. These findings provide valuable insights into the evolution, functional diversification, and stress-responsive roles of the CHS gene family, identifying candidate genes for future studies targeting stress tolerance and flavonoid biosynthesis in yellowhorn.

Acyltransferases↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

High-resolution multipoint linkage-disequilibrium mapping in the context of a human genome sequence.

A new method is presented for fine-scale linkage disequilibrium (LD) mapping of a disease mutation; it uses multiple linked single-nucleotide polymorphisms, restriction-fragment-length polymorphisms, or microsatellite markers and incorporates information from an annotated human genome sequence (HGS) and from a human mutation database. The method takes account of population demographic effects, using Markov chain Monte Carlo methods to integrate over the unknown gene genealogy and gene coalescence times. Information about the relative frequency of disease mutations in exons, introns, and other regions, from mutational databases, as well as assumptions about the completeness of the gene annotation, are used with an annotated HGS, to generate a prior probability that a mutation lies at any particular position in a specified region of the genome. This information is updated with information about mutation location, from LD at a set of linked markers in the region, to generate the posterior probability density of the mutation location. The performance of the method is evaluated by simulation and by analysis of a data set for diastrophic dysplasia (DTD) in Finland. The DTD disease gene has been positionally cloned, so the actual location of the mutation is known and can be compared with the position predicted by our method. For the DTD data, the addition of information from an HGS results in disease-gene localization at a resolution that is much higher than that which would be possible by LD mapping alone. In this case, the gene would be found by sequencing a region < or =7 kb in size.

Algorithms↗

Experimental RNomics: identification of 140 candidates for small non-messenger RNAs in the plant Arabidopsis thaliana.

BACKGROUND: Genomes from all organisms known to date express two types of RNA molecules: messenger RNAs (mRNAs), which are translated into proteins, and non-messenger RNAs, which function at the RNA level and do not serve as templates for translation. RESULTS: We have generated a specialized cDNA library from Arabidopsis thaliana to investigate the population of small non-messenger RNAs (snmRNAs) sized 50-500 nt in a plant. From this library, we identified 140 candidates for novel snmRNAs and investigated their expression, abundance, and developmental regulation. Based on conserved sequence and structure motifs, 104 snmRNA species can be assigned to novel members of known classes of RNAs (designated Class I snmRNAs), namely, small nucleolar RNAs (snoRNAs), 7SL RNA, U snRNAs, as well as a tRNA-like RNA. For the first time, 39 novel members of H/ACA box snoRNAs could be identified in a plant species. Of the remaining 36 snmRNA candidates (designated Class II snmRNAs), no sequence or structure motifs were present that would enable an assignment to a known class of RNAs. These RNAs were classified based on their location on the A. thaliana genome. From these, 29 snmRNA species located to intergenic regions, 3 located to intronic sequences of protein coding genes, and 4 snmRNA candidates were derived from annotated open reading frames. Surprisingly, 15 of the Class II snmRNA candidates were shown to be tissue-specifically expressed, while 12 are encoded by the mitochondrial or chloroplast genome. CONCLUSIONS: Our study has identified 140 novel candidates for small non-messenger RNA species in the plant A. thaliana and thereby sets the stage for their functional analysis.

Arabidopsis↗

Non-EST based prediction of exon skipping and intron retention events using Pfam information.

Most of the known alternative splice events have been detected by the comparison of expressed sequence tags (ESTs) and cDNAs. However, not all splice events are represented in EST databases since ESTs have several biases. Therefore, non-EST based approaches are needed to extend our view of a transcriptome. Here, we describe a novel method for the ab initio prediction of alternative splice events that is solely based on the annotation of Pfam domains. Furthermore, we applied this approach in a genome-wide manner to all human RefSeq transcripts and predicted a total of 321 exon skipping and intron retention events. We show that this method is very reliable as 78% (250 of 321) of our predictions are confirmed by ESTs or cDNAs. Subsequent analyses of splice events within Pfam domains revealed a significant preference of alternative exon junctions to be located at the protein surface and to avoid secondary structure elements. Thus, splice events within Pfams are probable to alter the structure and function of a domain which makes them highly interesting for detailed biological investigation. As Pfam domains are annotated in many other species, our strategy to predict exon skipping and intron retention events might be important for species with a lower number of ESTs.

Algorithms↗

Surrogate splicing for functional analysis of sesquiterpene synthase genes.

A method for the recovery of full-length cDNAs from predicted terpene synthase genes containing introns is described. The approach utilizes Agrobacterium-mediated transient expression coupled with a reverse transcription-polydeoxyribonucleotide chain reaction assay to facilitate expression cloning of processed transcripts. Subsequent expression of intronless cDNAs in a suitable prokaryotic host provides for direct functional testing of the encoded gene product. The method was optimized by examining the expression of an intron-containing beta-glucuronidase gene agroinfiltrated into petunia (Petunia hybrida) leaves, and its utility was demonstrated by defining the function of two previously uncharacterized terpene synthases. A tobacco (Nicotiana tabacum) terpene synthase-like gene containing six predicted introns was characterized as having 5-epi-aristolochene synthase activity, while an Arabidopsis (Arabidopsis thaliana) gene previously annotated as a terpene synthase was shown to possess a novel sesquiterpene synthase activity for alpha-barbatene, thujopsene, and beta-chamigrene biosynthesis.

Alternative Splicing↗

A comparative genomic analysis of the cow, pig, and human CFTR genes identifies potential intronic regulatory elements.

The identification of sequences within noncoding regions of genes that are conserved between several species may indicate potential regulatory elements. This is important for genes with complex control mechanisms such as the cystic fibrosis transmembrane conductance regulator (CFTR). CFTR demonstrates similar patterns of temporal and spatial expression in human and sheep, but these differ significantly in mouse cftr. The complete sheep CFTR sequence is unavailable so we annotated BAC clones encompassing the CFTR gene from two other artiodactyl species (cow and pig) for comparative sequence analysis. Regions of introns 2, 3, 10, 17a, 18, and 21 and 3' flanking sequence corresponding to human CFTR DNase I hypersensitive sites (DHS) showed high homology in the cow and pig. Cross-species sequence conservation also enabled finer mapping of other human DHS, including those in introns 1, 16, and 20. Additional potential regulatory elements not associated with human DHS were also identified.

Animals↗