Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Mass spectrometric genomic data mining: Novel insights into bioenergetic pathways in Chlamydomonas reinhardtii.

A new high-throughput computational strategy was established that improves genomic data mining from MS experiments. The MS/MS data were analyzed by the SEQUEST search algorithm and a combination of de novo amino acid sequencing in conjunction with an error-tolerant database search tool, operating on a 256 processor computer cluster. The error-tolerant search tool, previously established as GenomicPeptideFinder (GPF), enables detection of intron-split and/or alternatively spliced peptides from MS/MS data when deduced from genomic DNA. Isolated thylakoid membranes from the eukaryotic green alga Chlamydomonas reinhardtii were separated by 1-D SDS gel electrophoresis, protein bands were excised from the gel, digested in-gel with trypsin and analyzed by coupling nano-flow LC with MS/MS. The concerted action of SEQUEST and GPF allowed identification of 2622 distinct peptides. In total 448 peptides were identified by GPF analysis alone, including 98 intron-split peptides, resulting in the identification of novel proteins, improved annotation of gene models, and evidence of alternative splicing.

Algorithms↗

Circular RNA profiling reveals an abundant circLMO7 that regulates myoblasts differentiation and survival by sponging miR-378a-3p.

Circular RNAs (circRNAs) have been identified from various tissues and species, but their regulatory functions during developmental processes are not well understood. We examined circRNA expression profiles of two developmental stages of bovine skeletal muscle (embryonic and adult musculus longissimus) to provide first insights into their potential involvement in bovine myogenesis. We identified 12 981 circRNAs and annotated them to the Bos taurus reference genome, including 530 circular intronic RNAs (ciRNAs). One parental gene could generate multiple circRNA isoforms, with only one or two isoforms being expressed at higher expression levels. Also, several host genes produced different isoforms when comparing development stages. Most circRNA candidates contained two to seven exons, and genomic distances to back-splicing sites were usually less than 50 kb. The length of upstream or downstream flanking introns was usually less than 105 nt (mean≈11 000 nt). Several circRNAs differed in abundance between developmental stages, and real-time quantitative PCR (qPCR) analysis largely confirmed differential expression of the 17 circRNAs included in this analysis. The second part of our study characterized the role of circLMO7-one of the most down-regulated circRNAs when comparing adult to embryonic muscle tissue-in bovine muscle development. Overexpression of circLMO7 inhibited the differentiation of primary bovine myoblasts, and it appears to function as a competing endogenous RNA for miR-378a-3p, whose involvement in bovine muscle development has been characterized beforehand. Congruent with our interpretation, circLMO7 increased the number of myoblasts in the S-phase of the cell cycle and decreased the proportion of cells in the G0/G1 phase. Moreover, it promoted the proliferation of myoblasts and protected them from apoptosis. Our study provides novel insights into the regulatory mechanisms underlying skeletal muscle development and identifies a number of circRNAs whose regulatory potential will need to be explored in the future.

Animals↗

Efficient evidence-based genome annotation with EviAnn.

For many years, machine learning-based ab initio gene finding approaches have been central components of eukaryotic genome annotation pipelines, and they remain so today. The reliance on these approaches was originally sustained by the high cost and low availability of gene expression data, a primary source of evidence for gene annotation along with protein homology. However, innovations in modern sequencing technologies have revolutionized the acquisition of gene expression data, allowing scientists to rely more heavily on this class of evidence. In addition, proteins found in a multitude of well-annotated genomes represent another invaluable resource for gene annotation. Existing annotation packages often underutilize these data sources, which prompted us to develop EviAnn (Evidence-based Annotator), a novel evidence-based eukaryotic gene annotation system. EviAnn takes a strongly data-driven approach, building the exon-intron structure of genes from transcript alignments or protein-sequence homology rather than from purely ab initio gene finding techniques. We show that when provided with the same input data, EviAnn consistently outperforms current state-of-the-art packages including BRAKER3, MAKER2, and FINDER, while utilizing considerably less computer time. Annotation of a mammalian genome can be completed in less than an hour on a single multi-core server. EviAnn is freely available under an open-source license from https://github.com/alekseyzimin/EviAnn_release and from Bioconda as "eviann".

Journal Article↗

Prediction of many new exons and introns in Plasmodium falciparum chromosome 2.

The current prediction of genes in the Plasmodium falciparum genome database relies upon a limited number of specially developed computer algorithms. We have re-annotated the sequence of chromosome 2 of P. falciparum by a computer-assisted manual analysis, which is described here. Of 161 newly predicted introns, we have experimentally confirmed 98. We regard 110 introns from the previously published analyses as probable, we delete 3, change 26 and add 135. We recognise 214 genes in chromosome 2. We have predicted introns in 121 genes. The increased complexity of gene structure on chromosome 2 is likely to be mirrored by the entire genome.

Algorithms↗

The complete set of tRNA species in Nanoarchaeum equitans.

The archaeal parasite Nanoarchaeum equitans was found to generate five tRNA species via a unique process requiring the assembly of seperate 5' and 3' tRNA halves [Randau, L., Munch, R., Hohn, M.J., Jahn, D. and Soll, D. (2005) Nanoarchaeum equitans creates functional tRNAs from separate genes for their 5'- and 3'-halves. Nature 433, 537-541]. Biochemical evidence was missing for one of the computationally-predicted, joined tRNAs designated as tRNA(Trp). Our RT-PCR and sequencing results identify this tRNA as tRNA(Lys) (CUU) joined at the alternative position between bases 30 and 31. We show that the intron-containing tRNA(Trp) was misidentified in the initial Nanoarchaeum equitans genome annotation [E. Waters et al. (2003) The genome of Nanoarchaeum equitans: insights into early archaeal evolution and derived parasitism. Proc. Natl. Acad. Sci. USA 100, 12984-12988]. Along with a previously unidentified joined tRNA(Gln) (UUG), Nanoarchaeum equitans exhibits 44 tRNAs and is enabled to read all 61 sense codons. Features unique to this set of tRNA molecules are discussed.

Base Sequence↗

Pangenome-wide identification and expression analysis of the chalcone synthase (CHS) gene family in five yellowhorn spp.

Chalcone synthase (CHS) is a pivotal enzyme in flavonoid biosynthesis involved in plant development, defense, and secondary metabolism. Xanthoceras sorbifolium (yellowhorn) is a medicinal and ornamental species with high resistance to environmental stresses, but its CHS gene family remains uncharacterized. We performed a pangenome-wide identification of CHS genes across five yellowhorn genomes (Xzs4, Xwf8, Xjg, Xg11, and Xzg2). Across the five yellowhorn genomes, 27 CHS genes were identified and classified into four core pangenes, present in all five genomes, and two dispensable genes, present only in a subset of genomes. Phylogenetic analysis grouped these genes into three major clades, and chromosomal mapping and duplication analyses identified four tandemly duplicated gene pairs under purifying selection. The analyses of conserved structural features, including protein motifs and exon-intron organization, together with promoter cis-regulatory elements and gene ontology annotation, further indicated the potential involvement of CHS genes in flavonoid biosynthesis and stress-responsive mechanisms. Gene expression profiling identified significant upregulation of Xg11_CHS1 and Xg11_CHS3 under cold and drought stress, with tissue-specific expression patterns. These findings provide valuable insights into the evolution, functional diversification, and stress-responsive roles of the CHS gene family, identifying candidate genes for future studies targeting stress tolerance and flavonoid biosynthesis in yellowhorn.

Acyltransferases↗

InSatDb: a microsatellite database of fully sequenced insect genomes.

InSatDb presents an interactive interface to query information regarding microsatellite characteristics per se of five fully sequenced insect genomes (fruit-fly, honeybee, malarial mosquito, red-flour beetle and silkworm). InSatDb allows users to obtain microsatellites annotated with size (in base pairs and repeat units); genomic location (exon, intron, up-stream or transposon); nature (perfect or imperfect); and sequence composition (repeat motif and GC%). One can access microsatellite cluster (compound repeats) information and a list of microsatellites with conserved flanking sequences (microsatellite family or paralogs). InSatDb is complete with the insects information, web links to find details, methodology and a tutorial. A separate 'Analysis' section illustrates the comparative genomic analysis that can be carried out using the output. InSatDb is available at www.cdfd.org.in/insatdb.

Animals↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

High-resolution multipoint linkage-disequilibrium mapping in the context of a human genome sequence.

A new method is presented for fine-scale linkage disequilibrium (LD) mapping of a disease mutation; it uses multiple linked single-nucleotide polymorphisms, restriction-fragment-length polymorphisms, or microsatellite markers and incorporates information from an annotated human genome sequence (HGS) and from a human mutation database. The method takes account of population demographic effects, using Markov chain Monte Carlo methods to integrate over the unknown gene genealogy and gene coalescence times. Information about the relative frequency of disease mutations in exons, introns, and other regions, from mutational databases, as well as assumptions about the completeness of the gene annotation, are used with an annotated HGS, to generate a prior probability that a mutation lies at any particular position in a specified region of the genome. This information is updated with information about mutation location, from LD at a set of linked markers in the region, to generate the posterior probability density of the mutation location. The performance of the method is evaluated by simulation and by analysis of a data set for diastrophic dysplasia (DTD) in Finland. The DTD disease gene has been positionally cloned, so the actual location of the mutation is known and can be compared with the position predicted by our method. For the DTD data, the addition of information from an HGS results in disease-gene localization at a resolution that is much higher than that which would be possible by LD mapping alone. In this case, the gene would be found by sequencing a region < or =7 kb in size.

Algorithms↗

Experimental RNomics: identification of 140 candidates for small non-messenger RNAs in the plant Arabidopsis thaliana.

BACKGROUND: Genomes from all organisms known to date express two types of RNA molecules: messenger RNAs (mRNAs), which are translated into proteins, and non-messenger RNAs, which function at the RNA level and do not serve as templates for translation. RESULTS: We have generated a specialized cDNA library from Arabidopsis thaliana to investigate the population of small non-messenger RNAs (snmRNAs) sized 50-500 nt in a plant. From this library, we identified 140 candidates for novel snmRNAs and investigated their expression, abundance, and developmental regulation. Based on conserved sequence and structure motifs, 104 snmRNA species can be assigned to novel members of known classes of RNAs (designated Class I snmRNAs), namely, small nucleolar RNAs (snoRNAs), 7SL RNA, U snRNAs, as well as a tRNA-like RNA. For the first time, 39 novel members of H/ACA box snoRNAs could be identified in a plant species. Of the remaining 36 snmRNA candidates (designated Class II snmRNAs), no sequence or structure motifs were present that would enable an assignment to a known class of RNAs. These RNAs were classified based on their location on the A. thaliana genome. From these, 29 snmRNA species located to intergenic regions, 3 located to intronic sequences of protein coding genes, and 4 snmRNA candidates were derived from annotated open reading frames. Surprisingly, 15 of the Class II snmRNA candidates were shown to be tissue-specifically expressed, while 12 are encoded by the mitochondrial or chloroplast genome. CONCLUSIONS: Our study has identified 140 novel candidates for small non-messenger RNA species in the plant A. thaliana and thereby sets the stage for their functional analysis.

Arabidopsis↗

Non-EST based prediction of exon skipping and intron retention events using Pfam information.

Most of the known alternative splice events have been detected by the comparison of expressed sequence tags (ESTs) and cDNAs. However, not all splice events are represented in EST databases since ESTs have several biases. Therefore, non-EST based approaches are needed to extend our view of a transcriptome. Here, we describe a novel method for the ab initio prediction of alternative splice events that is solely based on the annotation of Pfam domains. Furthermore, we applied this approach in a genome-wide manner to all human RefSeq transcripts and predicted a total of 321 exon skipping and intron retention events. We show that this method is very reliable as 78% (250 of 321) of our predictions are confirmed by ESTs or cDNAs. Subsequent analyses of splice events within Pfam domains revealed a significant preference of alternative exon junctions to be located at the protein surface and to avoid secondary structure elements. Thus, splice events within Pfams are probable to alter the structure and function of a domain which makes them highly interesting for detailed biological investigation. As Pfam domains are annotated in many other species, our strategy to predict exon skipping and intron retention events might be important for species with a lower number of ESTs.

Algorithms↗

Oligonucleotide frequencies in DNA follow a Yule distribution.

We show that ranked oligonucleotide frequencies in both protein-coding and non-coding regions from several genomes fit poorly to the Zipf distribution, but that the same frequency data give excellent fit to the Yule distribution. The parameters of the Yule distribution for oligonucleotide frequencies in exons are the same (within error limits) as the parameters for introns. This precludes application of Yule or Zipf distribution of ranked oligonucleotide frequencies to annotating new genomic sequences.

DNA↗

Identification and analysis of gene families from the duplicated genome of soybean using EST sequences.

BACKGROUND: Large scale gene analysis of most organisms is hampered by incomplete genomic sequences. In many organisms, such as soybean, the best source of sequence information is the existence of expressed sequence tag (EST) libraries. Soybean has a large (1115 Mbp) genome that has yet to be fully sequenced. However it does have the 6th largest EST collection comprised of ESTs from a variety of soybean genotypes. Many EST libraries were constructed from RNA extracted from various genetic backgrounds, thus gene identification from these sources is complicated by the existence of both gene and allele sequence differences. We used the ESTminer suite of programs to identify potential soybean gene transcripts from a single genetic background allowing us to observe functional classifications between gene families as well as structural differences between genes and gene paralogs within families. The identification of potential gene sequences (pHaps) from soybean allows us to begin to get a picture of the genomic history of the organism as well as begin to observe the evolutionary fates of gene copies in this highly duplicated genome. RESULTS: We identified approximately 45,000 potential gene sequences (pHaps) from EST sequences of Williams/Williams82, an inbred genotype of soybean (Glycine max L. Merr.) using a redundancy criterion to identify reproducible sequence differences between related genes within gene families. Analysis of these sequences revealed single base substitutions and single base indels are the most frequently observed form of sequence variation between genes within families in the dataset. Genomic sequencing of selected loci indicate that intron-like intervening sequences are numerous and are approximately 220 bp in length. Functional annotation of gene sequences indicate functional classifications are not randomly distributed among gene families containing few or many genes. CONCLUSION: The predominance of single nucleotide insertion/deletions and substitution events between genes within families (individual genes and gene paralogs) is consistent with a model of gene amplification followed by single base random mutational events expected under the classical model of duplicated gene evolution. Molecular functions of small and large gene families appear to be non-randomly distributed possibly indicating a difference in retention of duplicates or local expansion.

Evolution, Molecular↗

An intronic insertion in KPL2 results in aberrant splicing and causes the immotile short-tail sperm defect in the pig.

The immotile short-tail sperm defect is an autosomal recessive disease within the Finnish Yorkshire pig population. This disease specifically affects the axoneme structure of sperm flagella, whereas cilia in other tissues appear unaffected. Recently, the disease locus was mapped to a 3-cM region on porcine chromosome 16. To facilitate identification of candidate genes, we constructed a porcine-human comparative map, which anchored the disease locus to a region on human chromosome 5p13.2 containing eight annotated genes. Sequence analysis of a candidate gene KPL2 revealed the presence of an inserted retrotransposon within an intron. The insertion affects splicing of the KPL2 transcript in two ways; it either causes skipping of the upstream exon, or causes the inclusion of an intronic sequence as well as part of the insertion in the transcript. Both changes alter the reading frame leading to premature termination of translation. Further work revealed that the aberrantly spliced exon is expressed predominantly in testicular tissue, which explains the tissue-specificity of the immotile short-tail sperm defect. These findings show that the KPL2 gene is important for correct axoneme development and provide insight into abnormal sperm development and infertility disorders.

Alternative Splicing↗

Surrogate splicing for functional analysis of sesquiterpene synthase genes.

A method for the recovery of full-length cDNAs from predicted terpene synthase genes containing introns is described. The approach utilizes Agrobacterium-mediated transient expression coupled with a reverse transcription-polydeoxyribonucleotide chain reaction assay to facilitate expression cloning of processed transcripts. Subsequent expression of intronless cDNAs in a suitable prokaryotic host provides for direct functional testing of the encoded gene product. The method was optimized by examining the expression of an intron-containing beta-glucuronidase gene agroinfiltrated into petunia (Petunia hybrida) leaves, and its utility was demonstrated by defining the function of two previously uncharacterized terpene synthases. A tobacco (Nicotiana tabacum) terpene synthase-like gene containing six predicted introns was characterized as having 5-epi-aristolochene synthase activity, while an Arabidopsis (Arabidopsis thaliana) gene previously annotated as a terpene synthase was shown to possess a novel sesquiterpene synthase activity for alpha-barbatene, thujopsene, and beta-chamigrene biosynthesis.

Alternative Splicing↗

A comparative genomic analysis of the cow, pig, and human CFTR genes identifies potential intronic regulatory elements.

The identification of sequences within noncoding regions of genes that are conserved between several species may indicate potential regulatory elements. This is important for genes with complex control mechanisms such as the cystic fibrosis transmembrane conductance regulator (CFTR). CFTR demonstrates similar patterns of temporal and spatial expression in human and sheep, but these differ significantly in mouse cftr. The complete sheep CFTR sequence is unavailable so we annotated BAC clones encompassing the CFTR gene from two other artiodactyl species (cow and pig) for comparative sequence analysis. Regions of introns 2, 3, 10, 17a, 18, and 21 and 3' flanking sequence corresponding to human CFTR DNase I hypersensitive sites (DHS) showed high homology in the cow and pig. Cross-species sequence conservation also enabled finer mapping of other human DHS, including those in introns 1, 16, and 20. Additional potential regulatory elements not associated with human DHS were also identified.

Animals↗

The WRKY family of transcription factors in rice and Arabidopsis and their origins.

WRKY transcription factors, originally isolated from plants contain one or two conserved WRKY domains, about 60 amino acid residues with the WRKYGQK sequence followed by a C2H2 or C2HC zinc finger motif. Evidence is accumulating to suggest that the WRKY proteins play significant roles in responses to biotic and abiotic stresses, and in development. In this research, we identified 102 putative WRKY genes from the rice genome and compared them with those from Arabidopsis. The WRKY genes from rice and Arabidopsis were divided into three groups with several subgroups on the basis of phylogenies and the basic structure of the WRKY domains (WDs). The phylogenetic trees generated from the WDs and the genes indicate that the WRKY gene family arose during evolution through duplication and that the dramatic amplification of rice WRKY genes in group III is due to tandem and segmental gene duplication compared with those of Arabidopsis. The result suggests that some of the rice WRKY genes in group III are evolutionarily more active than those in Arabidopsis, and may have specific roles in monocotyledonous plants. Further, it was possible to identify the presence of WRKY-like genes in protists (Giardia lamblia and Dictyostelium discoideum) and green algae Chlamydomonas reinhardtii through database research, demonstrating the ancient origin of the gene family. The results obtained by alignments of the WDs from different species and other analysis imply that domain gain and loss is a divergent force for expansion of the WRKY gene family, and that a rapid amplification of the WRKY genes predate the divergence of monocots and dicots. On the basis of these results, we believe that genes encoding a single WD may have been derived from the C-terminal WD of the genes harboring two WDs. The conserved intron splicing positions in the WDs of higher plants offer clues about WRKY gene evolution, annotation, and classification.

Amino Acid Motifs↗

Evolutionary implications of three novel members of the human sarcomeric myosin heavy chain gene family.

Sarcomeric myosin heavy chain (MyHC) is the major contractile protein of striated muscle. Six tandemly linked skeletal MyHC genes on chromosome 17 and two cardiac MyHC genes on chromosome 14 have been previously described in the human genome. We report the identification of three novel human sarcomeric MyHC genes on chromosomes 3, 7, and 20, which are notable for their atypical size and intron-exon structure. Two of the encoded proteins are structurally most like the slow-beta MyHC, whereas the third one is closest to the adult fast IIb isoform. Data from pairwise comparisons of aligned coding sequences imply the existence of ancestral genomes with four sarcomeric genes before the emergence of a dedicated smooth muscle MyHC gene. To further address the evolutionary relationships of the distinct sarcomeric and nonsarcomeric rod sequences, we have identified and further annotated human genomic DNA sequences corresponding to 14 class-II MyHCs. An extensive analysis provides a timeline for intron gain and loss, gene contraction and expansion, and gene conversion among genes encoding class-II myosins. One of the novel human genes is found to have introns at positions shared only with the molluscan catchin/MyHC gene, providing evidence for the structure of a pre-Cambrian ancestral gene.

Amino Acid Sequence↗