Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Dark matter in the genome: evidence of widespread transcription detected by microarray tiling experiments.

Microarrays provide the opportunity to measure transcription from regions of the genome without bias towards the location of known genes. This technology thus offers an important source of genomic sequence annotation that is complementary to cDNA sequencing and computational gene-finding methods. Recent "tiling" microarray experiments that assay transcription at regular intervals throughout the genome have shown evidence of large amounts of transcription outside the boundaries of known genes. This transcription is observed in polyadenylated RNA samples and appears to be derived from intergenic regions, from introns of known genes and from sequences antisense to known transcripts. In this article, we discuss different explanations for this phenomenon.

Animals↗

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays.

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis↗

Mitochondrial genomic characteristics and phylogenetic analysis of Cunninghamella elegans (Mucorales: Cunninghamellaceae).

Cunninghamella, a filamentous fungal genus with important biomedical and biochemical value, lacks any fully annotated mitochondrial genome to date. Herein, we presented the first complete mitogenome of Cunninghamella elegans, a circular 41,552 bp molecule (GC 27.86%) encoding 14 conserved protein-coding genes, 2 rRNA genes, 24 tRNA genes, and 6 non-conserved ORFs. Structural comparison with related species (Absidia glauca and Gongronella sp. w5) revealed dynamic evolution in intron and repeat elements. Phylogenetics places C. elegans within Cunninghamellaceae, with Gongronella as its closest relative. This reference mitogenome will underpin future evolutionary and taxonomic investigations of this industrially and medically significant lineage.

Cunninghamella elegans↗

Dynamite: a flexible code generating language for dynamic programming methods used in sequence comparison.

We have developed a code generating language, called Dynamite, specialised for the production and subsequent manipulation of complex dynamic programming methods for biological sequence comparison. From a relatively simple text definition file Dynamite will produce a variety of implementations of a dynamic programming method, including database searches and linear space alignments. The speed of the generated code is comparable to hand written code, and the additional flexibility has proved invaluable in designing and testing new algorithms. An innovation is a flexible labelling system, which can be used to annotate the original sequences with biological information. We illustrate the Dynamite syntax and flexibility by showing definitions for dynamic programming routines (i) to align two protein sequences under the assumption that they are both poly-topic transmembrane proteins, with the simultaneous assignment of transmembrane helices and (ii) to align protein information to genomic DNA, allowing for introns and sequencing error.

Algorithms↗

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525 bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2 kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial↗

Construction of an approximately 700-kb transcript map around the familial Mediterranean fever locus on human chromosome 16p13.3.

We used a combination of cDNA selection, exon amplification, and computational prediction from genomic sequence to isolate transcribed sequences from genomic DNA surrounding the familial Mediterranean fever (FMF) locus. Eighty-seven kb of genomic DNA around D16S3370, a marker showing a high degree of linkage disequilibrium with FMF, was sequenced to completion, and the sequence annotated. A transcript map reflecting the minimal number of genes encoded within the approximately 700 kb of genomic DNA surrounding the FMF locus was assembled. This map consists of 27 genes with discreet messages detectable on Northerns, in addition to three olfactory-receptor genes, a cluster of 18 tRNA genes, and two putative transcriptional units that have typical intron-exon splice junctions yet do not detect messages on Northerns. Four of the transcripts are identical to genes described previously, seven have been independently identified by the French FMF Consortium, and the others are novel. Six related zinc-finger genes, a cluster of tRNAs, and three olfactory receptors account for the majority of transcribed sequences isolated from a 315-kb FMF central region (between D16S468/D16S3070 and cosmid 377A12). Interspersed among them are several genes that may be important in inflammation. This transcript map not only has permitted the identification of the FMF gene (MEFV), but also has provided us an opportunity to probe the structural and functional features of this region of chromosome 16.

Amino Acid Sequence↗

u-Genome: a database on genome design in unicellular genomes.

Unicellular eukaryotes were among the first ones to be selected for complete genome sequencing because of the small size of their genomes and their interactions with humans and a broad range of animals and plants. Currently, ten completely sequenced unicellular genome sequences have been publicly released and as the number of available unicellular genomes increases, comparative genomics analysis within this group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named u-Genome, the unicellular genome design database. The database is unique in featuring three datasets namely (1) orthologous proteins (2) paralogous proteins and (3) statistical distributions on exons, introns, intergenic DNA and correlations between them. A tool, Uniview, designed to visualize the gene structures for individual genes in the genome is also integrated. This database is of importance in understanding unicellular genome design and architecture and evolution related studies. The database is available through a web interface at http://sege.ntu.edu.sg/wester/ugenome.

Animals↗

Analysis of EST-driven gene annotation in human genomic sequence.

We have performed a systematic analysis of gene identification in genomic sequence by similarity search against expressed sequence tags (ESTs) to assess the suitability of this method for automated annotation of the human genome. A BLAST-based strategy was constructed to examine the potential of this approach, and was applied to test sets containing all human genomic sequences longer than 5 kb in public databases, plus 300 kb of exhaustively characterized benchmark sequence. At high stringency, 70%-90% of all annotated genes are detected by near-identity to EST sequence; >95% of ESTs aligning with well-annotated sequences overlap a gene. These ESTs provide immediate access to the corresponding cDNA clones for follow-up laboratory verification and subsequent biologic analysis. At lower stringency, up to 97% of annotated genes were identified by similarity to ESTs. The apparent false-positive rate rose to 55% of ESTs among all sequences and 20% among benchmark sequences at the lowest stringency, indicating that many genes in public database entries are unannotated. Approximately half of the alignments span multiple exons, and thus aid in the construction of gene predictions and elucidation of alternative splicing. In addition, ESTs from multiple cDNA libraries frequently cluster over genes, providing a starting point for crude expression profiles. Clone IDs may be used to form EST pairs, and particularly to extend models by associating alignments of lower stringency with high-quality alignments. These results demonstrate that EST similarity search is a practical general-purpose annotation technique that complements pattern recognition methods as a tool for gene characterization.

Base Sequence↗

Annotation, nomenclature and evolution of four novel homeobox genes expressed in the human germ line.

The homeobox genes comprise a large gene superfamily characterised by a conserved DNA motif encoding the homeodomain. Most homeodomain proteins function as transcription factors, and many have important roles in embryonic development and cell differentiation. Here we describe, annotate and name four novel homeobox genes in the human genome: ARGFX, DPRX, TPRX1 and DUXA. Each has generated multiple retrotransposed (processed) pseudogenes; these are reliable indicators of germ-line expression because only in germ-line cells can retrotransposition result in inheritance to the next generation. The retrotransposed sequences were exploited here as a novel means to deduce exon-intron boundaries. All four novel genes show accelerated rates of protein sequence evolution. This fast rate of sequence change may be connected with roles in human reproductive biology. Deducing the evolutionary origins of these genes is not straightforward, but we propose that TPRX1, DPRX and DUXA are highly divergent derivatives of the CRX gene, itself a member of the Otx homeobox gene family.

Evolution, Molecular↗

Gene targeting of ErbB3 using a Cre-mediated unidirectional DNA inversion strategy.

Recombinase-mediated unidirectional DNA inversion and transcriptional arrest is a promising strategy for high throughput conditional mutagenesis in the mouse. Banks of mouse embryonic stem cells with defined, transcriptionally silent insertions that can be activated by Cre recombinase would take advantage of existing transgenic Cre lines to rapidly produce hundreds of lineage specific and temporally controlled knockout mice for each gene, thereby introducing significant parallelism to functional gene annotation. However, the extent to which this strategy results in effective gene knockout has not been established. To test the feasibility of this strategy we targeted ErbB3, a member of the ErbB family of tyrosine kinase receptors, using this strategy. Insertion of a reversed "flipflox" vector consisting of a gene inactivation cassette (GI) and an internal ribosome entry site (IRES)-GFP reporter into intron 1 of ErbB3 was transcriptionally silent and did not affect ErbB3 expression. Crosses with ubiquitous and lineage specific Cre recombinase expressing lines permanently inverted the inserted GI cassette and blocked ErbB3 expression. Unidirectional DNA inversion by in vivo recombination is an effective strategy for targeted or ubiquitous gene knockout.

Animals↗

Molecular cloning of a putative Ciona intestinalis cionin receptor, a new member of the CCK/gastrin receptor family.

Cionin, a peptide showing similarities with cholecystokinin and gastrin has been shown to be expressed in the gut and neural ganglion of the protochordate Ciona intestinalis. The present report describes the cloning of a putative cionin receptor (CioR), a new member of the CCK/gastrin family from the gastrointestinal tract of C. intestinalis. mRNA from the stomach of C. intestinalis was isolated using a modified RNA extraction procedure and, subsequently, reverse-transcribed into single-stranded cDNA by means of rapid amplification of 5'- and 3'-cDNA ends (RACE-PCR), followed by full-length PCR amplification. The cloned full-length PCR amplicons contained a short upstream open-reading frame (uORF) coding for a putative 16 amino acid long peptide, followed by a long open reading frame encoding a 526 amino acid putative CioR protein. At the amino acid level, the putative CioR protein shared 35-40% homology with cloned mammalian, chicken, and Xenopus laevis CCK receptors. Phylogenetic analysis revealed that the chicken and X. laevis CCK receptors are orthologues of the mammalian CCK2 receptors whereas CioR protein forms a clade with vertebrate cholecystokinin receptors. Moreover, we found that the CioR cDNA and deduced amino acid sequences were found to correspond to the annotated CCK/gastrin-like receptor gene on Scaffold 117 (C. intestinalis draft genome project, Joint Genome Institute database; http://www.jgi.doe.gov).

Amino Acid Sequence↗

Computational identification and systematic analysis of the ACR gene family in Oryza sativa.

Based on sequence similarity search and domain detection, nine ACT domain repeat protein-coding genes (the "ACR" genes) in rice were identified, which were mainly distributed on the chromosomes 2, 3, 4, and 8. An InterPro database search indicated that four copies of the ACT domain linearly occupied the entire polypeptide. The first three ACT domains were linked by two different sequences. However, the fourth ACT domain was extremely close to ACT3. Gene structure comparisons showed large differences in exon numbers, from three to eight, among members of the rice ACR gene family. In addition, it appeared that gene duplication might be operative when the compositions of exons and introns were analyzed. Phylogenetic analysis divided the ACR gene family into five distinct groups, and this division was generally according to the expression patterns of the ACR genes. The Arabidopsis and rice ACR proteins were clustered across together, suggesting that these ACR genes might originate from an ancient common ancestor. Notably, the identification of orthologues and paralogues would be useful for rice gene functional annotation.

Chromosome Mapping↗

A new strategy to identify novel genes and gene isoforms: Analysis of human chromosomes 15, 21 and 22.

We present here a novel methodology for the identification of genome regions potentially spanning one or more protein coding genes. It is based on the detection of clusters of conserved sequence tags whose evolutionary dynamics, based on the observation of an excess bias of synonymous substitutions at nucleotide level and of conservative replacements at protein level, suggests a likely protein coding role. A benchmark test carried out on a 236 Mbp of human-mouse syntenic regions from human chromosomes 15, 21 and 22 identified 25 CST clusters potentially containing unannotated genes. A further annotation update of the human genome assembly revealed that 11/25 clusters actually contained a total of 20 validated genes and 10 of the remaining 14 clusters had several experimental evidence in support of the presence of protein coding genes. These findings demonstrate the effectiveness and high prediction reliability of the proposed methodology which could specifically be applied to the annotation of novel genome sequences.

Animals↗

Molecular cloning and functional expression of a Drosophila receptor for the neuropeptides capa-1 and -2.

The Drosophila Genome Project website contains an annotated gene (CG14575) for a G protein-coupled receptor. We cloned this receptor and found that the cloned cDNA did not correspond to the annotated gene; it partly contained different exons and additional exons located at the 5(')-end of the annotated gene. We expressed the coding part of the cloned cDNA in Chinese hamster ovary cells and found that the receptor was activated by two neuropeptides, capa-1 and -2, encoded by the Drosophila capability gene. Database searches led to the identification of a similar receptor in the genome from the malaria mosquito Anopheles gambiae (58% amino acid residue identities; 76% conserved residues; and 5 introns at identical positions within the two insect genes). Because capa-1 and -2 and related insect neuropeptides stimulate fluid secretion in insect Malpighian (renal) tubules, the identification of this first insect capa receptor will advance our knowledge on insect renal function.

Amino Acid Sequence↗

The human gene CXorf17 encodes a member of a novel family of putative transmembrane proteins: cDNA cloning and characterization of CXorf17 and its mouse ortholog orf34.

We report the identification and cloning of a novel human gene, CXorf17, together with its mouse ortholog, orf34. The human and mouse transcripts were cloned from brain cDNA and encode deduced proteins of 1096 and 1091 amino acids, respectively. These proteins are 92% identical and 95% similar at the protein level. CXorf17 appears to be expressed at low levels and could be detected by RT-PCR in several adult and fetal human tissues. Analysis of the deduced amino acid sequence identified five putative transmembrane domains but no significant homology to previously described protein domains or sequence motifs. The CXorf17 protein has homology to two other non-annotated human proteins, C9orf10 and BC012177, the sequence similarity between them being strongest across two discrete domains of 250-270 amino acids in the N- and C-terminal parts of their sequences. We propose that these proteins belong to a previously undescribed family of putative transmembrane proteins. The identification of ESTs coding for similar proteins in other chordates but not lower eukaryotes suggests that these proteins may have first evolved during early chordate evolution. CXorf17 consists of 16 coding exons and maps to Xp11.22, approximately 14 kb telomeric to PRKWNK3 and 27 kb centromeric to KIAA1111. Its identification contributes to the annotation of expressed genes in the proximal part of the X chromosome.

Adult↗

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R. longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712 bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome↗

Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.

The spliced alignment of expressed sequence data to genomic sequence has proven a key tool in the comprehensive annotation of genes in eukaryotic genomes. A novel algorithm was developed to assemble clusters of overlapping transcript alignments (ESTs and full-length cDNAs) into maximal alignment assemblies, thereby comprehensively incorporating all available transcript data and capturing subtle splicing variations. Complete and partial gene structures identified by this method were used to improve The Institute for Genomic Research Arabidopsis genome annotation (TIGR release v.4.0). The alignment assemblies permitted the automated modeling of several novel genes and >1000 alternative splicing variations as well as updates (including UTR annotations) to nearly half of the approximately 27 000 annotated protein coding genes. The algorithm of the Program to Assemble Spliced Alignments (PASA) tool is described, as well as the results of automated updates to Arabidopsis gene annotations.

Algorithms↗

The genome sequence of Schizosaccharomyces pombe.

We have sequenced and annotated the genome of fission yeast (Schizosaccharomyces pombe), which contains the smallest number of protein-coding genes yet recorded for a eukaryote: 4,824. The centromeres are between 35 and 110 kilobases (kb) and contain related repeats including a highly conserved 1.8-kb element. Regions upstream of genes are longer than in budding yeast (Saccharomyces cerevisiae), possibly reflecting more-extended control regions. Some 43% of the genes contain introns, of which there are 4,730. Fifty genes have significant similarity with human disease genes; half of these are cancer related. We identify highly conserved genes important for eukaryotic cell organization including those required for the cytoskeleton, compartmentation, cell-cycle control, proteolysis, protein phosphorylation and RNA splicing. These genes may have originated with the appearance of eukaryotic life. Few similarly conserved genes that are important for multicellular organization were identified, suggesting that the transition from prokaryotes to eukaryotes required more new genes than did the transition from unicellular to multicellular organization.

Base Sequence↗