Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Theileria parva genomics reveals an atypical apicomplexan genome.

The discipline of genomics is setting new paradigms in research approaches to resolving problems in human and animal health. We propose to determine the genome sequence of Theileria parva, a pathogen of cattle, using the random shotgun approach pioneered at The Institute for Genomic Research (TIGR). A number of features of the T. parva genome make it particularly suitable for this approach. The G+C content of genomic DNA is about 31%, non-coding repetitive DNA constitutes less than 1% of total DNA and a framework for the 10-12 Mbp genome is available in the form of a physical map for all four chromosomes. Minisatellite sequences are the only dispersed repetitive sequences identified so far, but they are limited in distribution to 13 of 33 SfiI fragments. Telomere and sub-telomeric non-coding sequences occupy less than 10 kbp at each chromosomal end and there are only two units encoding cytoplasmic rRNAs. Three sets of distinct multicopy sequences encoding ORFs have been identified but it is not known if these are associated with expression of parasite antigenic diversity. Protein coding genes exhibit a bias in codon usage and introns when present are unusually short. Like other apicomplexan organisms, T. parva contains two extrachromosomal DNAs, a mitochondrial DNA and a plastid DNA molecule. By annotating the genome sequence, in combination with the use of microarray technology and comparative genomics, we expect to gain significant insights into unique aspects of the biology of T. parva. We believe that the data will underpin future research to aid in the identification of targets of protective CD8+ cell mediated immune responses, and parasite molecules involved in inducing reversible host leukocyte transformation and tumour-like behaviour of transformed parasitised cells.

Animals↗

Extensive search for discriminative features of alternative splicing.

Alternative pre-mRNA splicing events can be classified into various types, including cassette, mutually exclusive, alternative 3' splice site, alternative 5' splice site, retained intron. The detection of features of a particular type of alternative splicing events is an important and challenging problem in understanding the mechanism of alternative splicing. In this paper, we consider the problem of finding regulatory sequence patterns, which are specific to a particular type of alternative splicing events, on alternative exons and their flanking introns. For this problem, we have designed various pattern features and evaluated them on the alternative splicing data compiled in Lee's ASAP (Alternative Splicing Annotation Project) database. Through our work, we have succeeded in finding features with practically high accuracies.

Alternative Splicing↗

Organization of the MASP2 locus and its expression profile in mouse and rat.

The mouse, rat, and human MASP2 loci are situated on syntenic chromosome regions and are highly conserved. They comprise the genes for MASP-2/ MAp19, TAR DNA binding protein of 43 kDa, FRAP kinase, CDT6, Polymyositis-Scleroderma 100-kDa autoantigen, spermidine synthase, and TERE which were analyzed by annotation of available gene transcript data and cross-species comparison of available genomic sequences. The human and rat genes for spermidine synthase have an additional intron compared to the mouse gene. The mouse and rat genes for Polymyositis-Scleroderma 100-kDa autoantigen have an additional exon compared to the human gene. We find support for the hypothesis that the MAp19-specific exon within the MASP2 gene may have originated in a transposable element. Blocks of highly conserved intronic sequences were found in the MASP2 gene and the TARDBP gene. The expression of all genes within the MASP2 locus was analyzed in mouse and rat. The restricted expression of MASP-2 and MAp19 mRNA in liver contrasts with the ubiquitous expression of all neighboring genes studied.

Animals↗

Frame: detection of genomic sequencing errors.

MOTIVATION: The underlying error rate for genomic sequencing sometimes results in the introduction of artificial frameshifts and in-frame stop codons into putative protein encoding genes. Severe errors are then introduced into the inferred transcripts through mis-translation or premature termination. RESULTS: We describe a system for screening segments of DNA for frameshift and in-frame stop errors in coding regions. The method is based on homology matching using blastx to compare all six reading frames of the query nucleotide sequence against selected protein sequence databases. Fragments of protein matching neighbouring regions of the query DNA are united and extended laterally to define candidate open reading frames, within which, frameshifts and stops are identified. Suitable targets include prokaryotic or other intron-free genomic sequence and complementary DNAs. As an example of its use, we report here two frameshifted ORFs that deviate from the original TIGR sequence annotations for the recently released Helicobacter pylori genome. AVAILABILITY: The tool is accessible via the URL http://www.sander.ebi.ac.uk/frame/. CONTACT: brown@ebi.ac.uk.

Amino Acid Sequence↗

Gene structure conservation aids similarity based gene prediction.

One of the primary tasks in deciphering the functional contents of a newly sequenced genome is the identification of its protein coding genes. Existing computational methods for gene prediction include ab initio methods which use the DNA sequence itself as the only source of information, comparative methods using multiple genomic sequences, and similarity based methods which employ the cDNA or protein sequences of related genes to aid the gene prediction. We present here an algorithm implemented in a computer program called Projector which combines comparative and similarity approaches. Projector employs similarity information at the genomic DNA level by directly using known genes annotated on one DNA sequence to predict the corresponding related genes on another DNA sequence. It therefore makes explicit use of the conservation of the exon-intron structure between two related genes in addition to the similarity of their encoded amino acid sequences. We evaluate the performance of Projector by comparing it with the program Genewise on a test set of 491 pairs of independently confirmed mouse and human genes. It is more accurate than Genewise for genes whose proteins are <80% identical, and is suitable for use in a combined gene prediction system where other methods identify well conserved and non-conserved genes, and pseudogenes.

Algorithms↗

Human cytosolic sulfotransferase database mining: identification of seven novel genes and pseudogenes.

A total of 10 SULT genes are presently known to be expressed in human tissues. We performed a comprehensive genome-wide search for novel SULT genes using two different but complementary approaches, and developed a novel graphical display to aid in the annotation of the hits. Seven novel human SULT genes were identified, five of which were predicted to be pseudogenes, including two processed pseudogenes and three pseudogenes that contained introns. Those five pseudogenes represent the first unambiguous SULT pseudogenes described in any species. Expression-profiling studies were conducted for one novel gene, SULT6B1, and a series of alternatively spliced transcripts were identified in the human testis. SULT6B1 was also present in chimpanzee and gorilla, differing at only seven encoded amino-acid residues among the three species. The results of these database mining studies will aid in studies of the regulation of these SULT genes, provide insights into the evolution of this gene family in humans, and serve as a starting point for comparative genomic studies of SULT genes.

Amino Acid Sequence↗

Identification and evolutionary analysis of novel exons and alternative splicing events using cross-species EST-to-genome comparisons in human, mouse and rat.

BACKGROUND: Alternative splicing (AS) is important for evolution and major biological functions in complex organisms. However, the extent of AS in mammals other than human and mouse is largely unknown, making it difficult to study AS evolution in mammals and its biomedical implications. RESULTS: Here we describe a cross-species EST-to-genome comparison algorithm (ENACE) that can identify novel exons for EST-scanty species and distinguish conserved and lineage-specific exons. The identified exons represent not only novel exons but also evolutionarily meaningful AS events that are not previously annotated. A genome-wide AS analysis in human, mouse and rat using ENACE reveals a total of 758 novel cassette-on exons and 167 novel retained introns that have no EST evidence from the same species. RT-PCR-sequencing experiments validated approximately 50 approximately 80% of the tested exons, indicating high presence of exons predicted by ENACE. ENACE is particularly powerful when applied to closely related species. In addition, our analysis shows that the ENACE-identified AS exons tend not to pass the nonsynonymous-to-synonymous substitution ratio test and not to contain protein domain, implying that such exons may be under positive selection or relaxed negative selection. These AS exons may contribute to considerable inter-species functional divergence. Our analysis further indicates that a large number of exons may have been gained or lost during mammalian evolution. Moreover, a functional analysis shows that inter-species divergence of AS events may be substantial in protein carriers and receptor proteins in mammals. These exons may be of interest to studies of AS evolution. The ENACE programs and sequences of the ENACE-identified AS events are available for download. CONCLUSION: ENACE can identify potential novel cassette exons and retained introns between closely related species using a comparative approach. It can also provide information regarding lineage- or species-specificity in transcript isoforms, which are important for evolutionary and functional studies.

Algorithms↗

SpliceInfo: an information repository for mRNA alternative splicing in human genome.

We have developed an information repository named SpliceInfo to collect the occurrences of the four major alternative-splicing (AS) modes in human genome; these include exon skipping, 5'-alternative splicing, 3'-alternative splicing and intron retention. The dataset is derived by comparing the nucleotide and protein sequences available for a given gene for evidence of AS. Additional features such as the tissue specificity of the mRNA, the protein domain contained by exons, the GC-ratio of exons, the repeats contained within the exons, and the Gene Ontology are annotated computationally for each exonic region that is alternatively spliced. Motivated by a previous investigation of AS-related motifs such as exonic splicing enhancer and exonic splicing silencer, this resource also provides a means of identifying motifs candidates and this should help to identify potential regulatory mechanisms within a particular exonic sequence set and its two flanking intronic sequence sets. This is carried out using motif discovery tools to identify motif candidates related to alternative splicing regulation and together with a secondary structure prediction tool, will help in the identification of the structural properties of such regulatory motifs. The integrated resource is now available on http://SpliceInfo.mbc.NCTU.edu.tw/.

Alternative Splicing↗

Identification of new human cadherin genes using a combination of protein motif search and gene finding methods.

We have combined protein motif search and gene finding methods to identify genes encoding proteins containing specific domains. Particularly, we have focused on finding new human genes of the cadherin superfamily proteins, which represent a major group of cell-cell adhesion receptors contributing to embryonic neuronal morphogenesis. Models for three cadherin protein motifs were generated from over 100 already annotated cadherin domains and used to search the complete translated human genome. The genomic sequence regions containing motif "hits" were analyzed by eukaryotic GeneMark.hmm to identify the exon-intron structure of new genes. Three new genes CDH-J, PCDH-J and FAT-J were found. The predicted proteins PCDH-J and FAT-J were classified into protocadherin and FAT-like subfamilies, respectively, based on the number and organization of cadherin domains and presence of subfamily-specific conserved amino acid residues. Expression of FAT-J was shown in almost all tested tissues. The exon-intron organization of CDH-J was experimentally verified by PCR with specifically designed primers and its tissue-specific expression was demonstrated. The described methodology can be applied to discover new genes encoding proteins from families with well-characterized structural and functional domains.

Amino Acid Motifs↗

DNAView: a quality assessment tool for the visualization of large sequenced regions.

This communication describes DNAView, a graphical tool for the visualization and printing of large nucleic acid sequences. DNAView uses color coding to compactly display genomic segments of up to 100 kb on a single printed page. The specific color schemes integrated into DNAView can highlight 'local aggregate' properties of large segments of DNA. We have also incorporated a confidence expression for the assigned sequence. This is represented by base color intensity that is proportional to the number of times that base was sequenced. Areas of interest, such as exons, introns, repetitive elements and splice sites, can be emphasized using overlays. The colored image can be saved in a standard TIFF image file format that may be imported and annotated by other application software.

Base Sequence↗

AceView: a comprehensive cDNA-supported gene and transcripts annotation.

BACKGROUND: Regions covering one percent of the genome, selected by ENCODE for extensive analysis, were annotated by the HAVANA/Gencode group with high quality transcripts, thus defining a benchmark. The ENCODE Genome Annotation Assessment Project (EGASP) competition aimed at reproducing Gencode and finding new genes. The organizers evaluated the protein predictions in depth. We present a complementary analysis of the mRNAs, including alternative transcript variants. RESULTS: We evaluate 25 gene tracks from the University of California Santa Cruz (UCSC) genome browser. We either distinguish or collapse the alternative splice variants, and compare the genomic coordinates of exons, introns and nucleotides. Whole mRNA models, seen as chains of introns, are sorted to find the best matching pairs, and compared so that each mRNA is used only once. At the mRNA level, AceView is by far the closest to Gencode: the vast majority of transcripts of the two methods, including alternative variants, are identical. At the protein level, however, due to a lack of experimental data, our predictions differ: Gencode annotates proteins in only 41% of the mRNAs whereas AceView does so in virtually all. We describe the driving principles of AceView, and how, by performing hand-supervised automatic annotation, we solve the combinatorial splicing problem and summarize all of GenBank, dbEST and RefSeq into a genome-wide non-redundant but comprehensive cDNA-supported transcriptome. AceView accuracy is now validated by Gencode. CONCLUSION: Relative to a consensus mRNA catalog constructed from all evidence-based annotations, Gencode and AceView have 81% and 84% sensitivity, and 74% and 73% specificity, respectively. This close agreement validates a richer view of the human transcriptome, with three to five times more transcripts than in UCSC Known Genes (sensitivity 28%), RefSeq (sensitivity 21%) or Ensembl (sensitivity 19%).

Computational Biology↗

Multiple effects govern endogenous retrovirus survival patterns in human gene introns.

BACKGROUND: Endogenous retroviruses (ERVs) and solitary long terminal repeats (LTRs) have a significant antisense bias when located in gene introns, suggesting strong negative selective pressure on such elements oriented in the same transcriptional direction as the enclosing gene. It has been assumed that this bias reflects the presence of strong transcriptional regulatory signals within LTRs but little work has been done to investigate this phenomenon further. RESULTS: In the analysis reported here, we found significant differences between individual human ERV families in their prevalence within genes and degree of antisense bias and show that, regardless of orientation, ERVs of most families are less likely to be found in introns than in intergenic regions. Examination of density profiles of ERVs across transcriptional units and the transcription signals present in the consensus ERVs suggests the importance of splice acceptor sites, in conjunction with splice donor and polyadenylation signals, as the major targets for selection against most families of ERVs/LTRs. Furthermore, analysis of annotated human mRNA splicing events involving ERV sequence revealed that the relatively young human ERVs (HERVs), HERV9 and HERV-K (HML-2), are involved in no human mRNA splicing events at all when oriented antisense to gene transcription, while elements in the sense direction in transcribed regions show considerable bias for use of strong splice sites. CONCLUSION: Our observations suggest suppression of splicing among young intronic ERVs oriented antisense to gene transcription, which may account for their reduced mutagenicity and higher fixation rate in gene introns.

Base Sequence↗

Cloning and sequencing of cDNAs for hypothetical genes from chromosome 2 of Arabidopsis.

About 25% of the genes in the fully sequenced and annotated Arabidopsis genome have structures that are predicted solely by computer algorithms with no support from either nucleic acid or protein homologs from other species or expressed sequence matches from Arabidopsis. These are referred to as "hypothetical genes." On chromosome 2, sequenced by The Institute for Genomic Research, there are approximately 800 hypothetical genes among a total of approximately 4,100 genes. To test their expression under various growth conditions and in specific tissues, we used six cDNA populations prepared from cold-treated, heat-treated, and pathogen (Xanthomonas campestris pv campestris)-infected plants, callus, roots, and young seedlings. To date, 169 hypothetical genes were tested, and 138 of them are found to be expressed in one or more of the six cDNA populations. By sequencing multiple clones from each 5'- and 3'-rapid amplification of cDNA ends (RACE) product and assembling the sequences, we generated full-length sequences for 16 of these genes. For 14 genes, there was one full-length assembly that precisely supported the intron-exon boundaries of their gene predictions, adding only 5'- and 3'-untranslated region sequences. However, for three of these genes, the other assemblies represent additional exons and alternatively spliced or unspliced introns. For the remaining two genes, the cDNA sequences reveal major differences with predicted gene structures. In addition, a total of six genes displayed more than one polyadenylation site. These data will be used to update gene models in The Institute for Genomic Research annotation database ATH1.

Alternative Splicing↗

Evolutionary expansion, gene structure, and expression of the rice wall-associated kinase gene family.

The wall-associated kinase (WAK) gene family, one of the receptor-like kinase (RLK) gene families in plants, plays important roles in cell expansion, pathogen resistance, and heavy-metal stress tolerance in Arabidopsis (Arabidopsis thaliana). Through a reiterative database search and manual reannotation, we identified 125 OsWAK gene family members from rice (Oryza sativa) japonica cv Nipponbare; 37 (approximately 30%) OsWAKs were corrected/reannotated from earlier automated annotations. Of the 125 OsWAKs, 67 are receptor-like kinases, 28 receptor-like cytoplasmic kinases, 13 receptor-like proteins, 12 short genes, and five pseudogenes. The two-intron gene structure of the Arabidopsis WAK/WAK-Likes is generally conserved in OsWAKs; however, extra/missed introns were observed in some OsWAKs either in extracellular regions or in protein kinase domains. In addition to the 38 OsWAKs with full-length cDNA sequences and the 11 with rice expressed sequence tag sequences, gene expression analyses, using tiling-microarray analysis of the 20 OsWAKs on chromosome 10 and reverse transcription-PCR analysis for five OsWAKs, indicate that the majority of identified OsWAKs are likely expressed in rice. Phylogenetic analyses of OsWAKs, Arabidopsis WAK/WAK-Likes, and barley (Hordeum vulgare) HvWAKs show that the OsWAK gene family expanded in the rice genome due to lineage-specific expansion of the family in monocots. Localized gene duplications appear to be the primary genetic event in OsWAK gene family expansion and the 125 OsWAKs, present on all 12 chromosomes, are mostly clustered.

Arabidopsis↗

Genomic characterization of a repetitive motif strongly associated with developmental genes in Drosophila.

BACKGROUND: Non-coding DNA represents a high proportion of all metazoan genomes. Although an undetermined fraction of this DNA may be considered devoid of any function, it also contains important information residing in specific cis-regulatory sequences. RESULTS: We report a 27 bp motif that is overrepresented within the fly genome. This motif does not show any significant similarity with transposon sequences and is strongly associated with genes involved in development and/or signal transduction. The 27 bp motif is preferentially located within introns, and has a tendency to be present in multiple copies around genes. Furthermore, it is often found embedded in known non-coding regulatory regions. The regulatory network defined by this motif is partially shared in D. pseudoobscura. CONCLUSION: We have identified a 27 bp cis-regulatory sequence widely distributed within the Drosophila genome in association with developmental genes. This motif may be very useful towards the annotation of functional regulatory regions within the Drosophila genome and the construction of regulatory networks of Drosophila development.

Animals↗

Allelic imbalance at intragenic markers of Tbx18 is a hallmark of murine osteosarcoma.

We have recently identified a locus exhibiting a high frequency of allelic imbalance (AI) in both spontaneous human (HSA 6q14.1-15) and radiogenic murine (MMU9, 42 cM) osteosarcoma. Here we describe the fine mapping of the locus in osteosarcoma arising in (BALB/cxCBA) F(1) hybrid mice. These studies have allowed us to identify Tbx18, a member of the T-box transcriptional regulator gene family, as a candidate gene. Three intragenic Tbx18 polymorphisms were used to map the region of maximum AI to within the gene itself; 16 of 17 tumours exhibited imbalances of at least one of these markers. The highest frequency was found in exon 1, where 14 of 17 tumours were affected at a single nucleotide polymorphism at 541 nt. Two polymorphic CA repeat markers in intron 2 and intron 5 demonstrated overlapping regions of imbalance in several tumours. Both markers flanking the Tbx18 gene (D9Osm48 and D9Mit269) revealed significantly lower frequencies of imbalance and confirmed the limitation of the common interval to Tbx18. Examination of both the mouse and human annotated genomic sequences indicated Tbx18 to be the only gene within the interval. Sequence analysis of the Tbx18 coding region did not reveal any evidence of mutation. Given the haploinsufficiency phenotypes reported for other T-box genes, we speculate that AI may influence the function of Tbx18 during osteosarcomagenesis.

Alleles↗

Three-way translocation involves MLL, MLLT3, and a novel cell cycle control gene, FLJ10374, in the pathogenesis of acute myeloid leukemia with t(9;11;19)(p22;q23;p13.3).

The MLL gene, at 11q23, undergoes chromosomal translocation with a large number of partner genes in both acute lymphoblastic and acute myeloid leukemia (AML). We report a novel t(9;11;19)(p22;q23;p13.3) disrupting MLL in an infant AML patient. The 5' end of MLL fused to chromosome 9 sequences on the der(11), whereas the 3' end was translocated to chromosome 19. We developed long-distance inverse-polymerase chain reaction assays to investigate the localization of the breakpoints on der(11) and der(19). We found that intron 5 of MLL was fused to intron 5 of MLLT3 at the der(11) genomic breakpoint, resulting in a novel in-frame MLL exon 5-MLLT3 exon 6 fusion transcript. On the der(19), a novel gene annotated as FLJ10374 was disrupted by the breakpoint. Using reverse transcription-polymerase chain reaction analysis, we showed that FLJ10374 is ubiquitously expressed in human cells. Transfection of the FLJ10374 protein in different cell lines revealed that it localized exclusively to the nucleus. In serum-starved NIH-3T3 cells, the expression of FLJ10374 decreased the rate of the G1-to-S transition of the cell cycle, whereas the suppression of FLJ10374 through short interfering RNA increased cell proliferation. These results indicate that FLJ10374 negatively regulates cell cycle progression and proliferation. Thus, a single chromosomal rearrangement resulting in formation of the MLL-MLLT3 fusion gene and haplo-insufficiency of FLJ10374 may have cooperated to promote leukemogenesis in AML with t(9;11;19).

Acute Disease↗

Biological function of unannotated transcription during the early development of Drosophila melanogaster.

Many animal and plant genomes are transcribed much more extensively than current annotations predict. However, the biological function of these unannotated transcribed regions is largely unknown. Approximately 7% and 23% of the detected transcribed nucleotides during D. melanogaster embryogenesis map to unannotated intergenic and intronic regions, respectively. Based on computational analysis of coordinated transcription, we conservatively estimate that 29% of all unannotated transcribed sequences function as missed or alternative exons of well-characterized protein-coding genes. We estimate that 15.6% of intergenic transcribed regions function as missed or alternative transcription start sites (TSS) used by 11.4% of the expressed protein-coding genes. Identification of P element mutations within or near newly identified 5' exons provides a strategy for mapping previously uncharacterized mutations to their respective genes. Collectively, these data indicate that at least 85% of the fly genome is transcribed and processed into mature transcripts representing at least 30% of the fly genome.

Amino Acid Sequence↗