Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Dual Aberrant Splicing Caused by an Apparently Missense CHD7 Variant, c.5273A>G (p.Asp1758Gly), in CHARGE Syndrome.

CHARGE syndrome is a rare congenital disorder primarily attributed to heterozygous pathogenic variants of the CHD7 gene. Most pathogenic CHD7 variants are loss-of-function (LoF) variants, whereas the interpretation of missense variants remains challenging in the absence of functional evidence for their pathogenicity. We report a female infant presenting with clinical features characteristic of CHARGE syndrome. Targeted sequencing identified a heterozygous CHD7 variant (NM_017780.4:c.5273A>G), initially annotated as a missense substitution p.Asp1758Gly. This variant has been previously reported and registered with conflicting pathogenicity classifications; however, its transcript-level consequences remain unclear. Long-PCR-based RNA sequencing of total RNA from peripheral blood mononuclear cells revealed two aberrant splicing patterns associated with the variant: a predominant transcript carrying a 28-bp deletion due to cryptic donor splice-site activation, and a minor transcript with partial intron 24 retention. Both transcripts were predicted to result in premature termination codons. These findings demonstrate that c.5273A>G functions as a LoF variant through dual aberrant splicing rather than a simple missense substitution. This case underscores the importance of RNA-level splicing analysis for the accurate interpretation and classification of CHD7 missense variants.

CHD7↗

C. elegans ORFeome version 1.1: experimental verification of the genome annotation and resource for proteome-scale protein expression.

To verify the genome annotation and to create a resource to functionally characterize the proteome, we attempted to Gateway-clone all predicted protein-encoding open reading frames (ORFs), or the 'ORFeome,' of Caenorhabditis elegans. We successfully cloned approximately 12,000 ORFs (ORFeome 1.1), of which roughly 4,000 correspond to genes that are untouched by any cDNA or expressed-sequence tag (EST). More than 50% of predicted genes needed corrections in their intron-exon structures. Notably, approximately 11,000 C. elegans proteins can now be expressed under many conditions and characterized using various high-throughput strategies, including large-scale interactome mapping. We suggest that similar ORFeome projects will be valuable for other organisms, including humans.

Alternative Splicing↗

Improving gene annotation using peptide mass spectrometry.

Annotation of protein-coding genes is a key goal of genome sequencing projects. In spite of tremendous recent advances in computational gene finding, comprehensive annotation remains a challenge. Peptide mass spectrometry is a powerful tool for researching the dynamic proteome and suggests an attractive approach to discover and validate protein-coding genes. We present algorithms to construct and efficiently search spectra against a genomic database, with no prior knowledge of encoded proteins. By searching a corpus of 18.5 million tandem mass spectra (MS/MS) from human proteomic samples, we validate 39,000 exons and 11,000 introns at the level of translation. We present translation-level evidence for novel or extended exons in 16 genes, confirm translation of 224 hypothetical proteins, and discover or confirm over 40 alternative splicing events. Polymorphisms are efficiently encoded in our database, allowing us to observe variant alleles for 308 coding SNPs. Finally, we demonstrate the use of mass spectrometry to improve automated gene prediction, adding 800 correct exons to our predictions using a simple rescoring strategy. Our results demonstrate that proteomic profiling should play a role in any genome sequencing project.

Algorithms↗

Comparison of the oxidative phosphorylation (OXPHOS) nuclear genes in the genomes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

BACKGROUND: In eukaryotic cells, oxidative phosphorylation (OXPHOS) uses the products of both nuclear and mitochondrial genes to generate cellular ATP. Interspecies comparative analysis of these genes, which appear to be under strong functional constraints, may shed light on the evolutionary mechanisms that act on a set of genes correlated by function and subcellular localization of their products. RESULTS: We have identified and annotated the Drosophila melanogaster, D. pseudoobscura and Anopheles gambiae orthologs of 78 nuclear genes encoding mitochondrial proteins involved in oxidative phosphorylation by a comparative analysis of their genomic sequences and organization. We have also identified 47 genes in these three dipteran species each of which shares significant sequence homology with one of the above-mentioned OXPHOS orthologs, and which are likely to have originated by duplication during evolution. Gene structure and intron length are essentially conserved in the three species, although gain or loss of introns is common in A. gambiae. In most tissues of D. melanogaster and A. gambiae the expression level of the duplicate gene is much lower than that of the original gene, and in D. melanogaster at least, its expression is almost always strongly testis-biased, in contrast to the soma-biased expression of the parent gene. CONCLUSIONS: Quickly achieving an expression pattern different from the parent genes may be required for new OXPHOS gene duplicates to be maintained in the genome. This may be a general evolutionary mechanism for originating phenotypic changes that could lead to species differentiation.

Animals↗

Systematic identification of genes with coding microsatellites mutated in DNA mismatch repair-deficient cancer cells.

Microsatellite instability (MSI) caused by deficient DNA mismatch-repair functions is a hallmark of cancers associated with the hereditary nonpolyposis colorectal cancer (HNPCC) syndrome but is also found in about 15% of all sporadic tumors. Most affected microsatellites reside in untranslated intergenic or intronic sequences. However, recently few genes with coding microsatellites were also shown to be mutational targets in MSI-positive cancers and might represent important mutation targets in their pathogenesis. The systematic identification of such genes and the analysis of their mutation frequency in MSI-positive cancers might thus reveal major clues to their functional role in MSI-associated carcinogenesis. We therefore initiated a systematic database search in 33,595 distinctly annotated human genes and identified 17,654 potentially coding mononucleotide repeats (cMNRs) and 2,028 coding dinucleotide repeats (cDNRs), which consist of n > or = 6 and n > or = 4 repeat units, respectively. Expression pattern and mutation frequency of 19 of these genes with the longest repeats were compared between DNA mismatch repair-deficient (MSI(+)) and proficient (MSS) cancer cells. Instability frequencies in these coding microsatellite genes ranged from 10% to 100% in MSI-H tumor cells, whereas MSS cancer cells did not show mutations. RT-PCR analysis further showed that most of the affected genes (10/15) were highly expressed in tumor cells. The approach outlined here identified a new set of genes frequently affected by mutations in MSI-positive tumor cells. It will lead to novel and highly specific diagnostic and therapeutic targets for microsatellite unstable cancers.

Adenocarcinoma↗

Computational inference of homologous gene structures in the human genome.

With the human genome sequence approaching completion, a major challenge is to identify the locations and encoded protein sequences of all human genes. To address this problem we have developed a new gene identification algorithm, GenomeScan, which combines exon-intron and splice signal models with similarity to known protein sequences in an integrated model. Extensive testing shows that GenomeScan can accurately identify the exon-intron structures of genes in finished or draft human genome sequence with a low rate of false-positives. Application of GenomeScan to 2.7 billion bases of human genomic DNA identified at least 20,000-25,000 human genes out of an estimated 30,000-40,000 present in the genome. The results show an accurate and efficient automated approach for identifying genes in higher eukaryotic genomes and provide a first-level annotation of the draft human genome.

Algorithms↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

Compositional gradients in Gramineae genes.

In this study, we describe a property of Gramineae genes, and perhaps all monocot genes, that is not observed in eudicot genes. Along the direction of transcription, beginning at the junction of the 5'-UTR and the coding region, there are gradients in GC content, codon usage, and amino-acid usage. The magnitudes of these gradients are large enough to hinder the annotation of the rice genome and to confound the detection of protein homologies across the monocot-eudicot divide.

Amino Acids↗

Expressed sequence tag analysis in tef (Eragrostis tef (Zucc) Trotter).

Tef (Eragrostis tef (Zucc.) Trotter) is the most important cereal crop in Ethiopia; however, there is very little DNA sequence information available for this species. Expressed sequence tags (ESTs) were generated from 4 cDNA libraries: seedling leaf, seedling root, and inflorescence of E. tef and seedling leaf of Eragrostis pilosa, a wild relative of E. tef. Clustering of 3603 sequences produced 530 clusters and 1890 singletons, resulting in 2420 tef unigenes. Approximately 3/4 of tef unigenes matched protein or nucleotide sequences in public databases. Annotation of unigenes associated 68% of the putative tef genes with gene ontology categories. Identification of the translated unigenes for conserved protein domains revealed 389 protein family domains (Pfam), the most frequent of which was protein kinase. A total of 170 ESTs containing simple sequence repeats (EST-SSRs) were identified and 80 EST-SSR markers were developed. In addition, 19 single-nucleotide polymorphism (SNP) and (or) insertion-deletion (indel) and 34 intron fragment length polymorphism (IFLP) markers were developed. The EST database and molecular markers generated in this study will be valuable resources for further tef genetic research.

Databases, Nucleic Acid↗

Vertebrate gene predictions and the problem of large genes.

To find unknown protein-coding genes, annotation pipelines use a combination of ab initio gene prediction and similarity to experimentally confirmed genes or proteins. Here, we show that although the ab initio predictions have an intrinsically high false-positive rate, they also have a consistently low false-negative rate. The incorporation of similarity information is meant to reduce the false-positive rate, but in doing so it increases the false-negative rate. The crucial variable is gene size (including introns)--genes of the most extreme sizes, especially very large genes, are most likely to be incorrectly predicted.

Animals↗

SpliceMachine: predicting splice sites from high-dimensional local context representations.

MOTIVATION: In this age of complete genome sequencing, finding the location and structure of genes is crucial for further molecular research. The accurate prediction of intron boundaries largely facilitates the correct prediction of gene structure in nuclear genomes. Many tools for localizing these boundaries on DNA sequences have been developed and are available to researchers through the internet. Nevertheless, these tools still make many false positive predictions. RESULTS: This manuscript presents a novel publicly available splice site prediction tool named SpliceMachine that (i) shows state-of-the-art prediction performance on Arabidopsis thaliana and human sequences, (ii) performs a computationally fast annotation and (iii) can be trained by the user on its own data. AVAILABILITY: Results, figures and software are available at http://www.bioinformatics.psb.ugent.be/supplementary_data/ CONTACT: sven.degroeve@psb.ugent.be; yves.vandepeer@psb.ugent.be.

Algorithms↗

SMART 4.0: towards genomic data integration.

SMART (Simple Modular Architecture Research Tool) is a web tool (http://smart.embl.de/) for the identification and annotation of protein domains, and provides a platform for the comparative study of complex domain architectures in genes and proteins. The January 2004 release of SMART contains 685 protein domains. New developments in SMART are centred on the integration of data from completed metazoan genomes. SMART now uses predicted proteins from complete genomes in its source sequence databases, and integrates these with predictions of orthology. New visualization tools have been developed to allow analysis of gene intron-exon structure within the context of protein domain structure, and to align these displays to provide schematic comparisons of orthologous genes, or multiple transcripts from the same gene. Other improvements include the ability to query SMART by Gene Ontology terms, improved structure database searching and batch retrieval of multiple entries.

Algorithms↗

Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.

In contrast to the fairly reliable and complete annotation of the protein coding genes in the human genome, comparable information is lacking for noncoding RNAs (ncRNAs). We present a comparative screen of vertebrate genomes for structural noncoding RNAs, which evaluates conserved genomic DNA sequences for signatures of structural conservation of base-pairing patterns and exceptional thermodynamic stability. We predict more than 30,000 structured RNA elements in the human genome, almost 1,000 of which are conserved across all vertebrates. Roughly a third are found in introns of known genes, a sixth are potential regulatory elements in untranslated regions of protein-coding mRNAs and about half are located far away from any known gene. Only a small fraction of these sequences has been described previously. A comparison with recent tiling array data shows that more than 40% of the predicted structured RNAs overlap with experimentally detected sites of transcription. The widespread conservation of secondary structure points to a large number of functional ncRNAs and cis-acting mRNA structures in the human genome.

Animals↗

A study of the middle-scale nucleotide clustering in DNA sequences of various origin and functionality, by means of a method based on a modified standard deviation.

The deviation from randomness in the distribution of nucleotides in genomic sequences is quantified and studied, using a modified standard deviation (MSD). This method implies a "per block" computation of the standard deviation of the nucleotide frequencies of occurrence, using local means (means taken in a neighborhood of each block). This quantity may serve as a scale-dependent measure of the nucleotide clustering. In the present work, the meso-scale of tenths of nucleotides is principally explored, by means of suitably adjusted filter parameters. This length scale is of an order of magnitude not directly affected by the grammar and syntax rules of the protein-coding procedure, remaining shorter than the scale of appearance of large-scale characteristics of the genome. MSD has been found to distinguish systematically between the sequences of different origin and functionality. The most near-random are found to be coding sequences of prokaryotes, while in intronic and intergenic regions of eukaryotic genomes, extended clustering of similar nucleotides is observed. The distributions of MSD values of large collections of sequences are found to be in most cases characteristic of their biological role and origin. Protein- and non-coding, prokaryotic and eukaryotic DNA as well as promoter, rRNA, viral and organelle sequences have been examined. The presented results corroborate a recently proposed model for genome evolution. The method is also applied for an assessment of the annotation of ORFs taken from the complete genome of Saccharomyces cerevisiae.

Animals↗

In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists.

MOTIVATION: Prediction of the coding potential for stretches of DNA is crucial in gene calling and genome annotation, where it is used to identify potential exons and to position their boundaries in conjunction with functional sites, such as splice sites and translation initiation sites. The ability to discriminate between coding and non-coding sequences relates to the structure of coding sequences, which are organized in codons, and by their biased usage. For statistical reasons, the longer the sequences, the easier it is to detect this codon bias. However, in many eukaryotic genomes, where genes harbour many introns, both introns and exons might be small and hard to distinguish based on coding potential. RESULTS: Here, we present novel approaches that specifically aim at a better detection of coding potential in short sequences. The methods use complementary sequence features, combined with identification of which features are relevant in discriminating between coding and non-coding sequences. These newly developed methods are evaluated on different species, representative of four major eukaryotic kingdoms, and extensively compared to state-of-the-art Markov models, which are often used for predicting coding potential. The main conclusions drawn from our analyses are that (1) combining complementary sequence features clearly outperforms current Markov models for coding potential prediction in short sequence fragments, (2) coding potential prediction benefits from length-specific models, and these models are not necessarily the same for different sequence lengths and (3) comparing the results across several species indicates that, although our combined method consistently performs extremely well, there are important differences across genomes. SUPPLEMENTARY DATA: http://bioinformatics.psb.ugent.be/.

Algorithms↗

Diversity and relatedness among the type I interferons.

Type I interferons (IFNs) include the IFN-alpha family of subtypes, IFN-beta, IFN-omega, IFN-tau, IFN-kappa, IFN-lambda, and IFN-zeta. IFN genes lack introns and encode secretory signal peptide sequences that are proteolytically cleaved prior to secretion from the cell. In contrast to the approximately 50% amino acid sequence identity among the human IFN-alpha subtypes, human IFN-alphas share approximately 22% identity with human IFN-beta and 37% identity with human IFN-omega. Many of the conserved residues among the type I IFNs are implicated in receptor recognition and structural integrity. This report provides an update on the gene annotations for the mouse and human IFN gene clusters on chromosome 4 and 9, respectively, with accompanying amino acid sequence alignments. Based on sequence identities, a phylogenic tree analysis for the different mammalian Type I IFNs is also presented, showing the high degree of relatedness among these IFNs. Notably, sequence alignment of the different human and mouse IFN promoter regions reveals different signature patterns for transcription factor binding sites, implying different inducers might differentially activate the transcription of the different IFNs.

Amino Acid Sequence↗

Identification of rat genes by TWINSCAN gene prediction, RT-PCR, and direct sequencing.

The publication of a draft sequence of a third mammalian genome--that of the rat--suggests a need to rethink genome annotation. New mammalian sequences will not receive the kind of labor-intensive annotation efforts that are currently being devoted to human. In this paper, we demonstrate an alternative approach: reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on dual-genome de novo predictions from TWINSCAN. We tested 444 TWINSCAN-predicted rat genes that showed significant homology to known human genes implicated in disease but that were partially or completely missed by methods based on protein-to-genome mapping. Using primers in exons flanking a single predicted intron, we were able to verify the existence of 59% of these predicted genes. We then attempted to amplify the complete predicted open reading frames of 136 genes that were verified in the single-intron experiment. Spliced sequences were amplified in 46 cases (34%). We conclude that this procedure for elucidating gene structures with native cDNA sequences is cost-effective and will become even more so as it is further optimized.

Animals↗

ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.

Alternative splicing (AS) is now emerging as a major mechanism contributing to the expansion of the transcriptome and proteome complexity of multicellular organisms. The fact that a single gene locus may give rise to multiple mRNAs and protein isoforms, showing both major and subtle structural variations, is an exceptionally versatile tool in the optimization of the coding capacity of the eukaryotic genome. The huge and continuously increasing number of genome and transcript sequences provides an essential information source for the computational detection of genes AS pattern. However, much of this information is not optimally or comprehensively used in gene annotation by current genome annotation pipelines. We present here a web resource implementing the ASPIC algorithm which we developed previously for the investigation of AS of user submitted genes, based on comparative analysis of available transcript and genome data from a variety of species. The ASPIC web resource provides graphical and tabular views of the splicing patterns of all full-length mRNA isoforms compatible with the detected splice sites of genes under investigation as well as relevant structural and functional annotation. The ASPIC web resource-available at http://www.caspur.it/ASPIC/--is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility.

Algorithms↗