Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

265 records · Page 15Linked to original sources

Molecular identification of the first insect proctolin receptor.

The website of the Drosophila Genome Project (www.flybase.org) contains the sequence of an annotated gene CG6986, which is predicted to code for a G protein-coupled receptor. We cloned the cDNA of this gene and expressed it in Chinese hamster ovary cells. Screening of a neuropeptide library revealed that the expressed receptor was specific for the neuropeptide proctolin (EC(50), 6x10(-10)M). Proctolin (RYLPT) was the first invertebrate neuropeptide to be fully sequenced (already in 1975) and occurs with identical structure in both crustaceans and insects, where it has myo- and neurostimulatory actions. Northern blots showed that the Drosophila proctolin receptor was only weakly expressed in embryos, larvae, pupae, and in the thoraces and abdomina of adult flies, but strongly in the heads of adult animals. The Drosophila receptor reported here is the first invertebrate proctolin receptor to be identified.

Amino Acid Sequence↗

Recognizing shorter coding regions of human genes based on the statistics of stop codons.

With the quick progress of the Human Genome Project, a great amount of uncharacterized DNA sequences needs to be annotated copiously by better algorithms. Recognizing shorter coding sequences of human genes is one of the most important problems in gene recognition, which is not yet completely solved. This paper is devoted to solving the issue using a new method. The distributions of the three stop codons, i.e., TAA, TAG and TGA, in three phases along coding, noncoding, and intergenic sequences are studied in detail. Using the obtained distributions and other coding measures, a new algorithm for the recognition of shorter coding sequences of human genes is developed. The accuracy of the algorithm is tested based on a larger database of human genes. It is found that the average accuracy achieved is as high as 92.1% for the sequences with length of 192 base pairs, which is confirmed by sixfold cross-validation tests. It is hoped that by incorporating the present method with some existing algorithms, the accuracy for identifying human genes from unannotated sequences would be increased.

Algorithms↗

The Drosophila gene CG9918 codes for a pyrokinin-1 receptor.

The database from the Drosophila Genome Project contains a gene, CG9918, annotated to code for a G protein-coupled receptor. We cloned the cDNA of this gene and functionally expressed it in Chinese hamster ovary cells. We tested a library of about 25 Drosophila and other insect neuropeptides, and seven insect biogenic amines on the expressed receptor and found that it was activated by low concentrations of the Drosophila neuropeptide, pyrokinin-1 (TGPSASSGLWFGPRLamide; EC50, 5 x 10(-8) M). The receptor was also activated by other Drosophila neuropeptides, terminating with the sequence PRLamide (Hug-gamma, ecdysis-triggering-hormone-1, pyrokinin-2), but in these cases about six to eight times higher concentrations were needed. The receptor was not activated by Drosophila neuropeptides, containing a C-terminal PRIamide sequence (such as ecdysis-triggering-hormone-2), or PRVamide (such as capa-1 and -2), or other neuropeptides and biogenic amines not related to the pyrokinins. This paper is the first conclusive report that CG9918 is a Drosophila pyrokinin-1 receptor gene.

Amino Acid Sequence↗

Generation and analysis of 280,000 human expressed sequence tags.

We report the generation of 319,311 single-pass sequencing reactions (known as expressed sequence tags, or ESTs) obtained from the 5' and 3' ends of 194,031 human cDNA clones. Our goal has been to obtain tag sequences from many different genes and to deposit these in the publicly accessible Data Base for Expressed Sequence Tags. Highly efficient automatic screening of the data allows deposition of the annotated sequences without delay. Sequences have been generated from 26 oligo(dT) primed directionally cloned libraries, of which 18 were normalized. The libraries were constructed using mRNA isolated from 17 different tissues representing three developmental states. Comparisons of a subset of our data with nonredundant human mRNA and protein data bases show that the ESTs represent many known sequences and contain many that are novel. Analysis of protein families using Hidden Markov Models confirms this observation and supports the contention that although normalization reduces significantly the relative abundance of redundant cDNA clones, it does not result in the complete removal of members of gene families.

Adult↗

RNAi-induced phenotypes suggest a novel role for a chemosensory protein CSP5 in the development of embryonic integument in the honeybee (Apis mellifera).

Small chemosensory proteins (CSPs) belong to a conserved, but poorly understood, protein family found in insects and other arthropods. They exhibit both broad and restricted expression patterns during development. In this paper, we used a combination of genome annotation, transcriptional profiling and RNA interference to unravel the functional significance of a honeybee gene (csp5) belonging to the CSP family. We show that csp5 expression resembles the maternal-zygotic pattern that is characterized by the initiation of transcription in the ovary and the replacement of maternal mRNA with embryonic mRNA. Blocking the embryonic expression of csp5 with double-stranded RNA causes abnormalities in all body parts where csp5 is highly expressed. The treated embryos show a "diffuse", often grotesque morphology, and the head skeleton appears to be severely affected. They are 'unable-to-hatch' and cannot progress to the larval stages. Our findings reveal a novel, essential role for this gene family and suggest that csp5 (unable-to-hatch) is an ectodermal gene involved in embryonic integument formation. Our study confirms the utility of an RNAi approach to functional characterization of novel developmental genes uncovered by the honeybee genome project and provides a starting point for further studies on embryonic integument formation in this insect.

Amino Acid Sequence↗

Comparative analysis of the PCOLCE region in Fugu rubripes using a new automated annotation tool.

The Japanese pufferfish Fugu rubripes with a genome of about 400 Mb is becoming increasingly recognized as a vertebrate model organism for comparative gene analysis (see Elgar 1996 for review). We have isolated and sequenced two Fugu cosmids spanning a genomic region of 66 kb containing the Fugu homolog to the human PCOLCE-I (Glöckner et al. 1998). We then examined if RUMMAGE-DP, a newly developed analysis tool for gene discovery which was designed for human and mouse genomic DNA, can be used for automatic annotation of Fugu genomic sequence. The exon prediction programs contained in RUMMAGE-DP performed better overall for the human sequence than for the Fugu contig. The GENSCAN program was the only exon prediction programme that performed equally well for both organisms. We show that RUMMAGE-DP is very useful in automatic analysis of Fugu sequences. Comparative analysis of the genomic structure of the PCOLCE-I genes in Fugu and human reveals that the exon/intron structure throughout the protein coding region is almost identical. We defined an additional domain based on the high degree of similarity of 26 aa between mammals and Fugu. The PCOLCE-I protein in both organisms contains two highly conserved CUB domains. Exons 6 and 7 are the only coding exons that differ in length between the two species. We assume that these exons do not code for any catalytic domain of the protein. Analysis of the remaining five Fugu genes within the 66 kb interval revealed no conserved synteny with the corresponding human 7q22 region.

Amino Acid Sequence↗

The odorant-binding proteins of Drosophila melanogaster: annotation and characterization of a divergent gene family.

Insect odorant-binding proteins (OBPs) are thought to facilitate the delivery of hydrophobic odorants, such as sex pheromones or food odors, to receptors on sensory neurons. Increasingly, OBP family members are also being found in non-sensory tissues where they might carry other types of small hydrophobic molecules. They are identifiable by four or six conserved Cys residues and contain six alpha-helices which enclose a hydrophobic ligand-binding pocket. Through exhaustive BLAST searches we have increased the total number of OBPs identified in Drosophila melanogaster to 38, and have amplified the DNA complementary to RNA corresponding to 21 of these by reverse transcriptase polymerase chain reaction. Isoforms frequently share less than 30% amino acid identity and appear to have radically changed since the separation of the major insect orders. However, their sequences are consistent with known OBP structures. Most are located in clusters of between four and 14 genes and several were unusual in that they contained additions, deletions, or fusions. These hexa-helical insect OBPs are structurally unrelated to the functionally analogous lipocalin-like beta-barrel OBPs of vertebrates. As only two lipocalin-like proteins have been found in D. melanogaster, these helical proteins appear to be the dominant carrier of small hydrophobic molecules in insects.

Amino Acid Sequence↗

Cloning and characterization of multiple glycosyl hydrolase genes from Trichoderma virens.

Trichoderma virens is a widely distributed soil fungus that is parasitic on other soil fungi. The mycoparasitic activity of T. virens is correlated with the production of numerous antifungal activities, including the secretion of a considerable repertoire of fungal cell wall-degrading enzymes. Here, we report the characterization of a diverse set of chitinase and glucanase genes from T. virens. In each case, full-length genomic clones were isolated and characterized, while sequencing of the corresponding cDNA clones and manual annotation provided a basis for establishing gene structure. Based on homology of the deduced amino acid sequences, we have identified three members of the 42Kd endochitinase gene family, two 33Kd exochitinases, two exochitinases with homology to N-acetylglucosaminidases, and three glucanase genes predicted to encode beta-1,3- and beta-1,6-proteins. The majority of these genes appear to encode signal peptides, suggesting an extracellular location for the corresponding proteins. Despite their overall similarity, the 42Kd class of chitinases can be subdivided, based on the presence of distinct N-terminal domains, suggesting that the proteins may have distinct cellular roles, while Northern blot analysis confirms that these genes possess distinct patterns of gene regulation. Similarly, one of the 33Kd chitinase genes is unique, because it is predicted to encode a protein C-terminus with high homology to the conserved family I cellulose-binding domain. The expression patterns of the chitinase genes were analyzed in both a wild-type strain and a strain disrupted for the major 42Kd chitinase gene of T. virens. The results of these transcript analyses, together with enzymatic assay of the extracellular proteins, suggest interdependent regulation of this important gene family in T. virens.

Acetylglucosaminidase↗

DNA sequence and analysis of human chromosome 18.

Chromosome 18 appears to have the lowest gene density of any human chromosome and is one of only three chromosomes for which trisomic individuals survive to term. There are also a number of genetic disorders stemming from chromosome 18 trisomy and aneuploidy. Here we report the finished sequence and gene annotation of human chromosome 18, which will allow a better understanding of the normal and disease biology of this chromosome. Despite the low density of protein-coding genes on chromosome 18, we find that the proportion of non-protein-coding sequences evolutionarily conserved among mammals is close to the genome-wide average. Extending this analysis to the entire human genome, we find that the density of conserved non-protein-coding sequences is largely uncorrelated with gene density. This has important implications for the nature and roles of non-protein-coding sequence elements.

Aneuploidy↗

Human ARX gene: genomic characterization and expression.

Arx is a homeobox-containing gene with a high degree of sequence similarity between mouse and zebrafish. Arx is expressed in the forebrain and floor plate of the developing central nervous systems of these vertebrates and in the presumptive cortex of fetal mice. Our goal was to identify genes in Xp22.1-p21.3 involved in human neuronal development. Our in silico search for candidate genes noted that annotation of a human Xp22 PAC (RPCI1-258N20) sequence (GenBank Accession No. AC002504) identified putative exons consistent with an Arx homologue in Xp22. Northern blot analysis showed that a 3.3kb human ARX transcript was expressed at high levels in fetal brain. A 5.9kb transcript was expressed in adult heart, skeletal muscle, and liver with very faint expression in other adult tissues, including brain. In situ hybridization of ARX in human fetal brain sections at various developmental stages showed the highest expression in neuronal precursors in the germinal matrix of the ganglionic eminence and in the ventricular zone of the telencephalon. Expression was also observed in the hippocampus, cingulate, subventricular zone, cortical plate, caudate nucleus, and putamen. The expression pattern suggests that ARX is involved in the differentiation and maintenance of specific neuronal cell types in the human central nervous system. We also mapped the murine Arx gene to the mouse genome using a mouse/hamster radiation hybrid panel and showed that Arx and ARX are orthologues. Therefore, investigations in model vertebrates may provide insight into the role of ARX in development. The recent identification of ARX mutations in patients with various forms of mental retardation make such studies in model organisms even more compelling.

Amino Acid Sequence↗

Structural diversity and transcription of class III peroxidases from Arabidopsis thaliana.

Understanding peroxidase function in plants is complicated by the lack of substrate specificity, the high number of genes, their diversity in structure and our limited knowledge of peroxidase gene transcription and translation. In the present study we sequenced expressed sequence tags (ESTs) encoding novel heme-containing class III peroxidases from Arabidopsis thaliana and annotated 73 full-length genes identified in the genome. In total, transcripts of 58 of these genes have now been observed. The expression of individual peroxidase genes was assessed in organ-specific EST libraries and compared to the expression of 33 peroxidase genes which we analyzed in whole plants 3, 6, 15, 35 and 59 days after sowing. Expression was assessed in root, rosette leaf, stem, cauline leaf, flower bud and cell culture tissues using the gene-specific and highly sensitive reverse transcriptase-polymerase chain reaction (RT-PCR). We predicted that 71 genes could yield stable proteins folded similarly to horseradish peroxidase (HRP). The putative mature peroxidases derived from these genes showed 28-94% amino acid sequence identity and were all targeted to the endoplasmic reticulum by N-terminal signal peptides. In 20 peroxidases these signal peptides were followed by various N-terminal extensions of unknown function which are not present in HRP. Ten peroxidases showed a C-terminal extension indicating vacuolar targeting. We found that the majority of peroxidase genes were expressed in root. In total, class III peroxidases accounted for an impressive 2.2% of root ESTs. Rather few peroxidases showed organ specificity. Most importantly, genes expressed constitutively in all organs and genes with a preference for root represented structurally diverse peroxidases (< 70% sequence identity). Furthermore, genes appearing in tandem showed distinct expression profiles. The alignment of 73 Arabidopsis peroxidase sequences provides an easy access to the identification of orthologous peroxidases in other plant species and will provide a common platform for combining knowledge of peroxidase structure and function relationships obtained in various species.

Amino Acid Sequence↗

Large-scale transcriptome analyses reveal new genetic marker candidates of head, neck, and thyroid cancer.

A detailed genome mapping analysis of 213,636 expressed sequence tags (EST) derived from nontumor and tumor tissues of the oral cavity, larynx, pharynx, and thyroid was done. Transcripts matching known human genes were identified; potential new splice variants were flagged and subjected to manual curation, pointing to 788 putatively new alternative splicing isoforms, the majority (75%) being insertion events. A subset of 34 new splicing isoforms (5% of 788 events) was selected and 23 (68%) were confirmed by reverse transcription-PCR and DNA sequencing. Putative new genes were revealed, including six transcripts mapped to well-studied chromosomes such as 22, as well as transcripts that mapped to 253 intergenic regions. In addition, 2,251 noncoding intronic RNAs, eventually involved in transcriptional regulation, were found. A set of 250 candidate markers for loss of heterozygosis or gene amplification was selected by identifying transcripts that mapped to genomic regions previously known to be frequently amplified or deleted in head, neck, and thyroid tumors. Three of these markers were evaluated by quantitative reverse transcription-PCR in an independent set of individual samples. Along with detailed clinical data about tumor origin, the information reported here is now publicly available on a dedicated Web site as a resource for further biological investigation. This first in silico reconstruction of the head, neck, and thyroid transcriptomes points to a wealth of new candidate markers that can be used for future studies on the molecular basis of these tumors. Similar analysis is warranted for a number of other tumors for which large EST data sets are available.

Alternative Splicing↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗