Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

The planarian Schmidtea mediterranea as a model for epigenetic germ cell specification: analysis of ESTs from the hermaphroditic strain.

Freshwater planarians have prodigious regenerative abilities that enable them to form complete organisms from tiny body fragments. This plasticity is also exhibited by the planarian germ cell lineage. Unlike many model organisms in which germ cells are specified by localized determinants, planarian germ cells appear to be specified epigenetically, arising postembryonically from stem cells. The planarian Schmidtea mediterranea is well suited for investigating the mechanisms underlying epigenetic germ cell specification. Two strains of S. mediterranea exist: a hermaphroditic strain that reproduces sexually and an asexual strain that reproduces by means of transverse fission. To date, expressed sequence tags (ESTs) have been generated only from the asexual strain. To develop molecular reagents for studying epigenetic germ cell specification, we have sequenced 27,161 ESTs from two developmental stages of the hermaphroditic strain of S. mediterranea; this collection of ESTs represents approximately 10,000 unique transcripts. blast analysis of the assembled ESTs showed that 66% share similarity to sequences in public databases. We annotated the assembled ESTs using Gene Ontology terms as well as conserved protein domains and organized them in a relational database. To validate experimentally the Gene Ontology annotations, we used whole-mount in situ hybridization to examine the expression patterns of transcripts assigned to the biological process "reproduction." Of the 53 genes in this category, 87% were expressed in the reproductive organs. In addition to its utility for studying germ cell development, this EST collection will be an important resource for annotating the planarian genome and studying this animal's amazing regenerative abilities.

Animals↗

Arabidopsis thaliana: a model plant for genome analysis.

Arabidopsis thaliana is a small plant in the mustard family that has become the model system of choice for research in plant biology. Significant advances in understanding plant growth and development have been made by focusing on the molecular genetics of this simple angiosperm. The 120-megabase genome of Arabidopsis is organized into five chromosomes and contains an estimated 20,000 genes. More than 30 megabases of annotated genomic sequence has already been deposited in GenBank by a consortium of laboratories in Europe, Japan, and the United States. The entire genome is scheduled to be sequenced by the end of the year 2000. Reaching this milestone should enhance the value of Arabidopsis as a model for plant biology and the analysis of complex organisms in general.

Arabidopsis↗

Mapping of Myxococcus xanthus social motility dsp mutations to the dif genes.

Myxococcus xanthus dsp and dif mutants have similar phenotypes in that they are deficient in social motility and fruiting body development. We compared the two loci by genetic mapping, complementation with a cosmid clone, DNA sequencing, and gene disruption and found that 16 of the 18 dsp alleles map to the dif genes. Another dsp allele contains a mutation in the sglK gene. About 36.6 kb around the dsp-dif locus was sequenced and annotated, and 50% of the genes are novel.

Bacterial Proteins↗

Nineteen additional unpredicted transcripts from human chromosome 21.

The identification of all human chromosome 21 (HC21) genes is a necessary step in understanding the molecular pathogenesis of trisomy 21 (Down syndrome). The first analysis of the sequence of 21q included 127 previously characterized genes and predicted an additional 98 novel anonymous genes. Recently we evaluated the quality of this annotation by characterizing a set of HC21 open reading frames (C21orfs) identified by mapping spliced expressed sequence tags (ESTs) and predicted genes (PREDs), identified only in silico. This study underscored the limitations of in silico-only gene prediction, as many PREDs were incorrectly predicted. To refine the HC21 annotation, we have developed a reliable algorithm to extract and stringently map sequences that contain bona fide 3' transcript ends to the genome. We then created a specific 21q graphical display allowing an integrated view of the data that incorporates new ESTs as well as features such as CpG islands, repeats, and gene predictions. Using these tools we identified 27 new putative genes. To validate these, we sequenced previously cloned cDNAs and carried out RT-PCR, 5'- and 3'-RACE procedures, and comparative mapping. These approaches substantiated 19 new transcripts, thus increasing the HC21 gene count by 9.5%. These transcripts were likely not previously identified because they are small and encode small proteins. We also identified four transcriptional units that are spliced but contain no obvious open reading frame. The HC21 data presented here further emphasize that current gene prediction algorithms miss a substantial number of transcripts that nevertheless can be identified using a combination of experimental approaches and multiple refined algorithms.

Chromosomes, Human, Pair 21↗

Discovery of immune-related genes expressed in hemocytes of the tarantula spider Acanthoscurria gomesiana.

The present study reports the identification of immune related transcripts from hemocytes of the spider Acanthoscurria gomesiana by high throughput sequencing of expressed sequence tags (ESTs). To generate ESTs from hemocytes, two cDNA libraries were prepared: one by directional cloning (primary) and the other by the normalization of the first (normalized). A total of 7584 clones were sequenced and the identical ESTs were clustered, resulting in 3723 assembled sequences (AS). At least 20% of these sequences are putative novel genes. The automatic functional annotation of AS based on Gene Ontology revealed several abundant transcripts related to the following functional classes: hemocyanin, lectin, and structural constituents of ribosome and cytoskeleton. From this annotation, 73 transcripts possibly involved in immune response were also identified, suggesting the existence of several molecular processes not previously described for spiders, such as: pathogen recognition, coagulation, complement activation, cell adhesion and intracellular signaling pathway for the activation of cellular defenses.

Amino Acid Sequence↗

Comparative analysis of BAC and whole genome shotgun sequences from an Anopheles gambiae region related to Plasmodium encapsulation.

The only natural mechanism of malaria transmission in sub-Saharan Africa is the mosquito, generally Anopheles gambiae. Blocking malaria parasite transmission by stopping the development of Plasmodium in the insect vector would provide a useful alternative to the current methods of malaria control. Toward this end, it is important to understand the molecular basis of the malaria parasite refractory phenotype in An. gambiae mosquito strains. We have selected and sequenced six bacterial artificial chromosome (BAC) clones from the Pen-1 region that is the major quantitative trait locus involved in Plasmodium encapsulation. The sequence and the annotation of five overlapping BAC clones plus one adjacent, but not contiguous clone, totaling 585kb of genomic sequence from the centromeric end of the Pen-1 region of the PEST strain were compared to that of the genome sequence of the same strain produced by the whole genome shotgun technique. This project identified 23 putative mosquito genes plus putative copies of the retrotransposable elements BEL12 and TRANSIBN1_AG in the six BAC clones. Nineteen of the predicted genes are most similar to their Drosophila melanogaster homologs while one is more closely related to vertebrate genes. Comparison of these new BAC sequences plus previously published BAC sequences to the cognate region of the assembled genome sequence identified three retrotransposons present in one sequence version but not the other. One of these elements, Indy, has not been previously described. These observations provide evidence for the recent active transposition of these elements and demonstrate the plasticity of the Anopheles genome. The BAC sequences strongly support the public whole genome shotgun assembly and automatic annotation while also demonstrating the benefit of complementary genome sequences and of human curation. Importantly, the data demonstrate the differences in the genome sequence of an individual mosquito compared to that of a hypothetical, average genome sequence generated by whole genome shotgun assembly.

Amino Acid Sequence↗

Towards a two-dimensional proteome map of Mycoplasma pneumoniae.

A Proteome map of the bacterium Mycoplasma pneumoniae was constructed using two-dimensional (2-D) gel electrophoresis in combination with mass spectrometry (MS). M. pneumoniae is a human pathogen with a known genome sequence of 816 kbp coding for only 688 open reading frames, and is therefore an ideal model system to explore the scope and limits of the current technology. The soluble protein content of this bacterium grown under standard laboratory conditions was separated by 1-D or 2-D gel electrophoresis applying various pH gradients, different acrylamide concentrations and buffer systems. Proteins were identified using liquid chromatography-electrospray ionization ion trap and matrix-assisted laser desorption/ionization-MS. Mass spectrometric protein identification was supported and controlled using N-terminal sequencing and immunological methods. So far, proteins from about 350 spots were characterized with MS by determining the molecular weights and partial sequences of their tryptic peptides. Comparing these experimental data with the DNA sequence-derived predictions it was possible to assign these 350 proteins to 224 genes. The importance of proteomics for genome analysis was shown by the identification of four proteins, not annotated in the original publication. Although the proteome map is still incomplete, it is already a useful reference for comparative analyses of M. pneumoniae cells grown under modified conditions.

Acrylic Resins↗

Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.

MOTIVATION: Computer-based selection of entries from sequence databases with respect to a related functional description, e.g. with respect to a common cellular localization or contributing to the same phenotypic function, is a difficult task. Automatic semantic analysis of annotations is not only hampered by incomplete functional assignments. A major problem is that annotations are written in a rich, non-formalized language and are meant for reading by a human expert. This person can extract from the text considerably more information than is immediately apparent due to his extended biological background knowledge and logical reasoning. APPROACH: A technique of automated annotation evaluation based on a combination of lexical analysis and the usage of biological rule libraries has been developed. The proposed algorithm generates new functional descriptors from the annotation of a given entry using the semantic units of the annotation as prepositions for implications executed in accordance with the rule library. RESULTS: The prototype of a software system, the Meta_A(nnotator) program, is described and the results of its application to sequence attribute assignment and sequence selection problems, such as cellular localization and sequence domain annotation of SWISS-PROT entries, are presented. The current software version assigns useful subcellular localization qualifiers to approximately 88% of all SWISS-PROT entries. As shown by demonstrative examples, the combination of sequence and annotation analysis is a powerful approach for the detection of mutual annotation/sequence inconsistencies. AVAILABILITY: Results for the cellular localization assignment can be viewed at the URL http://www.bork. embl-heidelberg.de/CELL_LOC/CELL_LOC.html.

Algorithms↗

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis↗

Reannotation of Shewanella oneidensis genome.

As more and more complete bacterial genome sequences become available, the genome annotation of previously sequenced genomes may become quickly outdated. This is primarily due to the discovery and functional characterization of new genes. We have reannotated the recently published genome of Shewanella oneidensis with the following results: 51 new genes have been identified, and functional annotation has been added to the 97 genes, including 15 new and 82 existing ones with previously unassigned function. The identification of new genes was achieved by predicting the protein coding regions using the HMM-based program GeneMark.hmm. Subsequent comparison of the predicted gene products to the non-redundant protein database using BLAST and the COG (Clusters of Orthologous Groups) database using COGNITOR provided for the functional annotation.

Algorithms↗

Molecular cloning and functional expression of the first two specific insect myosuppressin receptors.

The Drosophila Genome Project database contains the sequences of two genes, CG8985 and CG13803, which are predicted to code for G protein-coupled receptors. We cloned the cDNAs corresponding to these genes and found that their gene structures had not been correctly annotated. We subsequently expressed the coding regions of the two corrected receptor genes in Chinese hamster ovary cells and found that each of them coded for a receptor that could be activated by low concentrations of Drosophila myosuppressin (EC50,4 x 10(-8) M). The insect myosuppressins are decapeptides that generally inhibit insect visceral muscles. Other tested Drosophila neuropeptides did not activate the two receptors. In addition to the two Drosophila myosuppressin receptors, we identified a sequence in the genomic database from the malaria mosquito Anopheles gambiae that also very likely codes for a myosuppressin receptor. To our knowledge, this paper is the first report on the molecular identification of specific insect myosuppressin receptors.

Amino Acid Sequence↗

The mouse genome: experimental examination of gene predictions and transcriptional start sites.

The completion of the mouse and other mammalian genome sequences will provide necessary, but not sufficient, knowledge for an understanding of much of mouse biology at the molecular level. As a requisite next step in this process, the genes in mouse and their structure must be elucidated. In particular, knowledge of the transcriptional start site of these genes will be necessary for further study of their regulatory regions. To assess the current state of mouse genome annotation to support this activity, we identified several hundred gene predictions in mouse with varying levels of supporting evidence and tested them using RACE-PCR. Modifications were made to the procedure allowing pooling of RNA samples, resulting in a scaleable procedure. The results illustrate potential errors or omissions in the current 5' end annotations in 58% of the genes detected. In testing experimentally unsupported gene predictions, we were able to identify 58 that are not usually annotated as genes but produced spliced transcripts (approximately 25% success rate). In addition, in many genes we were able to detect novel exons not predicted by any gene prediction algorithms. In 19.8% of the genes detected in this study, multiple transcript species were observed. These data show an urgent need to provide direct experimental validation of gene annotations. Moreover, these results show that direct validation using RACE-PCR can be an important component of genome-wide validation. This approach can be a useful tool in the ongoing efforts to increase the quality of gene annotations, especially transcriptional start sites, in complex genomes.

Animals↗

Identification and characterization of an interleukin-15 homologue from Tetraodon nigroviridis.

Interleukin-15 (IL-15) plays an important role in adaptive immune systems in vertebrates with similar bioactivities to interleukin-2 (IL-2). Here we report molecular cloning, sequence analysis and distribution of an IL-15 homologue from a pufferfish (Tetraodon nigroviridis). It is located within a 3,088 bp genomic fragment, transcribed into a 1,056 bp mRNA including 158 bp 5'UTR (untranslated region), 519 bp ORF (open reading frame) and 379 bp 3'UTR. T. nigroviridis IL-15 is constitutively detectable in tissues and organs selected. Levels of transcripts were observed after various stimulations. Gene organization is similar to mammals and birds, and a high degree of conservation of chromosome synteny exists between them. Systematic genomics search against Takifugu rubripes genome supports our conclusions. The T. nigroviridis IL-15 precursor with 172aa (amino acids) contains a putative 53aa signal peptide, while the mature peptide has a calculated molecular mass of 13.36 kDa and a theoretical pI of 4.67. The protein sequence shares 13.3-62.1% identity with reported IL-15s. Phylogenetic analysis grouped Tetraodon with other fish on a separated branch, excluded from mammalian and avian IL-15s. In addition, our analysis on another annotated T. nigroviridis IL-15 demonstrated that it may be a paralogue of IL-15. To differentiate it from the known IL-15s, we described it as IL-15x.

Amino Acid Sequence↗

probeBase: an online resource for rRNA-targeted oligonucleotide probes.

Ribosomal RNA-(rRNA)-targeted oligonucleotide probes are widely used for culture-independent identification of microorganisms in environmental and clinical samples. ProbeBase is a comprehensive database containing more than 700 published rRNA-targeted oligonucleotide probe sequences (status August 2002) with supporting bibliographic and biological annotation that can be accessed through the internet at http://www.probebase.net. Each oligonucleotide probe entry contains information on target organisms, target molecule (small- or large-subunit rRNA) and position, G+C content, predicted melting temperature, molecular weight, necessity of competitor probes, and the reference that originally described the oligonucleotide probe, including a link to the respective abstract at PubMed. In addition, probes successfully used for fluorescence in situ hybridization (FISH) are highlighted and the recommended hybridization conditions are listed. ProbeBase also offers difference alignments for 16S rRNA-targeted probes by using the probe match tool of the ARB software and the latest small-subunit rRNA ARB database (release June 2002). The option to directly submit probe sequences to the probe match tool of the Ribosomal Database Project II (RDP-II) further allows one to extract supplementary information on probe specificities. The two main features of probeBase, 'search probeBase' and 'find probe set', help researchers to find suitable, published oligonucleotide probes for microorganisms of interest or for rRNA gene sequences submitted by the user. Furthermore, the 'search target site' option provides guidance for the development of new FISH probes.

Base Sequence↗

The genome and proteome of coliphage T1.

The genome of enterobacterial phage T1 has been sequenced, revealing that its 50.7-kb terminally redundant, circularly permuted sequence contains 48,836 bp of nonredundant nucleotides. Seventy-seven open reading frames (ORFs) were identified, with a high percentage of small genes located at the termini of the genomes displaying no homology to existing phage or prophage proteins. Of the genes showing homologs (47%), we identified those involved in host DNA degradation (three endonucleases) and T1 replication (DNA helicase, primase, and single-stranded DNA-binding proteins) and recombination (RecE and Erf homologs). While the tail genes showed homology to those from temperate coliphage N15, the capsid biosynthetic genes were unique. Phage proteins were resolved by 2D gel electrophoresis, and mass spectrometry was used to identify several of the spots including the major head, portal, and tail proteins, thus verifying the annotation.

Amino Acid Sequence↗

Bioinformatic analysis of an unusual gene-enzyme relationship in the arginine biosynthetic pathway among marine gamma proteobacteria: implications concerning the formation of N-acetylated intermediates in prokaryotes.

BACKGROUND: The N-acetylation of L-glutamate is regarded as a universal metabolic strategy to commit glutamate towards arginine biosynthesis. Until recently, this reaction was thought to be catalyzed by either of two enzymes: (i) the classical N-acetylglutamate synthase (NAGS, gene argA) first characterized in Escherichia coli and Pseudomonas aeruginosa several decades ago and also present in vertebrates, or (ii) the bifunctional version of ornithine acetyltransferase (OAT, gene argJ) present in Bacteria, Archaea and many Eukaryotes. This paper focuses on a new and surprising aspect of glutamate acetylation. We recently showed that in Moritella abyssi and M. profunda, two marine gamma proteobacteria, the gene for the last enzyme in arginine biosynthesis (argH) is fused to a short sequence that corresponds to the C-terminal, N-acetyltransferase-encoding domain of NAGS and is able to complement an argA mutant of E. coli. Very recently, other authors identified in Mycobacterium tuberculosis an independent gene corresponding to this short C-terminal domain and coding for a new type of NAGS. We have investigated the two prokaryotic Domains for patterns of gene-enzyme relationships in the first committed step of arginine biosynthesis. RESULTS: The argH-A fusion, designated argH(A), and discovered in Moritella was found to be present in (and confined to) marine gamma proteobacteria of the Alteromonas- and Vibrio-like group. Most of them have a classical NAGS with the exception of Idiomarina loihiensis and Pseudoalteromonas haloplanktis which nevertheless can grow in the absence of arginine and therefore appear to rely on the arg(A) sequence for arginine biosynthesis. Screening prokaryotic genomes for virtual argH-X 'fusions' where X stands for a homologue of arg(A), we retrieved a large number of Bacteria and several Archaea, all of them devoid of a classical NAGS. In the case of Thermus thermophilus and Deinococcus radiodurans, the arg(A)-like sequence clusters with argH in an operon-like fashion. In this group of sequences, we find the short novel NAGS of the type identified in M. tuberculosis. Among these organisms, at least Thermus, Mycobacterium and Streptomyces species appear to rely on this short NAGS version for arginine biosynthesis. CONCLUSION: The gene-enzyme relationship for the first committed step of arginine biosynthesis should now be considered in a new perspective. In addition to bifunctional OAT, nature appears to implement at least three alternatives for the acetylation of glutamate. It is possible to propose evolutionary relationships between them starting from the same ancestral N-acetyltransferase domain. In M. tuberculosis and many other bacteria, this domain evolved as an independent enzyme, whereas it fused either with a carbamate kinase fold to give the classical NAGS (as in E. coli) or with argH as in marine gamma proteobacteria. Moreover, there is an urgent need to clarify the current nomenclature since the same gene name argA has been used to designate structurally different entities. Clarifying the confusion would help to prevent erroneous genomic annotation.

Acetylation↗

Systematic study of sequence motifs for RNA trans splicing in Trypanosoma brucei.

mRNA maturation in Trypanosoma brucei depends upon trans splicing, and variations in trans-splicing efficiency could be an important step in controlling the levels of individual mRNAs. RNA splicing requires specific sequence elements, including conserved 5' splice sites, branch points, pyrimidine-rich regions [poly(Y) tracts], 3' splice sites (3'SS), and sometimes enhancer elements. To analyze sequence requirements for efficient trans splicing in the poly(Y) tract and around the 3'SS, we constructed a luciferase-beta-galactosidase double-reporter system. By testing approximately 90 sequences, we demonstrated that the optimum poly(Y) tract length is approximately 25 nucleotides. Interspersing a purely uridine-containing poly(Y) tract with cytidine resulted in increased trans-splicing efficiency, whereas purines led to a large decrease. The position of the poly(Y) tract relative to the 3'SS is important, and an AC dinucleotide at positions -3 and -4 can lead to a 20-fold decrease in trans splicing. However, efficient trans splicing can be restored by inserting a second AG dinucleotide downstream, which does not function as a splice site but may aid in recruitment of the splicing machinery. These findings should assist in the development of improved algorithms for computationally identifying a 3'SS and help to discriminate noncoding open reading frames from true genes in current efforts to annotate the T. brucei genome.

Animals↗

POLYVIEW: a flexible visualization tool for structural and functional annotations of proteins.

UNLABELLED: The POLYVIEW visualization server can be used to generate protein sequence annotations, including secondary structures, relative solvent accessibilities, functional motifs and polymorphic sites. Two-dimensional graphical representations in a customizable format may be generated for both known protein structures and predictions obtained using protein structure prediction servers. POLYVIEW may be used for automated generation of pictures with structural and functional annotations for publications and proteomic on-line resources. AVAILABILITY: http://polyview.cchmc.org.

Algorithms↗