Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Role of context in the relationship between form and function: structural plasticity of some PROSITE patterns.

True positive hits of PROSITE sequence pattern are expected to have a characteristic three-dimensional structure. The combined sequence-structure attributes of PROSITE patterns can be used for function prediction of an uncharacterized protein with known primary and 3D structure, a situation that might arise in structural genomics projects. We have found specific examples of true hits of PROSITE patterns displaying structural plasticity by assuming significantly different local conformation, depending upon the context. Our work highlights the importance of taking into account all the known distinct conformations of PROSITE patterns, while creating a sensitive 3D template for the pattern, for use in functional annotation.

Amino Acid Motifs↗

DNAView: a quality assessment tool for the visualization of large sequenced regions.

This communication describes DNAView, a graphical tool for the visualization and printing of large nucleic acid sequences. DNAView uses color coding to compactly display genomic segments of up to 100 kb on a single printed page. The specific color schemes integrated into DNAView can highlight 'local aggregate' properties of large segments of DNA. We have also incorporated a confidence expression for the assigned sequence. This is represented by base color intensity that is proportional to the number of times that base was sequenced. Areas of interest, such as exons, introns, repetitive elements and splice sites, can be emphasized using overlays. The colored image can be saved in a standard TIFF image file format that may be imported and annotated by other application software.

Base Sequence↗

Computational approaches to identify leucine zippers.

The leucine zipper is a dimerization domain occurring mostly in regulatory and thus in many oncogenic proteins. The leucine repeat in the sequence has been traditionally used for identification, however with poor reliability. The coiled coil structure of a leucine zipper is required for dimerization and can be predicted with reasonable accuracy by existing algorithms. We exploit this fact for identification of leucine zippers from sequence alone. We present a program, 2ZIP, which combines a standard coiled coil prediction algorithm with an approximate search for the characteristic leucine repeat. No further information from homologues is required for prediction. This approach improves significantly over existing methods, especially in that the coiled coil prediction turns out to be highly informative and avoids large numbers of false positives. Many problems in predicting zippers or assessing prediction results stem from wrong sequence annotations in the database.

Algorithms↗

Annotation and evolutionary relationships of a small regulatory RNA gene micF and its target ompF in Yersinia species.

BACKGROUND: micF RNA, a small regulatory RNA found in bacteria, post-transcriptionally regulates expression of outer membrane protein F (OmpF) by interaction with the ompF mRNA 5'UTR. Phylogenetic data can be useful for RNA/RNA duplex structure analyses and aid in elucidation of mechanism of regulation. However micF and associated genes, ompF and ompC are difficult to annotate because of either similarities or divergences in nucleotide sequence. We report by using sequences that represent "gene signatures" as probes, e.g., mRNA 5'UTR sequences, closely related genes can be accurately located in genomic sequences. RESULTS: Alignment and search methods using NCBI BLAST programs have been used to identify micF, ompF and ompC in Yersinia pestis and Yersinia enterocolitica. By alignment with DNA sequences from other bacterial species, 5' start sites of genes and upstream transcriptional regulatory sites in promoter regions were predicted. Annotated genes from Yersinia species provide phylogenetic information on the micF regulatory system. High sequence conservation in binding sites of transcriptional regulatory factors are found in the promoter region upstream of micF and conservation in blocks of sequences as well as marked sequence variation is seen in segments of the micF RNA gene. Unexpected large differences in rates of evolution were found between the interacting RNA transcripts, micF RNA and the 5' UTR of the ompF mRNA. micF RNA/ompF mRNA 5' UTR duplex structures were modeled by the mfold program. Functional domains such as RNA/RNA interacting sites appear to display a minimum of evolutionary drift in sequence with the exception of a significant change in Y. enterocolitica micF RNA. CONCLUSIONS: Newly annotated Yersinia micF and ompF genes and the resultant RNA/RNA duplex structures add strong phylogenetic support for a generalized duplex model. The alignment and search approach using 5' UTR signatures may be a model to help define other genes and their start sites when annotated genes are available in well-defined reference organisms.

5' Untranslated Regions↗

Shotgun sample sequence comparisons between mouse and human genomes.

A mixed 'clone-by-clone' and 'whole-genome shotgun' strategy will be used to determine the genomic sequence of the mouse. This method will allow a phase of rapid annotation of the contemporaneous human sequence draft, through whole-genome 'sample sequence comparisons'.

Animals↗

UNITE: a database providing web-based methods for the molecular identification of ectomycorrhizal fungi.

Identification of ectomycorrhizal (ECM) fungi is often achieved through comparisons of ribosomal DNA internal transcribed spacer (ITS) sequences with accessioned sequences deposited in public databases. A major problem encountered is that annotation of the sequences in these databases is not always complete or trustworthy. In order to overcome this deficiency, we report on UNITE, an open-access database. UNITE comprises well annotated fungal ITS sequences from well defined herbarium specimens that include full herbarium reference identification data, collector/source and ecological data. At present UNITE contains 758 ITS sequences from 455 species and 67 genera of ECM fungi. UNITE can be searched by taxon name, via sequence similarity using blastn, and via phylogenetic sequence identification using galaxie. Following implementation, galaxie performs a phylogenetic analysis of the query sequence after alignment either to pre-existing generic alignments, or to matches retrieved from a blast search on the UNITE data. It should be noted that the current version of UNITE is dedicated to the reliable identification of ECM fungi. The UNITE database is accessible through the URL http://unite.zbi.ee

DNA, Ribosomal Spacer↗

The RESID Database of protein structure modifications and the NRL-3D Sequence-Structure Database.

The RESID Database is a comprehensive collection of annotations and structures for protein post-translational modifications including N-terminal, C-terminal and peptide chain cross-link modifications. The RESID Database includes systematic and frequently observed alternate names, Chemical Abstracts Service registry numbers, atomic formulas and weights, enzyme activities, taxonomic range, keywords, literature citations with database cross-references, structural diagrams and molecular models. The NRL-3D Sequence-Structure Database is derived from the three-dimensional structure of proteins deposited with the Research Collaboratory for Structural Bioinformatics Protein Data Bank. The NRL-3D Database includes standardized and frequently observed alternate names, sources, keywords, literature citations, experimental conditions and searchable sequences from model coordinates. These databases are freely accessible through the National Cancer Institute-Frederick Advanced Biomedical Computing Center at these web sites: http://www. ncifcrf.gov/RESID, http://www.ncifcrf.gov/NRL-3D; or at these National Biomedical Research Foundation Protein Information Resource web sites: http://pir.georgetown.edu/pirwww/dbinfo/resid .html, http://pir.georgetown.edu/pirwww/dbinfo/nrl3d .html

Amino Acids↗

Origin and neofunctionalization of a Drosophila paternal effect gene essential for zygote viability.

BACKGROUND: Although evolutionary novelty by gene duplication is well established, the origin and maintenance of essential genes that provide entirely new functions (neofunctionalization) is still largely unknown. Drosophila is a good model for the search of genes that are young enough to allow deciphering the molecular details of their evolutionary history. Recent years have seen increased interest in genes specifically required for male fertility because they often evolve rapidly. A special class of genes affecting male fertility, the paternal effect genes, have also become a focus of study to geneticists and reproductive biologists interested in fertilization and sperm-egg interactions. RESULTS: Using molecular genetics and the annotated Drosophila melanogaster genome, we identified CG14251 as the Drosophila paternal effect gene, ms(3)K81 (K81). This assignment was subsequently confirmed by P-element rescue of K81. A search for orthologous K81 sequences revealed that the distribution of K81 is surprisingly restricted to the 9 species comprising the melanogaster subgroup. Phylogenetic analyses indicate that K81 arose through duplication, most likely retroposition, of a ubiquitously expressed gene before the radiation of the melanogaster subgroup, followed by a period of rapid divergence and acquisition of a critical male germline-specific function. Interestingly, K81 has adopted the expression profile of a flanking gene suggesting that transcriptional coregulation may have been important in the neofunctionalization of K81. CONCLUSION: We present a detailed case history of the origin and evolution of a new essential gene and, in so doing, provide the first molecular identification of a Drosophila paternal effect gene, ms(3)K81 (K81).

Amino Acid Sequence↗

DNA microarrays in parasitology: strengths and limitations.

Genome sequencing efforts have provided a wealth of new biological information that promises to have a major impact on our understanding of parasites. Microarrays provide one of the major high-throughput platforms by which this information can be exploited in the laboratory. Many excellent reviews and technique articles have recently been published on applying microarrays to organisms for which fully annotated genomes are at hand. However, many parasitologists work on organisms whose genomes have been only partially sequenced and where little, if any, annotation is available. The focus of this review is on how to use and apply microarrays to these situations.

Animals↗

Novel candidate targets of beta-catenin/T-cell factor signaling identified by gene expression profiling of ovarian endometrioid adenocarcinomas.

The activity of beta-catenin (beta-cat), a key component of the Wnt signaling pathway, is deregulated in about 40% of ovarian endometrioid adenocarcinomas (OEAs), usually as a result of CTNNB1 gene mutations. The function of beta-cat in neoplastic transformation is dependent on T-cell factor (TCF) transcription factors, but specific genes activated by the interaction of beta-cat with TCFs in OEAs and other cancers with Wnt pathway defects are largely unclear. As a strategy to identify beta-cat/TCF transcriptional targets likely to contribute to OEA pathogenesis, we used oligonucleotide microarrays to compare gene expression in primary OEAs with mutational defects in beta-cat regulation (n = 11) to OEAs with intact regulation of beta-cat activity (n = 17). Both hierarchical clustering and principal component analysis based on global gene expression distinguished beta-cat-defective tumors from those with intact beta-cat regulation. We identified 81 potential beta-cat/TCF targets by selecting genes with at least 2-fold increased expression in beta-cat-defective versus beta-cat regulation-intact tumors and significance in a t test (P < 0.05). Seven of the 81 genes have been previously reported as Wnt/beta-cat pathway targets (i.e., BMP4, CCND1, CD44, FGF9, EPHB3, MMP7, and MSX2). Differential expression of several known and candidate target genes in the OEAs was confirmed. For the candidate target genes CST1 and EDN3, reporter and chromatin immunoprecipitation assays directly implicated beta-cat and TCF in their regulation. Analysis of presumptive regulatory elements in 67 of the 81 candidate genes for which complete genomic sequence data were available revealed an apparent difference in the location and abundance of consensus TCF-binding sites compared with the patterns seen in control genes. Our findings imply that analysis of gene expression profiling data from primary tumor samples annotated with detailed molecular information may be a powerful approach to identify key downstream targets of signaling pathways defective in cancer cells.

Binding Sites↗

Synonymous codon usage and gene function are strongly related in Oryza sativa.

The relationship between codon usage and gene function was investigated while considering a dataset of 2106 nuclear genes of Oryza sativa. The results of standard chi(2) test and F-statistic showed that for every 59 synonymous codons, a strongly significant association with gene functional categories existed in rice, indicating that codon usage was generally coordinated with gene function whether it was at the level of individual amino acids or at the level of nucleotides. However, it could not be directly said that the use of every codons differed significantly between any two functional categories. Notably, there existed large difference both in selection for biased codons or selection intensity among functional categories. Therefore, we identified at least two classes of genes: one group of genes, mainly belonging to the "METABOLISM" category, was tended to use G- and/or C-ending codons while the other was more biased to choose codons ending with A and/or U. The latter group contained genes of various functions, especially those genes classified into the "Nuclear Structure" category. These observations will be more important for molecular genetic engineering and genome functional annotation.

Chromosome Mapping↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43&#x2009;Mb, with a Scaffold N50 length of 19.05&#x2009;Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75&#x2009;Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals↗

Identifying the major proteome components of Haemophilus influenzae type-strain NCTC 8143.

With the completion of the Haemophilus influenzae Rd genomic sequence, we know the identity of most of the theoretical proteins in the proteome of this bacterium. However, the most abundant components of the actual proteome are unknown. Using mass spectrometry and two-dimensional gel electrophoresis (2-DE), we sequenced and analyzed the most abundant proteins observed in the ATCC reference strain of H. influenzae, NCTC 8143 (303 of approximately 400 Coomassie-stained 2-DE spots). To automate the identification of 2-DE spots, we coupled a liquid autosampler to a microcolumn liquid chromatography electrospray ionization tandem mass spectrometer capable of identifying 22 spots per day. From the 303 sequenced spots, we identified 263 unique proteins. Most of the abundant proteins lie in an isoelectric point range of pH 4-7 and a molecular mass range of 10-100 kDa. Of the observed proteins, the most abundant is the outer membrane protein P2. Based on variety and abundance, proteins involved in energy metabolism and macromolecular synthesis are the dominant classes of proteins. Unexpectedly, tryptophanase was identified as a highly abundant protein in the strain NCTC 8143 whose sequence is not present in the genome of the Rd strain. By searching the tandem mass spectra against the translated genomic sequence, we identified several proteins which were not annotated in the genomic sequence. Surprisingly, 22% of the identified 2-DE spots represent isoforms in which gene products with the same primary sequence have different observed pI and M(r), indicating that these proteins are post-translationally processed. Although most proteins' predicted and observed isoelectric points and molecular masses show reasonable concordance, the observed values for several proteins deviate significantly from the predicted values. These anomalies may represent either highly processed proteins or misinterpretations of the genomic sequence. Using the technology developed in this project, the protein expression of other strains of H. influenzae grown under different environmental conditions can be compared to identify differences in their proteomes.

Bacterial Proteins↗

A method for the improvement of threading-based protein models.

A new method for the homology-based modeling of protein three-dimensional structures is proposed and evaluated. The alignment of a query sequence to a structural template produced by threading algorithms usually produces low-resolution molecular models. The proposed method attempts to improve these models. In the first stage, a high-coordination lattice approximation of the query protein fold is built by suitable tracking of the incomplete alignment of the structural template and connection of the alignment gaps. These initial lattice folds are very similar to the structures resulting from standard molecular modeling protocols. Then, a Monte Carlo simulated annealing procedure is used to refine the initial structure. The process is controlled by the model's internal force field and a set of loosely defined restraints that keep the lattice chain in the vicinity of the template conformation. The internal force field consists of several knowledge-based statistical potentials that are enhanced by a proper analysis of multiple sequence alignments. The template restraints are implemented such that the model chain can slide along the template structure or even ignore a substantial fraction of the initial alignment. The resulting lattice models are, in most cases, closer (sometimes much closer) to the target structure than the initial threading-based models. All atom models could easily be built from the lattice chains. The method is illustrated on 12 examples of target/template pairs whose initial threading alignments are of varying quality. Possible applications of the proposed method for use in protein function annotation are briefly discussed.

Amino Acid Sequence↗

The membrane skeleton in Paramecium: Molecular characterization of a novel epiplasmin family and preliminary GFP expression results.

Previous attempts to identify the membrane skeleton of Paramecium cells have revealed a protein pattern that is both complex and specific. The most prominent structural elements, epiplasmic scales, are centered around ciliary units and are closely apposed to the cytoplasmic side of the inner alveolar membrane. We sought to characterize epiplasmic scale proteins (epiplasmins) at the molecular level. PCR approaches enabled the cloning and sequencing of two closely related genes by amplifications of sequences from a macronuclear genomic library. Using these two genes (EPI-1 and EPI-2), we have contributed to the annotation of the Paramecium tetraurelia macronuclear genome and identified 39 additional (paralogous) sequences. Two orthologous sequences were found in the Tetrahymena thermophila genome. Structural analysis of the 43 sequences indicates that the hallmark of this new multigenic family is a 79 aa domain flanked by two Q-, P- and V-rich stretches of sequence that are much more variable in amino-acid composition. Such features clearly distinguish members of the multigenic family from epiplasmic proteins previously sequenced in other ciliates. The expression of Green Fluorescent Protein (GFP)-tagged epiplasmin showed significant labeling of epiplasmic scales as well as oral structures. We expect that the GFP construct described herein will prove to be a useful tool for comparative subcellular localization of different putative epiplasmins in Paramecium.

Amino Acid Sequence↗

Recent advances in gene structure prediction.

De novo gene predictors are programs that predict the exon-intron structures of genes using the sequences of one or more genomes as their only input. In the past two years, dual-genome de novo predictors, which exploit local rates and patterns of mutation inferred from alignments between two genomes, have led to significant improvements in accuracy. Systems that exploit more than two genomes simultaneously have only recently begun to appear and are not yet competitive on practical tasks, but offer the greatest hope for near-term improvements. Dual-genome de novo prediction for compact eukaryotic genomes such as those of Arabidopsis thaliana and Caenorhabditis elegans is already quite accurate. Although mammalian gene prediction lags behind in accuracy, it is yielding ever more useful results. Coupled with significant improvements in pseudogene detection methods, which have eliminated many false positives, we have reached the point where de novo gene predictions are being used as hypotheses to drive experimental annotation via systematic RT-PCR and sequencing.

Animals↗

Schematic representation of residue-based protein context-dependent data: an application to transmembrane proteins.

An algorithmic method for drawing residue-based schematic diagrams of proteins on a 2D page is presented and illustrated. The method allows the creation of rendering engines dedicated to a given family of sequences, or fold. The initial implementation provides an engine that can produce a 2D diagram representing secondary structure for any transmembrane protein sequence. We present the details of the strategy for automating the drawing of these diagrams. The most important part of this strategy is the development of an algorithm for laying out residues of a loop that connects to arbitrary points of a 2D plane. As implemented, this algorithm is suitable for real-time modification of the loop layout. This work is of interest for the representation and analysis of data from (1) protein databases, (2) mutagenesis results, or (3) various kinds of protein context-dependent annotations or data.

Algorithms↗