Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Systematic study of sequence motifs for RNA trans splicing in Trypanosoma brucei.

mRNA maturation in Trypanosoma brucei depends upon trans splicing, and variations in trans-splicing efficiency could be an important step in controlling the levels of individual mRNAs. RNA splicing requires specific sequence elements, including conserved 5' splice sites, branch points, pyrimidine-rich regions [poly(Y) tracts], 3' splice sites (3'SS), and sometimes enhancer elements. To analyze sequence requirements for efficient trans splicing in the poly(Y) tract and around the 3'SS, we constructed a luciferase-beta-galactosidase double-reporter system. By testing approximately 90 sequences, we demonstrated that the optimum poly(Y) tract length is approximately 25 nucleotides. Interspersing a purely uridine-containing poly(Y) tract with cytidine resulted in increased trans-splicing efficiency, whereas purines led to a large decrease. The position of the poly(Y) tract relative to the 3'SS is important, and an AC dinucleotide at positions -3 and -4 can lead to a 20-fold decrease in trans splicing. However, efficient trans splicing can be restored by inserting a second AG dinucleotide downstream, which does not function as a splice site but may aid in recruitment of the splicing machinery. These findings should assist in the development of improved algorithms for computationally identifying a 3'SS and help to discriminate noncoding open reading frames from true genes in current efforts to annotate the T. brucei genome.

Animals↗

POLYVIEW: a flexible visualization tool for structural and functional annotations of proteins.

UNLABELLED: The POLYVIEW visualization server can be used to generate protein sequence annotations, including secondary structures, relative solvent accessibilities, functional motifs and polymorphic sites. Two-dimensional graphical representations in a customizable format may be generated for both known protein structures and predictions obtained using protein structure prediction servers. POLYVIEW may be used for automated generation of pictures with structural and functional annotations for publications and proteomic on-line resources. AVAILABILITY: http://polyview.cchmc.org.

Algorithms↗

Molecular characterization of the developmental gene in eyes: through data-mining on integrated transcriptome databases.

OBJECTIVES: Our aim was to utilize publicly available and proprietary sources to discover candidate genes important for ocular development. DESIGN AND METHODS: The collated information on our 5092 non-redundant clusters was grouped and functional annotation was conducted using gene ontology (FatiGO) for categorizing them with respect to molecular function. The web-based viewer technological platform (H-InvDB) was employed for transcription analyses of in-house high quality fetal eye Expressed Sequence Tags (ESTs). Eye-specific ESTs were also analyzed across species by using EMBEST. RESULTS: According to adult eye cDNA libraries, nucleic acid binding and cell structure/cytoskeletal protein genes were the most abundant among the ESTs of fetal eyes. Using cDNA assembly in H-InvDB, 20 (80%) of the 25 most commonly expressed genes in the human eye are also expressed in extraocular tissues. The crystalline gamma S gene is highly expressed in the eye, but not in other tissues. We used EMBEST to compare human fetal eye and octopus eye ESTs and the expression similarity was low (1.6%). This indicated that our fetal eye library contains genes necessary for the developmental process and biological function of the eye, which may not be expressed in the fully developed octopus eyes. The human fetal eye cDNA library also contained highly abundant eye tissue genes, including alphaA-crystallin, eukaryotic translation elongation factor 1 alpha 1 (EEF1A1), bestrophin (VMD2), cystatin C, and transforming growth factor, beta-induced (BIGH3). CONCLUSIONS: Our annotated EST set provides a valuable resource for gene discovery and functional genomic analysis. This display will help to appreciate the strengths and weaknesses of the different technological platforms, so that in future studies the maximum amount of beneficial information can be derived from the appropriate use of each method.

Animals↗

Similarities and differences in genome-wide expression data of six organisms.

Comparing genomic properties of different organisms is of fundamental importance in the study of biological and evolutionary principles. Although differences among organisms are often attributed to differential gene expression, genome-wide comparative analysis thus far has been based primarily on genomic sequence information. We present a comparative study of large datasets of expression profiles from six evolutionarily distant organisms: S. cerevisiae, C. elegans, E. coli, A. thaliana, D. melanogaster, and H. sapiens. We use genomic sequence information to connect these data and compare global and modular properties of the transcription programs. Linking genes whose expression profiles are similar, we find that for all organisms the connectivity distribution follows a power-law, highly connected genes tend to be essential and conserved, and the expression program is highly modular. We reveal the modular structure by decomposing each set of expression data into coexpressed modules. Functionally related sets of genes are frequently coexpressed in multiple organisms. Yet their relative importance to the transcription program and their regulatory relationships vary among organisms. Our results demonstrate the potential of combining sequence and expression data for improving functional gene annotation and expanding our understanding of how gene expression and diversity evolved.

Animals↗

Clinical and taxonomic status of pathogenic nonpigmented or late-pigmenting rapidly growing mycobacteria.

The history, taxonomy, geographic distribution, clinical disease, and therapy of the pathogenic nonpigmented or late-pigmenting rapidly growing mycobacteria (RGM) are reviewed. Community-acquired disease and health care-associated disease are highlighted for each species. The latter grouping includes health care-associated outbreaks and pseudo-outbreaks as well as sporadic disease cases. Treatment recommendations for each species and type of disease are also described. Special emphasis is on the Mycobacterium fortuitum group, including M. fortuitum, M. peregrinum, and the unnamed third biovariant complex with its recent taxonomic changes and newly recognized species (including M. septicum, M. mageritense, and proposed species M. houstonense and M. bonickei). The clinical and taxonomic status of M. chelonae, M. abscessus, and M. mucogenicum is also detailed, along with that of the closely related new species, M. immunogenum. Additionally, newly recognized species, M. wolinskyi and M. goodii, as well as M. smegmatis sensu stricto, are included in a discussion of the M. smegmatis group. Laboratory diagnosis of RGM using phenotypic methods such as biochemical testing and high-performance liquid chromatography and molecular methods of diagnosis are also discussed. The latter includes PCR-restriction fragment length polymorphism analysis, hybridization, ribotyping, and sequence analysis. Susceptibility testing and antibiotic susceptibility patterns of the RGM are also annotated, along with the current recommendations from the National Committee for Clinical Laboratory Standards (NCCLS) for mycobacterial susceptibility testing.

Anti-Bacterial Agents↗

Molecular fossils in the human genome: identification and analysis of the pseudogenes in chromosomes 21 and 22.

We have developed an initial approach for annotating and surveying pseudogenes in the human genome. We search human genomic DNA for regions that are similar to known protein sequences and contain obvious disablements (i.e., mid-sequence stop codons or frameshifts), while ensuring minimal overlap with annotations of known genes. Pseudogenes can be divided into "processed" and "nonprocessed"; the former are reverse transcribed from mRNA (and therefore have no intron structure), whereas the latter presumably arise from genomic duplications. We annotate putative processed pseudogenes based on whether there is a continuous span of homology that is >70% of the length of the closest matching human protein (i.e., with introns removed), or whether there is evidence of polyadenylation. We have applied our approach to chromosomes 21 and 22, the first parts of the human genome completely sequenced, finding 190 new pseudogene annotations beyond the 264 reported by the sequencing centers. In total, on chromosomes 21 and 22, there are 189 processed pseudogenes, 195 nonprocessed pseudogenes, and, additionally, 70 pseudogenic immunoglobulin gene segments. (Detailed assignments are available at http://bioinfo.mbb.yale.edu/genome/pseudogene or http://genecensus.org/pseudogene.) By extrapolation, we predict that there could be up to approximately 20,000 pseudogenes in the whole human genome, with a little more than half of them processed. We have determined the main populations and clusters of pseudogenes on chromosomes 21 and 22. There are notable excesses of pseudogenes relative to genes near the centromeres of both chromosomes, indicating the existence of pseudogenic "hot-spots" in the genome. We have looked at the distribution of InterPro families and Gene Ontology (GO) functional categories in our pseudogenes. Overall, the families in both processed and nonprocessed pseudogene populations occur according to a similar power-law distribution as that found for the occurrence of gene families, with a few big families and many small ones. The processed population is, in particular, enriched in highly expressed ribosomal-protein sequences (approximately 20%), which appear fairly evenly distributed across the chromosomes. We compared processed pseudogenes of different evolutionary ages, observing a high degree of similarity between "ancient" and "modern" subpopulations. This may be attributable to the consistently high expression of ribosomal proteins over evolutionary time. Finally, we find that chromosome 22 pseudogene population is dominated by immunoglobulin segments, which have a greater rate of disablement per amino acid than the other pseudogene populations and are also substantially more diverged.

Chromosome Mapping↗

Expression profiling of human idiopathic dilated cardiomyopathy.

OBJECTIVE: To investigate the global changes accompanying human dilated cardiomyopathy (DCM) we performed a large-scale expression screen using myocardial biopsies from a group of DCM patients with moderate heart failure. By hierarchical clustering and functional annotation of the deregulated genes we examined extensive changes in the cellular and molecular processes associated to DCM. METHODS: The expression profiles were obtained using a whole genome covering library (UniGene RZPD1) comprising 30336 cDNA clones and amplified RNA from myocardiac biopsies from 10 DCM patients in comparison to tissue samples from four non-failing, healthy donors. RESULTS: By setting stringent selection criteria 364 differentially expressed, sequence-verified non-redundant transcripts were identified with a false discovery rate of <0.001. Numerous genes and ESTs were identified representing previously recognised, as well as novel DCM-associated transcripts. Many of them were found to be upregulated and involved in cardiomyocyte energetics, muscle contraction or signalling. Two hundred and twenty-two deregulated transcripts were functionally annotated and hierarchically clustered providing an insight into the pathophysiology of DCM. Data was validated using the MLP-deficient mouse, in which several differentially expressed transcripts identified in the human DCM biopsies could be confirmed. CONCLUSIONS: We report the first genome-wide expression profile analysis using cardiac biopsies from DCM patients at various stages of the disease. Although there is a diversity of links between the cytoskeleton and the initiation of DCM, we speculate that genes implicated in intracellular signalling and in muscle contraction are associated with early stages of the disease. Altogether this study represents the most comprehensive and inclusive molecular portrait of human cardiomyopathy to date.

Adult↗

Isolation and characterization of a thermostable RNA ligase 1 from a Thermus scotoductus bacteriophage TS2126 with good single-stranded DNA ligation properties.

We have recently sequenced the genome of a novel thermophilic bacteriophage designated as TS2126 that infects the thermophilic eubacterium Thermus scotoductus. One of the annotated open reading frames (ORFs) shows homology to T4 RNA ligase 1, an enzyme of great importance in molecular biology, owing to its ability to ligate single-stranded nucleic acids. The ORF was cloned, and recombinant protein was expressed, purified and characterized. The recombinant enzyme ligates single-stranded nucleic acids in an ATP-dependent manner and is moderately thermostable. The recombinant enzyme exhibits extremely high activity and high ligation efficiency. It can be used for various molecular biology applications including RNA ligase-mediated rapid amplification of cDNA ends (RLM-RACE). The TS2126 RNA ligase catalyzed both inter- and intra-molecular single-stranded DNA ligation to >50% completion in a matter of hours at an elevated temperature, although favoring intra-molecular ligation on RNA and single-stranded DNA substrates. The properties of TS2126 RNA ligase 1 makes it very attractive for processes like adaptor ligation, and single-stranded solid phase gene synthesis.

Amino Acid Sequence↗

Predicting protein function from sequence and structural data.

When a protein's function cannot be experimentally determined, it can often be inferred from sequence similarity. Should this process fail, analysis of the protein structure can provide functional clues or confirm tentative functional assignments inferred from the sequence. Many structure-based approaches exist (e.g. fold similarity, three-dimensional templates), but as no single method can be expected to be successful in all cases, a more prudent approach involves combining multiple methods. Several automated servers that integrate evidence from multiple sources have been released this year and particular improvements have been seen with methods utilizing the Gene Ontology functional annotation schema.

Binding Sites↗

GPMAW--a software tool for analyzing proteins and peptides.

General Protein/Mass Analysis for Windows (GPMAW) is a valuable piece of software for any molecular biologist, biochemist or mass spectrometrist wishing to analyze protein or peptide sequences. All steps from the acquisition of protein sequence from a built-in web interface, to proteolytic digests, theoretical peptide fragmentation, detailed annotation of sequences and secondary structure prediction, can be performed rapidly and intuitively without first having to spend days reading manuals.

Amino Acid Sequence↗

Efficient recognition of protein fold at low sequence identity by conservative application of Psi-BLAST: validation.

A substantial fraction of protein sequences derived from genomic analyses is currently classified as representing 'hypothetical proteins of unknown function'. In part, this reflects the limitations of methods for comparison of sequences with very low identity. We evaluated the effectiveness of a Psi-BLAST search strategy to identify proteins of similar fold at low sequence identity. Psi-BLAST searches for structurally characterized low-sequence-identity matches were carried out on a set of over 300 proteins of known structure. Searches were conducted in NCBI's non-redundant database and were limited to three rounds. Some 614 potential homologs with 25% or lower sequence identity to 166 members of the search set were obtained. Disregarding the expect value, level of sequence identity and span of alignment, correspondence of fold between the target and potential homolog was found in more than 95% of the Psi-BLAST matches. Restrictions on expect value or span of alignment improved the false positive rate at the expense of eliminating many true homologs. Approximately three-quarters of the putative homologs obtained by three rounds of Psi-BLAST revealed no significant sequence similarity to the target protein upon direct sequence comparison by BLAST, and therefore could not be found by a conventional search. Although three rounds of Psi-BLAST identified many more homologs than a standard BLAST search, most homologs were undetected. It appears that more than 80% of all homologs to a target protein may be characterized by a lack of significant sequence similarity. We suggest that conservative use of Psi-BLAST has the potential to propose experimentally testable functions for the majority of proteins currently annotated as 'hypothetical proteins of unknown function'.

Algorithms↗

An automated annotation tool for genomic DNA sequences using GeneScan and BLAST.

Genomic sequence data are often available well before the annotated sequence is published. We present a method for analysis of genomic DNA to identify coding sequences using the GeneScan algorithm and characterize these resultant sequences by BLAST. The routines are used to develop a system for automated annotation of genome DNA sequences.

Algorithms↗

SISYPHUS--structural alignments for proteins with non-trivial relationships.

With the increasing amount of structural data, the number of homologous protein structures bearing topological irregularities is steadily growing. These include proteins with circular permutations, segment-swapping, context-dependent folding or chameleon sequences that can adopt alternative secondary structures. Their non-trivial structural relationships are readily identified during expert analysis but their automatic identification using the existing computational tools still remains difficult or impossible. Such non-trivial cases of protein relationships are known to pose a problem to multiple alignment algorithms and to impede comparative modeling studies. They support a new emerging concept of evolutionary changeable protein fold, which creates practical difficulties for the hierarchical classifications of protein structures.To facilitate the understanding of, and to provide a comprehensive annotation of proteins with such non-trivial structural relationships we have created SISYPHUS ([Sigmaomeganuphiomicronzeta]--in Greek crafty), a compendium to the SCOP database. The SISYPHUS database contains a collection of manually curated structural alignments and their inter-relationships. The multiple alignments are constructed for protein structural regions that range from oligomeric biological units, or individual domains to fragments of different size. The SISYPHUS multiple alignments are displayed with SPICE, a browser that provides an integrated view of protein sequences, structures and their annotations. The database is available from http://sisyphus.mrc-cpe.cam.ac.uk.

Databases, Protein↗

Automated DNA chip annotation tables at IFOM: the importance of synchronisation and cross-referencing of sequence databases.

The increasing popularity of DNA chip technology for the study of gene expression is producing, for each experiment, a sizable quantity of numerical data to analyse and an accompanying large number of gene identifiers that should be associated with the relevant biological annotation. We describe here a website at IFOM (FIRC Institute of Molecular Oncology) where we release regularly updated annotation tables for the most used Affymetrix oligonucleotide DNA chips and for the whole Research Genetics 46K clone collection for cDNA arrays. These tables are synchronised with every new release of the mouse and human UniGene databases (NCBI; National Center for Biotechnology Information), allowing fast and easy preliminary annotation of DNA array experiments. We also report some comparative evidence about the importance of biological database synchronisation and cross-references in the process of generating annotation tables for DNA chips.

Abstracting and Indexing↗

The genome sequence of an anaerobic aromatic-degrading denitrifying bacterium, strain EbN1.

Recent research on microbial degradation of aromatic and other refractory compounds in anoxic waters and soils has revealed that nitrate-reducing bacteria belonging to the Betaproteobacteria contribute substantially to this process. Here we present the first complete genome of a metabolically versatile representative, strain EbN1, which metabolizes various aromatic compounds, including hydrocarbons. A circular chromosome (4.3 Mb) and two plasmids (0.21 and 0.22 Mb) encode 4603 predicted proteins. Ten anaerobic and four aerobic aromatic degradation pathways were recognized, with the encoding genes mostly forming clusters. The presence of paralogous gene clusters (e.g., for anaerobic phenylacetate oxidation), high sequence similarities to orthologs from other strains (e.g., for anaerobic phenol metabolism) and frequent mobile genetic elements (e.g., more than 200 genes for transposases) suggest high genome plasticity and extensive lateral gene transfer during metabolic evolution of strain EbN1. Metabolic versatility is also reflected by the presence of multiple respiratory complexes. A large number of regulators, including more than 30 two-component and several FNR-type regulators, indicate a finely tuned regulatory network able to respond to the fluctuating availability of organic substrates and electron acceptors in the environment. The absence of genes required for nitrogen fixation and specific interaction with plants separates strain EbN1 ecophysiologically from the closely related nitrogen-fixing plant symbionts of the Azoarcus cluster. Supplementary material on sequence and annotation are provided at the Web page http://www.micro-genomes.mpg.de/ebn1/.

Adaptation, Physiological↗

Taverna: a tool for building and running workflows of services.

Taverna is an application that eases the use and integration of the growing number of molecular biology tools and databases available on the web, especially web services. It allows bioinformaticians to construct workflows or pipelines of services to perform a range of different analyses, such as sequence analysis and genome annotation. These high-level workflows can integrate many different resources into a single analysis. Taverna is available freely under the terms of the GNU Lesser General Public License (LGPL) from http://taverna.sourceforge.net/.

Computational Biology↗

The Comparative Toxicogenomics Database (CTD).

The Mount Desert Island Biological Laboratory in Salsbury Cove, Maine, USA, is developing the Comparative Toxicogenomics Database (CTD), a community-supported genomic resource devoted to genes and proteins of human toxicologic significance. CTD will be the first publicly available database to a) provide annotated associations among genes, proteins, references, and toxic agents, with a focus on annotating data from aquatic and mammalian organisms; b) include nucleotide and protein sequences from diverse species; c) offer a range of analysis tools for customized comparative studies; and d) provide information to investigators on available molecular reagents. This combination of features will facilitate cross-species comparisons of toxicologically significant genes and proteins. These comparisons will promote understanding of molecular evolution, the significance of conserved sequences, the genetic basis of variable sensitivity to environmental agents, and the complex interactions between the environment and human health. CTD is currently under development, and the planned scope and functions of the database are described herein. The intent of this report is to invite community participation in the development of CTD to ensure that it will be a valuable resource for environmental health, molecular biology, and toxicology research.

Animals↗

The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.

For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha+beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the functions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the database overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc++ +.

Enzymes↗