Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Complete genomic nucleotide sequence and analysis of the temperate bacteriophage VWB.

The entire double-stranded DNA genome of the Streptomyces venezuelae bacteriophage VWB was sequenced and analyzed. Its size is 49,220 bp with an overall molar G + C content of 71.2 mol%. Sixty-one potential open reading frames were identified and annotated using several complementary bioinformatics tools. Clusters of functionally related putative genes were defined, supporting a refined version of the modular theory of phage evolution.

Amino Acid Sequence↗

The role of protein structure in genomics.

The genome projects produce an enormous amount of sequence data that needs to be annotated in terms of molecular structure and biological function. These tasks have triggered additional initiatives like structural genomics. The intention is to determine as many protein structures as possible, in the most efficient way, and to exploit the solved structures for the assignment of biological function to hypothetical proteins. We discuss the impact of these developments on protein classification, gene function prediction, and protein structure prediction.

Databases, Factual↗

Computational methods and evaluation of RNA stabilization reagents for genome-wide expression studies.

Gene expression studies require high quality messenger RNA (mRNA) in addition to other factors such as efficient primers and labeling reagents. To prevent RNA degradation and to improve the quality of gene array expression data, several commercial reagents have become available. We examined a conventional hot-phenol lysis method and RNA stabilization reagents, and generated comparative gene expression profiles from Escherichia coli cells grown on minimal medium. Our data indicate that certain RNA stabilization reagents induce stress responses and proper caution must be exercised during their use. We observed that the laboratory reagent (phenol/EtOH, 5:95, v/v) worked efficiently in isolating high quality mRNA and reproducibility was such that reliable gene expression profiles were generated. To assist in the analysis of gene expression data, we wrote a number of macros that use the most recent gene annotation and process data in accordance with gene function. Scripts were also written to examine the occurrence of artifacts, based on GC content, length of the individual open reading frame (ORF), its distribution on plus and minus DNA strands, and the distance from the replication origin.

Base Composition↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

Positional candidate gene selection from livestock EST databases using Gene Ontology.

MOTIVATION: The number of expressed sequence tags (ESTs) in GenBank has now surpassed 200,000 for cattle and 100,000 for swine. The Institute of Genome Research (TIGR) has organized these sequences into approximately 60,000 non-redundant consensus sequences (identified by TIGR Gene Indices) for cattle and 40,000 for swine. Anonymous ESTs are of limited value unless they are connected to function. Functional information is difficult to manage electronically because of heterogeneity of meaning and form among databases. The Gene Ontology (GO) Consortium has produced ontologies for gene function with consistent meaning and form across species. Linking livestock EST to gene function through similarity with sequences from other annotation-rich mammals could accelerate: (1) the discovery of positional candidate genes underlying a livestock quantitative trait locus (QTL) and (2) comparative mapping between livestock and other mammals (e.g. humans, mouse and rat). We initiated this investigation to determine if incorporation of the GO into the annotation process could accelerate livestock positional candidate gene discovery. RESULTS: We have associated livestock ESTs with GO nodes through sequence similarity to the NCBI Reference Sequences (RefSeq). Positional candidate genes are identified within minutes that otherwise required days. The schema described here accommodates queries that return GO nodes from terms familiar to biologists, such as gene name, alternate/alias symbol, and OMIM phenotype. AVAILABILITY: Scripts and schema are available on request from the authors.

Animals↗

The SBASE protein domain library, Release 4.0: a collection of annotated protein sequence segments.

SBASE 4.0 is the fourth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database search tools, such as FASTA and BLAST3. The present release contains 61 137 entries provided with standardized names and cross-referenced to all major protein, nucleic acid and sequence pattern collections. The entries are clustered into 13 155 groups in order to facilitate detection of distant similarities. SBASE 4.0 is freely available by anonymous ftp file transfer from ftp.icgeb.trieste.it. Individual records can be retrieved with the gopher server at icgeb.trieste.it and with a World Wide Web server at http://www.icgeb.trieste.it. Automated searching of SBASE with BLAST can be carried out with the electronic mail server sbase@icgeb.trieste.it, which now also provides a graphic representation of the homologies. A related mail server, domain@hubi.abc.hu, assigns SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

The SBASE protein domain library, release 5.0: a collection of annotated protein sequence segments.

SBASE 5.0 is the fifth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand-binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database-search programs. The present release contains over 79863 entries provided with standardized names and is cross-referenced to all major sequence databases and sequence pattern collections. The information is assigned to individual domains rather than to entire protein sequences, thus SBASE contains substantially more cross-references and links than do the protein sequence databases. The entries are clustered into >16 000 groups in order to facilitate the detection of distant similarities. SBASE 5.0 is freely available by anonymous 'ftp' file transfer from . Automated searching of SBASE with BLAST can be carried out with the WWW-server . and with the electronic mail server which now also provides a graphic representation of the homologies. A related WWW-server and e-mail server predicts SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

EchoBASE: an integrated post-genomic database for Escherichia coli.

EchoBASE (http://www.ecoli-york.org) is a relational database designed to contain and manipulate information from post-genomic experiments using the model bacterium Escherichia coli K-12. Its aim is to collate information from a wide range of sources to provide clues to the functions of the approximately 1500 gene products that have no confirmed cellular function. The database is built on an enhanced annotation of the updated genome sequence of strain MG1655 and the association of experimental data with the E.coli genes and their products. Experiments that can be held within EchoBASE include proteomics studies, microarray data, protein-protein interaction data, structural data and bioinformatics studies. EchoBASE also contains annotated information on 'orphan' enzyme activities from this microbe to aid characterization of the proteins that catalyse these elusive biochemical reactions.

Databases, Genetic↗

Argonaute--a database for gene regulation by mammalian microRNAs.

MicroRNAs (miRNAs) constitute a recently discovered class of small non-coding RNAs that regulate expression of target genes either by decreasing the stability of the target mRNA or by translational inhibition. They are involved in diverse processes, including cellular differentiation, proliferation and apoptosis. Recent evidence also suggests their importance for cancerogenesis. By far the most important model systems in cancer research are mammalian organisms. Thus, we decided to compile comprehensive information on mammalian miRNAs, their origin and regulated target genes in an exhaustive, curated database called Argonaute (http://www.ma.uni-heidelberg.de/apps/zmf/argonaute/interface). Argonaute collects latest information from both literature and other databases. In contrast to current databases on miRNAs like miRBase::Sequences, NONCODE or RNAdb, Argonaute hosts additional information on the origin of an miRNA, i.e. in which host gene it is encoded, its expression in different tissues and its known or proposed function, its potential target genes including Gene Ontology annotation, as well as miRNA families and proteins known to be involved in miRNA processing. Additionally, target genes are linked to an information retrieval system that provides comprehensive information from sequence databases and a simultaneous search of MEDLINE with all synonyms of a given gene. The web interface allows the user to get information for a single or multiple miRNAs, either selected or uploaded through a text file. Argonaute currently has information on 839 miRNAs from human, mouse and rat.

Animals↗

Cyanidioschyzon merolae genome. A tool for facilitating comparable studies on organelle biogenesis in photosynthetic eukaryotes.

The ultrasmall unicellular red alga Cyanidioschyzon merolae lives in the extreme environment of acidic hot springs and is thought to retain primitive features of cellular and genome organization. We determined the 16.5-Mb nuclear genome sequence of C. merolae 10D as the first complete algal genome. BLASTs and annotation results showed that C. merolae has a mixed gene repertoire of plants and animals, also implying a relationship with prokaryotes, although its photosynthetic components were comparable to other phototrophs. The unicellular green alga Chlamydomonas reinhardtii has been used as a model system for molecular biology research on, for example, photosynthesis, motility, and sexual reproduction. Though both algae are unicellular, the genome size, number of organelles, and surface structures are remarkably different. Here, we report the characteristics of double membrane- and single membrane-bound organelles and their related genes in C. merolae and conduct comparative analyses of predicted protein sequences encoded by the genomes of C. merolae and C. reinhardtii. We examine the predicted proteins of both algae by reciprocal BLASTP analysis, KOG assignment, and gene annotation. The results suggest that most core biological functions are carried out by orthologous proteins that occur in comparable numbers. Although the fundamental gene organizations resembled each other, the genes for organization of chromatin, cytoskeletal components, and flagellar movement remarkably increased in C. reinhardtii. Molecular phylogenetic analyses suggested that the tubulin is close to plant tubulin rather than that of animals and fungi. These results reflect the increase in genome size, the acquisition of complicated cellular structures, and kinematic devices in C. reinhardtii.

Algal Proteins↗

The molecular structure of Rv1873, a conserved hypothetical protein from Mycobacterium tuberculosis, at 1.38 A resolution.

The X-ray crystal structure of the gene product encoded by open reading frame Rv1873 of Mycobacterium tuberculosis has been determined by single isomorphous replacement with anomalous scattering (SIRAS) phasing techniques at 1.38 A resolution from monoclinic crystals with unit-cell parameters a = 33.44, b = 31.63, c = 53.19 A, beta = 90.8 degrees. The 16.2 kDa Rv1873 is a monomer that adopts a primarily alpha-helical fold with limited structural similarity to previously determined tertiary structures. It has been annotated as a conserved hypothetical protein of unknown function and is classified by the Clusters of Orthologous Groups (COG) database as belonging to COG5579. The three-dimensional structure of the Rv1873 gene product reveals limited similarity to a repeated motif that is found in a variety of other proteins. While not a novel fold, it serves as a model for orthologues predicted to be related by sequence and it is hoped that knowledge of the structure of Rv1873 will aid in determining a possible function for this protein.

Amino Acid Sequence↗

Electron transport in the pathway of acetate conversion to methane in the marine archaeon Methanosarcina acetivorans.

A liquid chromatography-hybrid linear ion trap-Fourier transform ion cyclotron resonance mass spectrometry approach was used to determine the differential abundance of proteins in acetate-grown cells compared to that of proteins in methanol-grown cells of the marine isolate Methanosarcina acetivorans metabolically labeled with 14N versus 15N. The 246 differentially abundant proteins in M. acetivorans were compared with the previously reported 240 differentially expressed genes of the freshwater isolate Methanosarcina mazei determined by transcriptional profiling of acetate-grown cells compared to methanol-grown cells. Profound differences were revealed for proteins involved in electron transport and energy conservation. Compared to methanol-grown cells, acetate-grown M. acetivorans synthesized greater amounts of subunits encoded in an eight-gene transcriptional unit homologous to operons encoding the ion-translocating Rnf electron transport complex previously characterized from the Bacteria domain. Combined with sequence and physiological analyses, these results suggest that M. acetivorans replaces the H2-evolving Ech hydrogenase complex of freshwater Methanosarcina species with the Rnf complex, which generates a transmembrane ion gradient for ATP synthesis. Compared to methanol-grown cells, acetate-grown M. acetivorans synthesized a greater abundance of proteins encoded in a seven-gene transcriptional unit annotated for the Mrp complex previously reported to function as a sodium/proton antiporter in the Bacteria domain. The differences reported here between M. acetivorans and M. mazei can be attributed to an adaptation of M. acetivorans to the marine environment.

Acetates↗

MADCAP: isolation of novel nAb-naïve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV↗

Discovery of eight novel divergent homologs expressed in cattle placenta.

Ten divergent homologs were identified using a subtractive bioinformatic analysis of 12,614 cattle placenta expressed sequence tags followed by comparative, evolutionary, and gene expression studies. Among the 10 divergent homologs, 8 have not been identified previously. These were named as follows: cattle cerebrum and skeletal muscle-specific transcript 1 (CSSMST1), cattle intestine-specific transcript 1 (CIST1), hepatitis A virus cellular receptor 1 amino-terminal domain-containing protein (HAVCRNDP), prolactin-related proteins 8, 9, and 11 (PRP8, PRP9, and PRP11, respectively) and secreted and transmembrane protein 1A and 1B (SECTM1A and SECTM1B, respectively). In addition, two previously known divergent genes were identified, trophoblast Kunitz domain protein 1 (TKDP1) and a new splice variant of TKDP4. Nucleotide substitution analysis provided evidence for positive selection in members of the PRP gene family, SECTM1A and SECTM1B. Gene expression profiles, motif predictions, and annotations of homologous sequences indicate immunological and reproductive functions of the divergent homologs. The genes identified in this study are thus of evolutionary and physiological importance and may have a role in placental adaptations.

Amino Acid Sequence↗

Fast and sensitive multiple alignment of large genomic sequences.

BACKGROUND: Genomic sequence alignment is a powerful method for genome analysis and annotation, as alignments are routinely used to identify functional sites such as genes or regulatory elements. With a growing number of partially or completely sequenced genomes, multiple alignment is playing an increasingly important role in these studies. In recent years, various tools for pair-wise and multiple genomic alignment have been proposed. Some of them are extremely fast, but often efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast search program identifies a chain of strong local sequence similarities. In a second step, regions between these anchor points are aligned using a slower but more accurate method. RESULTS: Herein, we present CHAOS, a novel algorithm for rapid identification of chains of local pair-wise sequence similarities. Local alignments calculated by CHAOS are used as anchor points to improve the running time of DIALIGN, a slow but sensitive multiple-alignment tool. We show that this way, the running time of DIALIGN can be reduced by more than 95% for BAC-sized and longer sequences, without affecting the quality of the resulting alignments. We apply our approach to a set of five genomic sequences around the stem-cell-leukemia (SCL) gene and demonstrate that exons and small regulatory elements can be identified by our multiple-alignment procedure. CONCLUSION: We conclude that the novel CHAOS local alignment tool is an effective way to significantly speed up global alignment tools such as DIALIGN without reducing the alignment quality. We likewise demonstrate that the DIALIGN/CHAOS combination is able to accurately align short regulatory sequences in distant orthologues.

Algorithms↗

Application of magnetic techniques in the field of drug discovery and biomedicine.

Magnetic separation technology, using magnetic particles, is quick and easy method for sensitive and reliable capture of specific proteins, genetic material and other biomolecules. The technique offers an advantage in terms of subjecting the analyte to very little mechanical stress compared to other methods. Secondly, these methods are non-laborious, cheap and often highly scalable. Moreover, techniques employing magnetism are more amenable to automation and miniaturization. Now that the human genome is sequenced and about 30,000 genes are annotated, the next step is to identify the function of these individual genes, carrying out genotyping studies for allelic variation and SNP analysis, ultimately leading to identification of novel drug targets. In this post-genomic era, technologies based on magnetic separation are becoming an integral part of todays biology laboratory. This article briefly reviews the selected applications of magnetic separation techniques in the field of biotechnology, biomedicine and drug discovery.

Journal Article↗

RNAi screening of uncharacterized genes identifies promising druggable targets in Schistosoma japonicum.

Schistosomiasis affects more than 250 million people worldwide and is one of the neglected tropical diseases. Currently, the treatment of schistosomiasis relies on a single drug-praziquantel-which has led to increasing pressure from drug resistance. Therefore, there is an urgent need to find new treatments. The development of genome sequencing has provided valuable information for understanding the biology of schistosomes. In the genome of Schistosoma japonicum, approximately 11% of the protein-coding sequences are uncharacterized genes (UGs) annotated as "hypothetical protein" or "protein of unknown function." These poorly understood genes have been unjustifiably neglected, although some may be essential for the survival of the parasites and serve as potential drug targets. In this study, we systematically mined the highly expressed UGs in both genders of this parasite throughout key developmental stages in their mammalian host, using our previously published S. japonicum genome and RNA-seq data. By employing in vitro RNA interference (RNAi), we screened 126 UGs that lack homologs in Homo sapiens and identified 8 that are essential for the parasite vitality. We further investigated two UGs, Sjc_0002003 and Sjc_0009272, which resulted in the most severe phenotypes. Fluorescence in situ hybridization demonstrated that both genes were expressed throughout the body without sex bias. Silencing either Sjc_0002003 or Sjc_0009272 reduced the cell proliferation in the body. Furthermore, in vivo RNAi indicated both genes are required for the growth and survival of the parasites in the mammalian host. For Sjc_0002003, we further characterize the underlying molecular cause of the observed phenotype. Through RNA-seq analysis and functional studies, we revealed that silencing Sjc_0002003 reduces the expression of a series of intestinal genes, including Sjc_0007312 (hypothetical protein), Sjc_0008276 (vha-17), Sjc_0002942 (PLA2G15), and Sjc_0003646 (SJCHGC09134 protein), leading to gut dilation. Our work highlights the importance of UGs in schistosomes as promising targets for drug development in the treatment of the schistosomiasis.

Schistosoma japonicum↗

Rapid proteome analysis of bronchoalveolar lavage samples of lifelong smokers and never-smokers by micro-scale liquid chromatography and mass spectrometry.

BACKGROUND: The aim of this study was to determine whether relative qualitative and quantitative differences in protein expression could be related to smoke exposure or smoke-induced airway inflammation. We therefore explored and characterized the protein components found in bronchoalveolar lavage (BAL) fluid sampled from either lifelong smokers or never-smokers. METHODS: BAL fluid samples obtained by bronchoscopy from 60-year-old healthy never-smokers (n = 18) and asymptomatic smokers (n = 30) were analyzed in either pooled or individual form. Initial global proteomic analysis used shotgun digestion approaches on unfractionated BAL fluid samples (after minimal sample preparation) and separation of peptides by gradient (90-min) liquid chromatography (LC) coupled with on-line linear ion trap quadropole mass spectrometry (LTQ MS) for identification and analysis. RESULTS: LTQ MS identified 481 high- to low-abundance proteins. Relative differences in patterns of BAL fluid proteins in smokers compared with never-smokers were observed in pooled and individual samples as well as by 2-dimensional gel analysis. Gene ontology categorization of all annotated proteins showed a wide spectrum of molecular functions and biological processes. CONCLUSIONS: The described method provides comprehensive qualitative proteomic analysis of BAL fluid protein expression from never-smokers and from smokers at risk of developing chronic obstructive pulmonary disease. Many of the proteins identified had not been detected in previous studies of BAL fluid; thus, the use of LC-tandem MS with LTQ may provide new information regarding potentially important patterns of protein expression associated with lifelong smoking.

Bronchoalveolar Lavage Fluid↗