Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

The complete genome sequence of Escherichia coli K-12.

The 4,639,221-base pair sequence of Escherichia coli K-12 is presented. Of 4288 protein-coding genes annotated, 38 percent have no attributed function. Comparison with five other sequenced microbes reveals ubiquitous as well as narrowly distributed gene families; many families of similar genes within E. coli are also evident. The largest family of paralogous proteins contains 80 ABC transporters. The genome as a whole is strikingly organized with respect to the local direction of replication; guanines, oligonucleotides possibly related to replication and recombination, and most genes are so oriented. The genome also contains insertion sequence (IS) elements, phage remnants, and many other patches of unusual composition indicating genome plasticity through horizontal transfer.

Bacterial Proteins↗

New recombination methods for Sinorhizobium meliloti genetics.

The availability of bacterial genome sequences has created a need for improved methods for sequence-based functional analysis to facilitate moving from annotated DNA sequence to genetic materials for analyzing the roles that postulated genes play in bacterial phenotypes. A powerful cloning method that uses lambda integrase recombination to clone and manipulate DNA sequences has been adapted for use with the gram-negative alpha-proteobacterium Sinorhizobium meliloti in two ways that increase the utility of the system. Adding plasmid oriT sequences to a set of vehicles allows the plasmids to be transferred to S. meliloti by conjugation and also allows cloned genes to be recombined from one plasmid to another in vivo by a pentaparental mating protocol, saving considerable time and expense. In addition, vehicles that contain yeast Flp recombinase target recombination sequences allow the construction of deletion mutations where the end points of the deletions are located at the ends of the cloned genes. Several deletions were constructed in a cluster of 60 genes on the symbiotic plasmid (pSymA) of S. meliloti, predicted to code for a denitrification pathway. The mutations do not affect the ability of the bacteria to form nitrogen-fixing nodules on Medicago sativa (alfalfa) roots.

Acetylene↗

Evaluation of the host transcriptional response to human cytomegalovirus infection.

Gene expression data from human cytomegalovirus (HCMV)-infected cells were analyzed using DNA-Chip Analyzer (dChip) followed by singular value decomposition (SVD) and compared with a previous analysis of the same data that employed GeneChip software and a fold change filtering approach. dChip and SVD analysis revealed two clusters of coexpressed human genes responding differently to HCMV infection: one containing some genes identified previously, and another that was largely unique to this analysis. Annotating these genes, we identified several functional categories important to host cell responses to HCMV infection. These categories included genes involved in transcriptional regulation, oncogenesis, and cell cycle regulation, which were more prevalent in cluster 1, and genes involved in immune system regulation, signal transduction, and cell adhesion, which were more prevalent in cluster 2. Within these categories, we found genes involved in the host response to HCMV infection (mainly in cluster 1), as well as genes targeted by HCMV's immune evasion strategies (mainly in cluster 2). As the second group of genes identified by the dChip and SVD approach was statistically and biologically significant, our results point out the advantages of using different methods to analyze gene expression data.

Cytomegalovirus Infections↗

A SNP-centric database for the investigation of the human genome.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are an increasingly important tool for genetic and biomedical research. Although current genomic databases contain information on several million SNPs and are growing at a very fast rate, the true value of a SNP in this context is a function of the quality of the annotations that characterize it. Retrieving and analyzing such data for a large number of SNPs often represents a major bottleneck in the design of large-scale association studies. DESCRIPTION: SNPper is a web-based application designed to facilitate the retrieval and use of human SNPs for high-throughput research purposes. It provides a rich local database generated by combining SNP data with the Human Genome sequence and with several other data sources, and offers the user a variety of querying, visualization and data export tools. In this paper we describe the structure and organization of the SNPper database, we review the available data export and visualization options, and we describe how the architecture of SNPper and its specialized data structures support high-volume SNP analysis. CONCLUSIONS: The rich annotation database and the powerful data manipulation and presentation facilities it offers make SNPper a very useful online resource for SNP research. Its success proves the great need for integrated and interoperable resources in the field of computational biology, and shows how such systems may play a critical role in supporting the large-scale computational analysis of our genome.

Databases, Genetic↗

LS-NMF: a modified non-negative matrix factorization algorithm utilizing uncertainty estimates.

BACKGROUND: Non-negative matrix factorisation (NMF), a machine learning algorithm, has been applied to the analysis of microarray data. A key feature of NMF is the ability to identify patterns that together explain the data as a linear combination of expression signatures. Microarray data generally includes individual estimates of uncertainty for each gene in each condition, however NMF does not exploit this information. Previous work has shown that such uncertainties can be extremely valuable for pattern recognition. RESULTS: We have created a new algorithm, least squares non-negative matrix factorization, LS-NMF, which integrates uncertainty measurements of gene expression data into NMF updating rules. While the LS-NMF algorithm maintains the advantages of original NMF algorithm, such as easy implementation and a guaranteed locally optimal solution, the performance in terms of linking functionally related genes has been improved. LS-NMF exceeds NMF significantly in terms of identifying functionally related genes as determined from annotations in the MIPS database. CONCLUSION: Uncertainty measurements on gene expression data provide valuable information for data analysis, and use of this information in the LS-NMF algorithm significantly improves the power of the NMF technique.

Algorithms↗

REEF: searching REgionally Enriched Features in genomes.

BACKGROUND: In Eukaryotic genomes, different features including genes are not uniformly distributed. The integration of annotation information and genomic position of functional DNA elements in the Eukaryotic genomes opened the way to test novel hypotheses of higher order genome organization and regulation of expression. RESULTS: REEF is a new tool, aimed at identifying genomic regions enriched in specific features, such as a class or group of genes homogeneous for expression and/or functional characteristics. The method for the calculation of local feature enrichment uses test statistic based on the Hypergeometric Distribution applied genome-wide by using a sliding window approach and adopting the False Discovery Rate for controlling multiplicity. REEF software, source code and documentation are freely available at http://telethon.bio.unipd.it/bioinfo/reef/. CONCLUSION: REEF can aid to shed light on the role of organization of specific genomic regions in the determination of their functional role.

Algorithms↗

An expressed sequence tag (EST) library from developing fruits of an Hawaiian endemic mint (Stenogyne rugosa, Lamiaceae): characterization and microsatellite markers.

BACKGROUND: The endemic Hawaiian mints represent a major island radiation that likely originated from hybridization between two North American polyploid lineages. In contrast with the extensive morphological and ecological diversity among taxa, ribosomal DNA sequence variation has been found to be remarkably low. In the past few years, expressed sequence tag (EST) projects on plant species have generated a vast amount of publicly available sequence data that can be mined for simple sequence repeats (SSRs). However, these EST projects have largely focused on crop or otherwise economically important plants, and so far only few studies have been published on the use of intragenic SSRs in natural plant populations. We constructed an EST library from developing fleshy nutlets of Stenogyne rugosa principally to identify genetic markers for the Hawaiian endemic mints. RESULTS: The Stenogyne fruit EST library consisted of 628 unique transcripts derived from 942 high quality ESTs, with 68% of unigenes matching Arabidopsis genes. Relative frequencies of Gene Ontology functional categories were broadly representative of the Arabidopsis proteome. Many unigenes were identified as putative homologs of genes that are active during plant reproductive development. A comparison between unigenes from Stenogyne and tomato (both asterid angiosperms) revealed many homologs that may be relevant for fruit development. Among the 628 unigenes, a total of 44 potentially useful microsatellite loci were predicted. Several of these were successfully tested for cross-transferability to other Hawaiian mint species, and at least five of these demonstrated interesting patterns of polymorphism across a large sample of Hawaiian mints as well as close North American relatives in the genus Stachys. CONCLUSION: Analysis of this relatively small EST library illustrated a broad GO functional representation. Many unigenes could be annotated to involvement in reproductive development. Furthermore, first tests of microsatellite primer pairs have proven promising for the use of Stenogyne rugosa EST SSRs for evolutionary and phylogeographic studies of the Hawaiian endemic mints and their close relatives. Given that allelic repeat length variation in developmental genes of other organisms has been linked with morphological evolution, these SSRs may also prove useful for analyses of phenotypic differences among Hawaiian mints.

5' Untranslated Regions↗

Construction of a full-length cDNA library from young spikelets of hexaploid wheat and its characterization by large-scale sequencing of expressed sequence tags.

The polyploid nature of wheat is a key characteristic of the plant. Full-length complementary DNAs (cDNAs) provide essential information that can be used to annotate the genes and provide a functional analysis of these genes and their products. We constructed a full-length cDNA library derived from young spikelets of common wheat, and obtained 24056 expressed sequence tags (ESTs) from both ends of the cDNA clones. These ESTs were grouped into 3605 contigs using the phrap method, representing expressed loci from each of the three genomes. Using BLAST, 3605 contigs were grouped into 1902 gene clusters, showing that loci of the three genomes are not always expressed. A homology search of these gene clusters against a wheat EST database (15964 gene clusters) and a rice full-length cDNA database (21447 gene clusters) revealed that a quarter of the wheat full-length cDNAs were novel. A protein database of Arabidopsis was used to examine the functional classification of these gene clusters. The GC-content in the 5 -UTR region of wheat cDNAs was compared to that of rice. Forty-three genes (3.5% of wheat cDNAs homologous to those of rice) possessed distinct GC-content in the 5 -UTR region, suggesting different breeding behaviors of wheat and rice.

5' Untranslated Regions↗

Random sequencing of cDNA library derived from partially-fed adult female Haemaphysalis longicornis salivary gland.

A cDNA library was constructed from salivary glands of partially-fed adult female Haemaphysalis longicornis (hard tick). Randomly selected clones were sequenced and a total of 633 sequences were analyzed by bioinformatic programs. The sequences were grouped into 213 clusters, with each cluster being considered to be composed of mRNAs derived from the same gene or closely related genes. About 36% of the mRNA sequences showed significant similarity to known proteins in the non-redundant protein database by the NCBI blastx program and appeared to be coding for functional predicted proteins, whereas the remaining 64% had no similar sequences. Two thirds of the predicted proteins were annotated as basic cellular proteins (housekeeping proteins). Among the functional predicted protein sequences, other than the housekeeping proteins, several protease inhibitors including anticoagulants, two metalloproteases and a potential immunosuppressive protein could be identified. These proteins may play important roles during tick feeding and could be novel anti-tick vaccine candidates.

Amino Acid Sequence↗

High genetic diversity in the chemoreceptor superfamily of Caenorhabditis elegans.

We investigated genetic polymorphism in the Caenorhabditis elegans srh and str chemoreceptor gene families, each of which consists of approximately 300 genes encoding seven-pass G-protein-coupled receptors. Almost one-third of the genes in each family are annotated as pseudogenes because of apparent functional defects in N2, the sequenced wild-type strain of C. elegans. More than half of these "pseudogenes" have only one apparent defect, usually a stop codon or deletion. We sequenced the defective region for 31 such genes in 22 wild isolates of C. elegans. For 10 of the 31 genes, we found an apparently functional allele in one or more wild isolates, suggesting that these are not pseudogenes but instead functional genes with a defective allele in N2. We suggest the term "flatliner" to describe genes whose functional vs. pseudogene status is unclear. Investigations of flatliner gene positions, d(N)/d(S) ratios, and phylogenetic trees indicate that they are not readily distinguished from functional genes in N2. We also report striking heterogeneity in the frequency of other polymorphisms among these genes. Finally, the large majority of polymorphism was found in just two strains from geographically isolated islands, Hawaii and Madeira. This suggests that our sampling of wild diversity in C. elegans is narrow and that identification of additional strains from similarly isolated regions will greatly expand the diversity available for study.

Alleles↗

A secure semantic interoperability infrastructure for inter-enterprise sharing of electronic healthcare records.

Healthcare professionals need access to accurate and complete healthcare records for effective assessment, diagnosis and treatment of patients. The non-interoperability of healthcare information systems means that interenterprise access to a patient's history over many distributed encounters is difficult to achieve. The ARTEMIS project has developed a secure semantic web service infrastructure for the interoperability of healthcare information systems. Healthcare professionals share services and medical information using a web service annotation and mediation environment based on functional and clinical semantics derived from healthcare standards. Healthcare professionals discover medical information about individuals using a patient identification protocol based on pseudonymous information. The management of care pathways and access to medical information is based on a well-defined business process allowing healthcare providers to negotiate collaboration and data access agreements within the context of strict legislative frameworks.

Computer Security↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

The Pyridoxal-5'-Phosphate-Dependent Enzymes of Mycobacterium tuberculosis.

Enzymes that depend on the cofactor pyridoxal 5'-phosphate (PLP) catalyze a remarkable variety of biochemical reactions in all organisms. In particular, the genome of Mycobacterium tuberculosis, the causative agent of tuberculosis (TB), encodes 45 bona fide PLP-dependent enzymes plus a few related proteins that presumably do not have enzymic function. The large majority of the 45 enzymes have been characterized in terms of catalytic activity and structure. Several of them have been shown to be central to the bacterium's survival and pathogenicity, while some of these enzymes are targets of an extant drug (d-cycloserine). Herein, the annotated catalog of the PLP-dependent enzymes in M. tuberculosis is presented and analyzed with three main goals in mind. The first will be to assess the specific aspects of mycobacterial metabolism that rely most on PLP-dependent enzymes. A second goal will be to signal those enzymes whose function is still uncertain and whose functional characterization may help to further understand the biology of M. tuberculosis. Finally, we will examine the potential and limitations of targeting the PLP-dependent enzymes for the development of new antimycobacterial drugs.

Mycobacterium tuberculosis↗

Preparation, crystallization and preliminary X-ray analysis of XC2382, an ApaG protein of unknown structure from Xanthomonas campestris.

Xanthomonas campestris pv. campestris is the causative agent of black rot, one of the major worldwide diseases of cruciferous crops. Its genome encodes approximately 4500 proteins, roughly one third of which have unknown function. XC2382 is one such protein, with a MW of 14.2 kDa. Based on a bioinformatics study, it was annotated as an ApaG gene product that serves multiple functions. The ApaG protein has been overexpressed in Escherichia coli, purified and crystallized using the hanging-drop vapour-diffusion method. The crystals diffracted to a resolution of at least 2.30 A. They are tetragonal and belong to space group P4(1/3), with unit-cell parameters a = b = 57.6, c = 122.9 A. There are two, three or four molecules in the asymmetric unit.

Bacterial Proteins↗

AthaMap: from in silico data to real transcription factor binding sites.

AthaMap generates a map for cis-regulatory sequences for the whole Arabidopsis thaliana genome. AthaMap was initially developed by matrix-based detection of putative transcription factor binding sites (TFBS) mostly determined from random binding site selection experiments. Now, also experimentally verified TFBS have been included for 48 different Arabidopsis thaliana transcription factors (TF). Based on these sequences, 89,416 very similar putative TFBS were determined within the genome of A. thaliana and annotated to AthaMap. Matrix- and single sequence-based binding sites can be included in colocalization analysis for the identification of combinatorial cis-regulatory elements. As an example, putative target genes of the WRKY18 transcription factor that is involved in plant-pathogen interaction were determined. New functions of AthaMap include descriptions for all annotated Arabidopsis thaliana genes and direct links to TAIR, TIGR and MIPS. Transcription factors used in the binding site determination are linked to TAIR and TRANSFAC databases. AthaMap is freely available at http://www.athamap.de.

Arabidopsis↗

Genomic analysis of regulatory mechanisms governing EPS66A biosynthesis in Streptomyces changanensis HL-66.

Streptomyces changanensis HL-66 produces the α-(1,4)/(1,6)-glucan exopolysaccharide EPS66A, a potent plant immune elicitor with promising applications in plant protection. However, its low native fermentation yield limits large-scale application. To investigate the biosynthetic potential and regulatory mechanisms underlying EPS66A production, the whole genome of HL-66 was sequenced and analyzed. The HL-66 genome is 6.82 Mb in size, with a GC content of 74%, and encodes 6081 predicted functional genes. Among these, 1390 genes were annotated to Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, 4187 were assigned to Gene Ontology (GO) terms, and 143 were classified into Clusters of Orthologous Groups (COG) categories. antiSMASH analysis identified 22 secondary metabolite biosynthetic gene clusters, including multiple polyketide synthase (PKS) and nonribosomal peptide synthetase (NRPS) clusters. Functional analyses revealed that the glycosyltransferase gene (GTy) and the global regulatory gene (bldD) are involved in EPS66A biosynthesis. bldD is involved in morphological development and EPS66A production, whereas GTy specifically regulates EPS66A production without affecting growth or development. In both in vivo and potted-plant experiments, EPS66A (200 μg/mL) significantly reduced the severity of tobacco mosaic virus, apple anthracnose leaf spot, walnut bacterial leaf spot, and jujube anthracnose, achieving control efficacies of 90.21%, 87.95%, 77.41%, and 68.55%, respectively, and outperforming a commercial chitosan oligosaccharide control. These findings provide new insights into the genetic architecture and regulatory mechanisms of EPS66A biosynthesis and support its development as a polysaccharide-based green pesticide.

Streptomyces↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: cross-references to additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT. The URLs for SWISS-PROT on the WWW are: http://www.expasy.ch/sprot and http://www. ebi.ac.uk/sprot

Amino Acid Sequence↗

Gene expression profiles underlying alternative caste phenotypes in a highly eusocial bee, Melipona quadrifasciata.

To evaluate caste-biased gene expression in Melipona quadrifasciata, a stingless bee, we generated 1278 ESTs using Representational Difference Analysis. Most annotated sequences were similar to honey bee genes of unknown function. Only few queen-biased sequences had their putative function assigned by sequence comparison, contrasting with the worker-biased ESTs. The expression of six annotated genes connected to caste specificity was validated by real time PCR. Interestingly, queens that were developmentally induced by treatment with a juvenile hormone analogue displayed an expression profile clearly different from natural queens for this set of genes. In summary, this study represents an important first step in applying a comparative genomic approach to queen/worker polyphenism in the bee.

Animals↗