Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

The ABC transporter gene family of Caenorhabditis elegans has implications for the evolutionary dynamics of multidrug resistance in eukaryotes.

BACKGROUND: Many drugs of natural origin are hydrophobic and can pass through cell membranes. Hydrophobic molecules must be susceptible to active efflux systems if they are to be maintained at lower concentrations in cells than in their environment. Multi-drug resistance (MDR), often mediated by intrinsic membrane proteins that couple energy to drug efflux, provides this function. All eukaryotic genomes encode several gene families capable of encoding MDR functions, among which the ABC transporters are the largest. The number of candidate MDR genes means that study of the drug-resistance properties of an organism cannot be effectively carried out without taking a genomic perspective. RESULTS: We have annotated sequences for all 60 ABC transporters from the Caenorhabditis elegans genome, and performed a phylogenetic analysis of these along with the 49 human, 30 yeast, and 57 fly ABC transporters currently available in GenBank. Classification according to a unified nomenclature is presented. Comparison between genomes reveals much gene duplication and loss, and surprisingly little orthology among analogous genes. Proteins capable of conferring MDR are found in several distinct subfamilies and are likely to have arisen independently multiple times. CONCLUSIONS: ABC transporter evolution fits a pattern expected from a process termed 'dynamic-coherence'. This is an unusual result for such a highly conserved gene family as this one, present in all domains of cellular life. Mechanistically, this may result from the broad substrate specificity of some ABC proteins, which both reduces selection against gene loss, and leads to the facile sorting of functions among paralogs following gene duplication.

ATP-Binding Cassette Transporters↗

Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation.

An algorithm is described for the systematic characterization of the physico-chemical properties seen at each position in a multiple protein sequence alignment. The new algorithm allows questions important in the design of mutagenesis experiments to be quickly answered since positions in the alignment that show unusual or interesting residue substitution patterns may be rapidly identified. The strategy is based on a flexible set-based description of amino acid properties, which is used to define the conservation between any group of amino acids. Sequences in the alignment are gathered into subgroups on the basis of sequence similarity, functional, evolutionary or other criteria. All pairs of subgroups are then compared to highlight positions that confer the unique features of each subgroup. The algorithm is encoded in the computer program AMAS (Analysis of Multiply Aligned Sequences) which provides a textual summary of the analysis and an annotated (boxed, shaded and/or coloured) multiple sequence alignment. The algorithm is illustrated by application to an alignment of 67 SH2 domains where patterns of conserved hydrophobic residues that constitute the protein core are highlighted. The analysis of charge conservation across annexin domains identifies the locations at which conserved charges change sign. The algorithm simplifies the analysis of multiple sequence data by condensing the mass of information present, and thus allows the rapid identification of substitutions of structural and functional importance.

Algorithms↗

DIAN: a novel algorithm for genome ontological classification.

Faced with the determination of many completely sequenced genomes, computational biology is now faced with the challenge of interpreting the significance of these data sets. A multiplicity of data-related problems impedes this goal: Biological annotations associated with raw data are often not normalized, and the data themselves are often poorly interrelated and their interpretation unclear. All of these problems make interpretation of genomic databases increasingly difficult. With the current explosion of sequences now available from the human genome as well as from model organisms, the importance of sorting this vast amount of conceptually unstructured source data into a limited universe of genes, proteins, functions, structures, and pathways has become a bottleneck for the field. To address this problem, we have developed a method of interrelating data sources by applying a novel method of associating biological objects to ontologies. We have developed an intelligent knowledge-based algorithm, to support biological knowledge mapping, and, in particular, to facilitate the interpretation of genomic data. In this respect, the method makes it possible to inventory genomes by collapsing multiple types of annotations and normalizing them to various ontologies. By relying on a conceptual view of the genome, researchers can now easily navigate the human genome in a biologically intuitive, scientifically accurate manner.

Algorithms↗

Predicting subcellular localization of proteins using machine-learned classifiers.

MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.

Algorithms↗

The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.

For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha+beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the functions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the database overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc++ +.

Enzymes↗

Gene expression patterns in the developing murine placenta.

OBJECTIVE: Successful placental development is crucial for optimal growth, maturation, and survival of the embryo/fetus. To examine genetic aspects of placental development, we investigated gene expression patterns in the murine placenta at embryonic day 10.5 (E10.5), E12.5, E15.5, and E17.5. METHODS: By use of the Affymetrix MU74A array (Affymetrix, Santa Clara, CA), we measured expression levels for 12,473 probe sets. Using pairwise analysis we selected 622 probe sets, corresponding to 599 genes, that were up- or down-regulated by more than fourfold between time points E10.5 and E12.5, E12.5 and E15.5, E15.5 and E17.5. We analyzed and functionally annotated those genes regulated during development. RESULTS: In comparing E10.5 to E12.5 we found that angiogenesis and fatty acid metabolism and transport related genes were up-regulated at E10.5, while genes involved in hormonal control and ribosomal proteins were up-regulated at E12.5. When comparing E12.5 to E15.5 we noted that genes involved in the cell cycle and RNA metabolism were strongly up-regulated at E12.5, while genes involved in cellular transport were up-regulated at E15.5. Finally, when comparing E15.5 to E17.5, we found genes related to cell cycle control, genes expressed in the nucleus and involved in RNA metabolism were up-regulated at E17.5. CONCLUSION: Microarray analysis has allowed us to describe gene expression patterns and profiles in the developing mouse placenta. Further analysis has demonstrated that several functional classes are up- and down-regulated at specific time points in placental development. These changes may have significant implications for placental development in the human.

Animals↗

T(E)Xtopo: shaded membrane protein topology plots in LAT(E)X2epsilon.

UNLABELLED: T(E)Xtopo is a LAT(E)X2epsilon macro package for plotting topology data directly from PHD predictions or SwissProt database files in publication-ready quality. The plot can be shaded automatically to emphasize conserved residues or functional properties of the residue sidechains. The addition of rich decorations, such as labels, annotations and legends, is easily accomplished. AVAILABILITY: The T(E)Xtopo macro package and a full on-line documentation are freely available at http://homepages.uni-tuebingen.de/beitz/ CONTACT: eric.beitz@uni-tuebingen.de

Amino Acid Sequence↗

Assigning function to yeast proteins by integration of technologies.

Interpreting genome sequences requires the functional analysis of thousands of predicted proteins, many of which are uncharacterized and without obvious homologs. To assess whether the roles of large sets of uncharacterized genes can be assigned by targeted application of a suite of technologies, we used four complementary protein-based methods to analyze a set of 100 uncharacterized but essential open reading frames (ORFs) of the yeast Saccharomyces cerevisiae. These proteins were subjected to affinity purification and mass spectrometry analysis to identify copurifying proteins, two-hybrid analysis to identify interacting proteins, fluorescence microscopy to localize the proteins, and structure prediction methodology to predict structural domains or identify remote homologies. Integration of the data assigned function to 48 ORFs using at least two of the Gene Ontology (GO) categories of biological process, molecular function, and cellular component; 77 ORFs were annotated by at least one method. This combination of technologies, coupled with annotation using GO, is a powerful approach to classifying genes.

Computational Biology↗

Deciphering the Role of LNX2 as a Potential Contributor to Neurodevelopmental Disorders.

BACKGROUND/OBJECTIVES: Attention-deficit/hyperactivity disorder (ADHD) is a common neurodevelopmental condition characterized by a complex and multifactorial genetic architecture. In this study, we report a male patient, born to non-consanguineous healthy parents, presenting with ADHD and oppositional defiant disorder (ODD). METHODS: Trio-based whole-exome sequencing (WES) was performed in the proband and both parents. Variant classification was performed according to American College of Medical Genetics and Genomics (ACMG) guidelines, and the potential pathogenicity of the identified variant was further assessed through multiple in silico prediction algorithms and protein structural analyses. RESULTS: WES identified a homozygous variant in the LNX2 gene (NM_153371.4: c.1165G>A, p.Ala389Thr), classified as a variant of uncertain significance (VUS) and supported by multiple in silico predictions. LNX2 is expressed during brain development and encodes an E3 ubiquitin ligase involved in neuronal differentiation and synaptic function. The identified variant is located within the PDZ2 domain, a functionally relevant region involved in protein-protein interactions. Although the variant is reported in population databases (gnomAD ID: rs148429804), it has not been associated with any clinical phenotype, and its presence in the homozygous state has been reported only once, remaining extremely rare and lacking clinical annotation. Structural modelling predicted localized rearrangement of the hydrogen-bonding network within the PDZ2 domain without major conformational changes. Integrative transcriptomic, and single-cell analyses further supported the biological relevance of LNX2 in neurodevelopment, highlighting its preferential association with neuronal projection-cell networks, synaptic vesicle trafficking pathways, and neuron-specific regulatory programs. CONCLUSION: Although the identified LNX2 variant cannot be considered causative for the patient's phenotype and a definitive disease-gene relationship cannot be established based on a single individual, the complementary genetic, structural, and transcriptomic findings support the biological plausibility of LNX2 as a candidate gene for neurodevelopmental disorders. Additional independent patients and functional studies will be required to clarify its contribution to human disease.

Child↗

SMART: identification and annotation of domains from signalling and extracellular protein sequences.

SMART is a simple modular architecture research tool and database that provides domain identification and annotation on the WWW (http://coot.embl-heidelberg.de/SMART). The tool compares query sequences with its databases of domain sequences and multiple alignments whilst concurrently identifying compositionally biased regions such as signal peptide, transmembrane and coiled coil segments. Annotated and unannotated regions of the sequence can be used as queries in searches of sequence databases. The SMART alignment collection represents more than 250 signalling and extracellular domains. Each alignment is curated to assign appropriate domain boundaries and to ensure its quality. In addition, each domain is annotated extensively with respect to cellular localisation, species distribution, functional class, tertiary structure and functionally important residues.

Amino Acid Sequence↗

Genomic Identification and Comparative Characterization of Chemosensory Genes in Two Walnut Pests.

Conogethes punctiferalis (generalist) and Atrijuglans aristata (specialist) are important pests of walnut fruits. Chemosensory genes play critical roles in host location and mating, making them promising candidates for pest management research. However, genome-wide identification of these gene families has not yet been performed for either species, and cross-species comparisons are often confounded by differences in annotation quality and methodology. Here, we employed a unified pipeline for genome annotation and gene family identification to systematically characterize the odorant receptor (OR), gustatory receptor (GR), odorant-binding protein (OBP), and chemosensory protein (CSP) genes of both species and further analyzed their physicochemical properties, chromosomal distribution, and phylogeny. Genome annotation identified 13,236 and 14,083 protein-coding genes in C. punctiferalis and A. aristata, respectively, with BUSCO completeness of 94.6% and 94.5%. We identified 121 candidate chemosensory genes in C. punctiferalis (46 ORs, 23 GRs, 32 OBPs, and 20 CSPs) and 137 in A. aristata (65 ORs, 26 GRs, 28 OBPs, and 18 CSPs). Notably, A. aristata possesses more OR genes (65) than C. punctiferalis (46). Both species possess a single conserved ORco, with three and four pheromone receptor (PR) genes in C. punctiferalis and A. aristata, respectively. Chemosensory genes were either dispersed or tandemly arrayed on chromosomes, with each family clustering into conserved functional branches. This study provides a reliable foundation for comparative chemosensory evolution studies and a catalog of candidate genes for future functional validation.

Atrijuglans aristata↗

Recent developments in microarray-based enzyme assays: from functional annotation to substrate/inhibitor fingerprinting.

Recent advances in proteomics have provided impetus towards the development of robust technologies for high-throughput studies of enzymes. The term "catalomics" defines an emerging '-omics' field in which high-throughput studies of enzymes are carried out by using advanced chemical proteomics approaches. Of the various available methods, microarrays have emerged as a powerful and versatile platform to accelerate not only the functional annotation but also the substrate and inhibitor specificity (e.g. substrate and inhibitor fingerprinting, respectively) of enzymes. Herein, we review recent developments in the fabrication of various types of microarray technologies (protein-, peptide- and small-molecule-based microarrays) and their applications in high-throughput characterizations of enzymes.

Animals↗

Comparative experiments on learning information extractors for proteins and their interactions.

OBJECTIVE: Automatically extracting information from biomedical text holds the promise of easily consolidating large amounts of biological knowledge in computer-accessible form. This strategy is particularly attractive for extracting data relevant to genes of the human genome from the 11 million abstracts in Medline. However, extraction efforts have been frustrated by the lack of conventions for describing human genes and proteins. We have developed and evaluated a variety of learned information extraction systems for identifying human protein names in Medline abstracts and subsequently extracting information on interactions between the proteins. METHODS AND MATERIAL: We used a variety of machine learning methods to automatically develop information extraction systems for extracting information on gene/protein name, function and interactions from Medline abstracts. We present cross-validated results on identifying human proteins and their interactions by training and testing on a set of approximately 1000 manually-annotated Medline abstracts that discuss human genes/proteins. RESULTS: We demonstrate that machine learning approaches using support vector machines and maximum entropy are able to identify human proteins with higher accuracy than several previous approaches. We also demonstrate that various rule induction methods are able to identify protein interactions with higher precision than manually-developed rules. CONCLUSION: Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus.

Algorithms↗

IRIS: a database surveying known human immune system genes.

We have compiled an online database of known human defense genes: the Immunogenetic Related Information Source (IRIS). As of October 1, 2004, there are 1562 immune genes recorded in IRIS, representing 7% of the human genome. This resource contains searchable information including chromosomal location, sequence data, and a curated functional annotation for each entry. We used IRIS as a basis for analyzing the composition and characteristics of the immune genome, such as gene clustering, polymorphism, and relationship to disease. High protein sequence similarity correlated inversely with distance between immune genes, consistent with clustering of duplicated loci. We also found that, even though some immune genes exhibit high levels of polymorphism, such as MHC class I, the range of levels of polymorphism in immune genes is similar to that of nonimmune genes. Approximately 20% of immune genes have a known disease association. IRIS is available online at .

Databases, Genetic↗

The nucleotide-sugar transporter family: a phylogenetic approach.

Nucleotide sugar transporters (NST) establish the functional link of membrane transport between the nucleotide sugars synthesized in the cytoplasm and nucleus, and the glycosylation processes that take place in the endoplasmic reticulum (ER) and Golgi apparatus. The aim of the present work was to perform a phylogenetic analysis of 87 bank annotated protein sequences comprising all the NST so far characterized and their homologues retrieved by BLAST searches, as well as the closely related triose-phosphate translocator (TPT) plant family. NST were classified in three comprehensive families by linking them to the available experimental data. This enabled us to point out both the possible ER subcellular targeting of these transporters mediated by the dy-lysine motif and the substrate recognition mechanisms specific to each family as well as an important acceptor site motif, establishing the role of evolution in the functional properties of each NST family.

Amino Acid Sequence↗

Analysis of the expressed genome of the lone star tick, Amblyomma americanum (Acari: Ixodidae) using an expressed sequence tag approach.

An expressed sequence tag (EST) approach was used to study the genome of two developmental stages of the lone star tick, Amblyomma americanum. cDNA libraries were constructed from the larval and adult stages of A. americanum. In total, 1942 ESTs were sequenced (1462 adult ESTs and 480 larval ESTs) and analyzed using bioinformatic programs. Contig assembly using the CAPII program revealed 11% and 15% redundancy of sequences in the larval and adult ESTs, respectively. Of the 1942 ESTs, 1738 sequences were considered quality sequences and of these, 771 or approximately 44.4% of the sequences were putatively identified based on amino acid identity using the protein Basic Local Alignment Search Tool (BLAST) algorithm. Putatively identified sequences were classified according to their predicted gene function. In total, 967 sequences, or 55.6% of the quality sequences, had limited or no protein similarity to previously identified gene products. Sequences lacking protein homology were analyzed using an automated sequence annotation system for predicted protein characteristics such as open reading frames, signal peptides, protein motifs, and transmembrane regions. In this paper we describe the sequencing of the largest number of ESTs obtained from an arachnid species to date and the subsequent detailed analysis of these sequences.

Algorithms↗

Protein classification using ontology classification.

MOTIVATION: The classification of proteins expressed by an organism is an important step in understanding the molecular biology of that organism. Traditionally, this classification has been performed by human experts. Human knowledge can recognise the functional properties that are sufficient to place an individual gene product into a particular protein family group. Automation of this task usually fails to meet the 'gold standard' of the human annotator because of the difficult recognition stage. The growing number of genomes, the rapid changes in knowledge and the central role of classification in the annotation process, however, motivates the need to automate this process. RESULTS: We capture human understanding of how to recognise members of the protein phosphatases family by domain architecture as an ontology. By describing protein instances in terms of the domains they contain, it is possible to use description logic reasoners and our ontology to assign those proteins to a protein family class. We have tested our system on classifying the protein phosphatases of the human and Aspergillus fumigatus genomes and found that our knowledge-based, automatic classification matches, and sometimes surpasses, that of the human annotators. We have made the classification process fast and reproducible and, where appropriate knowledge is available, the method can potentially be generalised for use with any protein family. AVAILABILITY: All components described in this paper are freely available. OWL ontology http://www.bioinf.man.ac.uk/phosphabase myGrid http://www.mygrid.org.uk Instance Store http://instancestore.man.ac.uk.

Algorithms↗

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Arabidopsis thaliana is the most widely-studied plant today. The concerted efforts of over 11 000 researchers and 4000 organizations around the world are generating a rich diversity and quantity of information and materials. This information is made available through a comprehensive on-line resource called the Arabidopsis Information Resource (TAIR) (http://arabidopsis.org), which is accessible via commonly used web browsers and can be searched and downloaded in a number of ways. In the last two years, efforts have been focused on increasing data content and diversity, functionally annotating genes and gene products with controlled vocabularies, and improving data retrieval, analysis and visualization tools. New information include sequence polymorphisms including alleles, germplasms and phenotypes, Gene Ontology annotations, gene families, protein information, metabolic pathways, gene expression data from microarray experiments and seed and DNA stocks. New data visualization and analysis tools include SeqViewer, which interactively displays the genome from the whole chromosome down to 10 kb of nucleotide sequence and AraCyc, a metabolic pathway database and map tool that allows overlaying expression data onto the pathway diagrams. Finally, we have recently incorporated seed and DNA stock information from the Arabidopsis Biological Resource Center (ABRC) and implemented a shopping-cart style on-line ordering system.

Arabidopsis↗