Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.

Most Internet online resources for investigating HIV biology contain either bioinformatics tools, protein information or sequence data. The objective of this study was to develop a comprehensive online proteomics resource that integrates bioinformatics with the latest information on HIV-1 protein structure, gene expression, post-transcriptional/post-translational modification, functional activity, and protein-macromolecule interactions. The BioAfrica HIV-1 Proteomics Resource http://bioafrica.mrc.ac.za/proteomics/index.html is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites. The HIV-1 Protein Data-mining Tool includes a set of 27 group M (subtypes A through K) reference sequences that can be used to assess the influence of genetic variation on immunological and functional domains of the protein. The BLAST Structure Tool identifies proteins with similar, experimentally determined topologies, and the Tools Directory provides a categorized list of websites and relevant software programs. This combined database and software repository is designed to facilitate the capture, retrieval and analysis of HIV-1 protein data, and to convert it into clinically useful information relating to the pathogenesis, transmission and therapeutic response of different HIV-1 variants. The HIV-1 Proteomics Resource is readily accessible through the BioAfrica website at: http://bioafrica.mrc.ac.za/proteomics/index.html.

Africa↗

Building dictionaries of 1D and 3D motifs by mining the Unaligned 1D sequences of 17 archaeal and bacterial genomes.

We have used the Teiresias algorithm to carry out unsupervised pattern discovery in a database containing the unaligned ORFs from the 17 publicly available complete archaeal and bacterial genomes and build a 1D dictionary of motifs. These motifs which we refer to as seqlets account for and cover 97.88% of this genomic input at the level of amino acid positions. Each of the seqlets in this 1D dictionary was located among the sequences in Release 38.0 of the Protein Data Bank and the structural fragments corresponding to each seqlet's instances were identified and aligned in three dimensions: those of the seqlets that resulted in RMSD errors below a pre-selected threshold of 2.5 Angstroms were entered in a 3D dictionary of structurally conserved seqlets. These two dictionaries can be thought of as cross-indices that facilitate the tackling of tasks such as automated functional annotation of genomic sequences, local homology identification, local structure characterization, comparative genomics, etc.

Algorithms↗

PROGEN: an automated modelling algorithm for the generation of complete protein structures from the alpha-carbon atomic coordinates.

A modelling algorithm (PROGEN) for the generation of complete protein atomic coordinates from only the alpha-carbon coordinates is described. PROGEN utilizes an optimal geometry parameter (OGP) database for the positioning of atoms for each amino acid of the polypeptide model. The OGP database was established by examining the statistical correlations between 23 different intra-peptide and inter-peptide geometric parameters relative to the alpha-carbon distances for each amino acid in a library of 19 known proteins from the Brookhaven Protein Database (BPDB). The OGP files for specific amino acids and peptides were used to generate the atomic positions, with respect to alpha-carbons, for main-chain and side-chain atoms in the modelled structure. Refinement of the initial model was accomplished using energy minimization (EM) and molecular dynamics techniques. PROGEN was tested using 60 known proteins in the BPDB, representing a wide spectrum of primary and secondary structures. Comparison between PROGEN models and BPDB crystal reference structures gave r.m.s.d. values for peptide main-chain atoms between 0.29 and 0.76 A, with a grand average of 0.53 A for all 60 models. The r.m.s.d. for all non-hydrogen atoms ranged between 1.44 and 1.93 A for the 60 polypeptide models. PROGEN was also able to make the correct assignment of cis- or trans-proline configurations in the protein structures examined. PROGEN offers a fully automatic building and refinement procedure and requires no special or specific structural considerations for the protein to be modelled.

Algorithms↗

CYGD: the Comprehensive Yeast Genome Database.

The Comprehensive Yeast Genome Database (CYGD) compiles a comprehensive data resource for information on the cellular functions of the yeast Saccharomyces cerevisiae and related species, chosen as the best understood model organism for eukaryotes. The database serves as a common resource generated by a European consortium, going beyond the provision of sequence information and functional annotations on individual genes and proteins. In addition, it provides information on the physical and functional interactions among proteins as well as other genetic elements. These cellular networks include metabolic and regulatory pathways, signal transduction and transport processes as well as co-regulated gene clusters. As more yeast genomes are published, their annotation becomes greatly facilitated using S.cerevisiae as a reference. CYGD provides a way of exploring related genomes with the aid of the S.cerevisiae genome as a backbone and SIMAP, the Similarity Matrix of Proteins. The comprehensive resource is available under http://mips.gsf.de/genre/proj/yeast/.

Binding Sites↗

Antiviral and antitumor peptides from insects.

Insects can rapidly clear microbial infections by producing a variety of immune-induced molecules including antibacterial and/or antifungal peptides/polypeptides. In this report, we present the isolation, structural characterization, and biological properties of two variants of a group of bioactive, slightly cationic peptides, referred to as alloferons. Two peptides were isolated from the blood of an experimentally infected insect, the blow fly Calliphora vicina (Diptera), with the following amino acid sequences: HGVSGHGQHGVHG (alloferon 1) and GVSGHGQHGVHG (alloferon 2). Although these peptides have no clear homologies with known immune response modifiers, protein database searches established some structural similarities with proteins containing amino acid stretches similar to alloferon. In vitro experiments reveal that the synthetic version of alloferon has stimulatory activities on natural killer lymphocytes, whereas in vivo trials indicate induction of IFN production in mice after treatments with synthetic alloferon. Additional in vivo experiments in mice indicate that alloferon has antiviral and antitumoral capabilities. Taken together, these results suggest that this peptide, which has immunomodulatory properties, may have therapeutic capacities. The fact that insects may produce cytokine-like materials modulating basic mechanisms for human immunity suggests a source of anti-infection and antitumoral biopharmaceuticals.

Amino Acid Sequence↗

Entrez Gene: gene-centered information at NCBI.

Entrez Gene (www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=gene) is NCBI's database for gene-specific information. It does not include all known or predicted genes; instead Entrez Gene focuses on the genomes that have been completely sequenced, that have an active research community to contribute gene-specific information, or that are scheduled for intense sequence analysis. The content of Entrez Gene represents the result of curation and automated integration of data from NCBI's Reference Sequence project (RefSeq), from collaborating model organism databases, and from many other databases available from NCBI. Records are assigned unique, stable and tracked integers as identifiers. The content (nomenclature, map location, gene products and their attributes, markers, phenotypes, and links to citations, sequences, variation details, maps, expression, homologs, protein domains and external databases) is updated as new information becomes available. Entrez Gene is a step forward from NCBI's LocusLink, with both a major increase in taxonomic scope and improved access through the many tools associated with NCBI Entrez.

Databases, Genetic↗

Different amino acid substitutions at the same position in rhodopsin lead to distinct phenotypes.

PURPOSE: Identification of a novel rhodopsin mutation in a family with retinitis pigmentosa and comparison of the clinical phenotype to a known mutation at the same amino acid position. METHODS: Screening for mutations in rhodopsin was performed in 78 patients with retinitis pigmentosa. All exons and flanking intronic regions were amplified by PCR, sequenced, and compared to the reference sequence derived from the National Center for Biotechnology Information (NCBI, Bethesda, MD) database. Patients were characterized clinically according to the results of best corrected visual acuity testing (BCVA), slit lamp examination (SLE), funduscopy, Goldmann perimetry (GP), dark adaptometry (DA), and electroretinography (ERG). Structural analyses of the rhodopsin protein were performed with the Swiss-Pdb Viewer program available on-line (http://www.expasy.org.spdvbv/ provided in the public domain by Swiss Institute of Bioinformatics, Geneva, Switzerland). RESULTS: A novel rhodopsin mutation (Gly90Val) was identified in a Swiss family of three generations. The pedigree indicated autosomal dominant inheritance. No additional mutation was found in this family in other autosomal dominant genes. The BCVA of affected family members ranged from 20/25 to 20/20. Fundus examination showed fine pigment mottling in patients of the third generation and well-defined bone spicules in patients of the second generation. GP showed concentric constriction. DA demonstrated monophasic cone adaptation only. ERG revealed severely reduced rod and cone signals. The clinical picture is compatible with retinitis pigmentosa. A previously reported amino acid substitution at the same position in rhodopsin leads to a phenotype resembling night blindness in mutation carriers, whereas patients reported in the current study showed the classic retinitis pigmentosa phenotype. The effect of different amino acid substitutions on the three-dimensional structure of rhodopsin was analyzed by homology modeling. Distinct distortions of position 90 (shifts in amino acids 112 and 113) and additional hydrogen bonds were found. CONCLUSIONS: Different amino acid substitutions at position 90 of rhodopsin can lead to night blindness or retinitis pigmentosa. The data suggest that the property of the substituted amino acid distinguishes between the phenotypes.

Adult↗

Study of transcripts from AC010088, a 199,485 bp fragment of the human Y chromosome located in the azoospermia factor region c.

Deletions on the long arm of the human Y chromosome are associated with male infertility. In this work, we studied transcripts of a 199,485 bp long fragment of the Yq11 region (GenBank accession number, AC010088) located in the AZFc (azoospermia factor region c), and characterized their gene structures. After masking repetitive elements, we searched human mRNA Refseqs (reference sequences), a dbEST (database of expressed sequence tags) and a non-redundant nucleic acid database for the mRNAs and ESTs corresponding to the AC010088 using the BLAST programs at the NCBI (National Center for Biotechnology Information) site. Our findings are summarized as follows: i) BPY2 (testis basic protein on Y, 2), DAZ1 (deleted in azoospermia 1), TTY4 (testis transcript Y 4) mRNAs and 23 ESTs were found; ii) Eighteen of 23 ESTs were transcripts of the DAZ gene(s), one EST was a transcript of TTY4 gene, and the remaining 4 probably corresponded to 4 different pseudogenes; iii) DAZ gene(s) were expressed not only in testis, but also in lung carcinoma cells, stomach and Ewing's sarcoma cells; iv) beta-satellite clusters were present around and within the BPY2 and TTY4 gene region; v) In this study, TTY4, BPY2 and DAZ1 genes were mapped precisely to the AC010088 region.

Chromosome Mapping↗

The phylogenetic position of the Theileria buffeli group in relation to other Theileria species.

Theileria parasites known as either T. buffeli, T. orientalis or T. sergenti share many characteristics and are referred to as the T. buffeli group. The 18S ribosomal RNA and merozoite-piroplasm surface protein-encoding genes display a wider genetic variation within the T. buffeli group than between well defined species like T. annulata and T. parva. Analysis of 18S rRNA gene sequences from the database showed that similar groups of related Theileria parasites occur in wild ruminants in Japan and the USA. Moreover, a recently discovered Theileria species pathogenic for small ruminants in China is phylogenetically related to these Japanese Theileria parasites. Phylogenetic analysis of all Theileria 18S rRNA genes reported thus far allowed a subdivision into eight clusters. Some of these clusters contain multiple different species, whereas others appear to contain parasites with similar biological properties whose true speciation remains at present unresolved. The consequences for 18S rRNA based diagnostic assays is discussed.

Animals↗

Auto-transporter A protein of Neisseria meningitidis: a potent CD4+ T-cell and B-cell stimulating antigen detected by expression cloning.

A meningococcal genomic expression library was screened for potent CD4+ T-cell antigens, using patients' peripheral blood lymphocytes (PBLs). One of the most promising positive clones was fully characterized. The recombinant meningococcal DNA contained a single, incomplete, open reading frame (ORF), which was fully reconstructed with reference to available genomic sequence data. The gene was designated autA (auto-transporter A) as its peptide sequence shares molecular characteristics of the auto-transporter family of proteins. Only a single copy of this gene was detected in the meningococcal, and none in the gonococcal, genomic sequence databases. The complete autA gene, when cloned into an expression vector, expressed a protein of approximately 68 kDa. Purified rAutA recalled strong secondary T-cell responses in PBLs of patients and some healthy donors, and induced strong primary T-cell responses in healthy donors. The human B-cell immunogenicity and cross-reactivity of AutA, purified under native conditions, was confirmed in dot immunoblot experiments. Immunoblots with rabbit polyclonal antibodies to rAutA demonstrated the conserved nature, antigenicity and cross-reactivity of AutA amongst meningococci of different serogroups and strains representing different hypervirulent lineages. AutA showed homology with another meningococcal and gonococcal ORF (designated AutB). AutB was cloned and expressed and used to raise an autB-specific antiserum. Immunoblot experiments indicated that AutB is not expressed in meningococci and does not cross-react with AutA. Thus, AutA, being a potent CD4+ T-cell and B-cell-stimulating antigen, which is highly conserved, deserves further investigation as a potential vaccine candidate.

Adolescent↗

Entrez Gene: gene-centered information at NCBI.

Entrez Gene (www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=gene) is NCBI's database for gene-specific information. Entrez Gene includes records from genomes that have been completely sequenced, that have an active research community to contribute gene-specific information or that are scheduled for intense sequence analysis. The content of Entrez Gene represents the result of both curation and automated integration of data from NCBI's Reference Sequence project (RefSeq), from collaborating model organism databases and from other databases within NCBI. Records in Entrez Gene are assigned unique, stable and tracked integers as identifiers. The content (nomenclature, map location, gene products and their attributes, markers, phenotypes and links to citations, sequences, variation details, maps, expression, homologs, protein domains and external databases) is provided via interactive browsing through NCBI's Entrez system, via NCBI's Entrez programing utilities (E-Utilities), and for bulk transfer by ftp.

Databases, Genetic↗

Shotgun proteomics: tools for the analysis of complex biological systems.

Recent interest in proteomics has been fueled by the completion of multiple genome projects and ignited by the common need of biologists to rapidly and comprehensively evaluate complex samples of proteins on a global level. 'Shotgun proteomics' refers to the direct analysis of complex protein mixtures to rapidly generate a global profile of the protein complement within the mixture. This approach has been facilitated by the use of multidimensional protein identification technology (MudPIT), which incorporates multidimensional high-pressure liquid chromatography (LC/LC), tandem mass spectrometry (MS/MS) and database-searching algorithms. This review will focus on the most recent advances in methodologies for shotgun proteomics and address the limitations of the application of each to real biological samples.

Algorithms↗

The XTH family of enzymes involved in xyloglucan endotransglucosylation and endohydrolysis: current perspectives and a new unifying nomenclature.

The polysaccharide xyloglucan is thought to play an important structural role in the primary cell wall of dicotyledons. Accordingly, there is considerable interest in understanding the biochemical basis and regulation of xyloglucan metabolism, and research over the last 16 years has identified a large family of cell wall proteins that specifically catalyze xyloglucan endohydrolysis and/or endotransglucosylation. However, a confusing and contradictory series of nomenclatures has emerged in the literature, of which xyloglucan endotransglycosylases (XETs) and endoxyloglucan transferases (EXGTs) are just two examples, to describe members of essentially the same class of genes/proteins. The completion of the first plant genome sequencing projects has revealed the full extent of this gene family and so this is an opportune time to resolve the many discrepancies in the database that include different names being assigned to the same gene. Following consultation with members of the scientific community involved in plant cell wall research, we propose a new unifying nomenclature that conveys an accurate description of the spectrum of biochemical activities that cumulative research has shown are catalyzed by these enzymes. Thus, a member of this class of genes/proteins will be referred to as a xyloglucan endotransglucosylase/hydrolase (XTH). The two known activities of XTH proteins are referred to enzymologically as xyloglucan endotransglucosylase (XET, which is hereby re-defined) activity and xyloglucan endohydrolase (XEH) activity. This review provides a summary of the biochemical and functional diversity of XTHs, including an overview of the structure and organization of the Arabidopsis XTH gene family, and highlights the potentially important roles that XTHs appear to play in numerous examples of plant growth and development.

Arabidopsis↗

Identification of garlic in old gildings by gas chromatography-mass spectrometry.

The proteinaceous content of garlic (Allium sativum) was characterised according to its amino acid composition by using a gas chromatography-mass spectrometry (GC-MS) analytical procedure. The procedure was tested on fresh and aged garlic samples as well as on reference gilding specimens prepared according to old recipes. The proteinaceous pattern showed a characteristic distribution of amino acids with glutamic acid being the major component. The average amino acidic composition was: glutamic acid (Glu; 29%), aspartic acid (Asp; 17%), serine (Ser; 11%), alanine, glycine, valine, leucine, lysine and phenylalanine (Ala, Gly, Val, Leu, Lys and Phe; 5-6%), isoleucine, proline and tyrosine (Ile, Pro and Tyr; 2-3%), methionine and hydroxyproline (Met and Hyp; 0.5%). In order to distinguish this material from animal glue and egg, which are the other proteinaceous media commonly used in gilding techniques, a database of amino acid percentages of the three proteins was built up and submitted to principal component analysis. Three separate clusters were obtained, allowing the protein identification. The application of the procedure on several gilding samples from Italian wall and easel paintings (13th-17th century) permitted to evidence the use of garlic as a gluing agent.

Garlic↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

Cell wall proteomics of the green alga Haematococcus pluvialis (Chlorophyceae).

The green microalga Haematococcus pluvialis can synthesize and accumulate large amounts of the ketocarotenoid astaxanthin, and undergo profound changes in cell wall composition and architecture during the cell cycle and in response to environmental stresses. In this study, cell wall proteins (CWPs) of H. pluvialis were systematically analyzed by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) coupled with peptide mass fingerprinting (PMF) and sequence-database analysis. In total, 163 protein bands were analyzed, which resulted in positive identification of 81 protein orthologues. The highly complex and dynamic composition of CWPs is manifested by the fact that the majority of identified CWPs are differentially expressed at specific stages of the cell cycle along with a number of common wall-associated 'housekeeping' proteins. The detection of cellulose synthase orthologue in the vegetative cells suggested that the biosynthesis of cellulose occurred during primary wall formation, in contrast to earlier observations that cellulose was exclusively present in the secondary wall of the organism. A transient accumulation of a putative cytokinin oxidase at the early stage of encystment pointed to a possible role in cytokinin degradation while facilitating secondary wall formation and/or assisting in cell expansion. This work represents the first attempt to use a proteomic approach to investigate CWPs of microalgae. The reference protein map constructed and the specific protein markers obtained from this study provide a framework for future characterization of the expression and physiological functions of the proteins involved in the biogenesis and modifications in the cell wall of Haematococcus and related organisms.

Cell Cycle↗

Eight cDNA encoding putative aquaporins in Vitis hybrid Richter-110 and their differential expression.

The nucleotide sequences of eight cDNAs encoding putative aquaporins obtained from a leaf Vitis hybrid Richter-110 cDNA library are reported. They encode proteins ranging from 249 to 287 amino acids with characteristic sequences that clearly include them within the MIP family. According to available database sequence homologies, they can be classified into four groups belonging to two subfamilies: PIP (PIP1 and PIP2) and TIP (gamma-TIP and delta-TIP). In order to elucidate the expression patterns of these putative aquaporins in the plant, specific probes were developed and tissue specific differential expression was tested by reverse Northern and compared with two reference genes (malic enzyme and glutamate dehydrogenase). Clearly, most of the putative aquaporins had higher expression in roots, whereas expression in shoot and leaves was generally weaker than the reference genes.

Aquaporins↗

Column chromatographic prefractionation leads to the detection of 543 different gene products in human fetal brain.

In a previous publication a large series of proteins were identified in fetal human brain by the use of two-dimensional electrophoresis (2-DE) with subsequent matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) and MALDI-tandem time-of-flight (TOF/TOF) analysis. Further identification of many more different spots by traditional 2-DE without additional step such as narrow immobilized ph gradient (IPG) strips or prefractionation seems unlikely and we therefore decided to separate extracted brain proteins by ion-exchange chromatography using a TSK gel DEAE-5PW column followed by 2-DE of individual fractions and analysis by MALDI-TOF/TOF with LIFT technology in fetal brain of the early second trimester. About 1880 protein spots corresponding to 543 different gene products were identified. These proteins included housekeeping, signaling, cytoskeletal, metabolic, antioxidant, and neuron/synaptosomal specific proteins. Among these, 314 gene products (314/543, 57.8%), which have never been detected in traditional 2-DE of human fetal brain, were observed by this method. This updated map of fetal brain proteins may serve as data base and reference map for fetal brain proteins, and the methodology applied may be used as a valuable analytical tool for the basis of protein expressional studies in health and disease.

Brain↗