A prescription which predicts functionally equivalent residues at given sites in protein sequences.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Present knowledge of human immunodeficiency virus type 1 (HIV-1) envelope immunobiology has been derived almost exclusively from analyses of subtype B viruses, yet such viruses represent only a minority of strains currently spreading worldwide. To generate a more representative panel of genetically diverse envelope genes, we PCR amplified, cloned, and sequenced complete gp160 coding regions of 35 primary (peripheral blood mononuclear cell-propagated) HIV-1 isolates collected at major epicenters of the current AIDS pandemic. Analysis of their deduced amino acid sequences revealed several important differences from prototypic subtype B strains, including changes in the number and distribution of cysteine residues, substantial length differences in hypervariable regions, and premature truncations in the gp41 domain. Moreover, transiently expressed glycoprotein precursor molecules varied considerably in both size and carbohydrate content. Phylogenetic analyses of full-length env sequences indicated that the panel included members of all major sequence subtypes of HIV-1 group M (clades A to G), as well as an intersubtype recombinant (F/B) from an infected individual in Brazil. In addition, all subtype E and three subtype G viruses initially classified on the basis of partial env sequences were found to cluster in subtype A in the 3' half of their gp41 coding region, suggesting that they are also recombinant. The biological activity of PCR-derived env genes was examined in a single-round virus infectivity assay. This analysis identified 20 clones, including 1 from each subtype (or recombinant), which expressed fully functional envelope glycoproteins. One of these, derived from a patient with rapid CD4 cell decline, contained an amino acid substitution in a highly conserved endocytosis signal (Y721C), as mediated virus entry with very poor efficiency, although they did not contain sequence changes predicted to alter protein function. These results indicate that the env genes of primary HIV-1 isolates collected worldwide can vary considerably in their genetic, phylogenetic, and biological properties. The panel of env constructs described here should prove valuable for future structure-function studies of naturally occurring envelope glycoproteins as well as AIDS vaccine development efforts targeted against a broader spectrum of viruses.
Adenovirus E1A and c-myc genes are known to be capable of transforming primary rat cells when they occur in combination with either polyoma middle-T or T24 Harvey-ras1 genes. There was a low level of amino acid sequence homology between the nuclear adenovirus-12 (Ad12)E1A protein product (289 amino acids) and the c-myc protein based on optimal alignment and percentage identity. In contrast to others [Ralston R, Bishop JM (1983) Nature 306:803-806], we concluded that this low level of amino acid sequence homology was not significant, since rabies glycoprotein (RGP), which has no transforming function and localizes to the cell surface, had a similar low level of amino acid sequence homology to the c-myc protein. Furthermore, dot-matrix analysis, when used to test the overall level of amino acid sequence homology, showed no significant homology between c-myc and Ad12E1A, E1B, or RGP. Thus, low levels of amino acid sequence homology between two proteins may not be sufficient to predict structural and functional similarities between them reliably, even if the two proteins appear to share a common function.
There has been increased interest in bacterial polyadenylation with the recent demonstration that 3' poly(A) tails are involved in RNA degradation. Poly(A) polymerase I (PAP I) of Escherichia coli is a member of the nucleotidyltransferase (Ntr) family that includes the functionally related tRNA CCA-adding enzymes. Thirty members of the Ntr family were detected in a search of the current database of eubacterial genomic sequences. Gram-negative organisms from the beta and gamma subdivisions of the purple bacteria have two genes encoding putative Ntr proteins, and it was possible to predict their activities as either PAP or CCA adding by sequence comparisons with the E. coli homologues. Prediction of the functions of proteins encoded by the genes from more distantly related bacteria was not reliable. The Bacillus subtilis papS gene encodes a protein that was predicted to have PAP activity. We have overexpressed and characterized this protein, demonstrating that it is a tRNA nucleotidyltransferase. We suggest that the papS gene should be renamed cca, following the notation for its E. coli counterpart. The available evidence indicates that cca is the only gene encoding an Ntr protein, despite previous suggestions that B. subtilis has a PAP similar to E. coli PAP I. Thus, the activity involved in RNA 3' polyadenylation in the gram-positive bacteria apparently resides in an enzyme distinct from its counterpart in gram-negative bacteria.
A large-scale effort to measure, detect and analyse protein-protein interactions using experimental methods is under way. These include biochemistry such as co-immunoprecipitation or crosslinking, molecular biology such as the two-hybrid system or phage display, and genetics such as unlinked noncomplementing mutant detection. Using the two-hybrid system, an international effort to analyse the complete yeast genome is in progress. Evidently, all these approaches are tedious, labour intensive and inaccurate. From a computational perspective, the question is how can we predict that two proteins interact from structure or sequence alone. Here we present a method that identifies gene-fusion events in complete genomes, solely based on sequence comparison. Because there must be selective pressure for certain genes to be fused over the course of evolution, we are able to predict functional associations of proteins. We show that 215 genes or proteins in the complete genomes of Escherichia coli, Haemophilus influenzae and Methanococcus jannaschii are involved in 64 unique fusion events. The approach is general, and can be applied even to genes of unknown function.
Most of the genes involved in the development of multicellular eukaryotes encode large, multidomain proteins. To decipher the major trends in the evolution of these proteins and make functional predictions for uncharacterized domains, we applied a strategy of sequence database search that includes construction of specialized data sets and iterative subsequence masking. This computational approach allowed us to detect previously unnoticed but potentially important sequence similarities. Developmental gene products are enriched in predicted nonglobular regions as compared to unbiased sets of eukaryotic and bacterial proteins. Developmental genes that act intracellularly, primarily at the level of transcription regulation, typically code for proteins containing highly conserved DNA-binding domains, most of which appear to have evolved before the radiation of bacteria and eukaryotes. We identified bacterial homologues, namely a protein family that includes the Escherichia coli universal stress protein UspA, for the MADS-box transcription regulators previously described only in eukaryotes. We also show that the FUS6 family of eukaryotic proteins contains a putative DNA-binding domain related to bacterial helix-turn-helix transcription regulators. Developmental proteins that act extracellularly are less conserved and often do not have bacterial homologues. Nevertheless, several provocative similarities between different groups of such proteins were detected.
PROBLEM STATEMENT: We have studied the relationships among SWISS-PROT, TrEMBL, and GenBank with two goals. First is to determine whether users can reliably identify those proteins in SWISS-PROT whose functions were determined experimentally, as opposed to proteins whose functions were predicted computationally. If this information was present in reasonable quantities, it would allow researchers to decrease the propagation of incorrect function predictions during sequence annotation, and to assemble training sets for developing the next generation of sequence-analysis algorithms. Second is to assess the consistency between translated GenBank sequences and sequences in SWISS-PROT and TrEMBL. RESULTS: (1) Contrary to claims by the SWISS-PROT authors, we conclude that SWISS-PROT does not identify a significant number of experimentally characterized proteins. (2) SWISS-PROT is more incomplete than we expected in that version 38.0 from July 1999 lacks many proteins from the full genomes of important organisms that were sequenced years earlier. (3) Even if we combine SWISS-PROT and TrEMBL, some sequences from the full genomes are missing from the combined dataset. (4) In many cases, translated GenBank genes do not exactly match the corresponding SWISS-PROT sequences, for reasons that include missing or removed methionines, differing translation start positions, individual amino-acid differences, and inclusion of sequence data from multiple sequencing projects. For example, results show that for Escherichia coli, 80.6% of the proteins in the GenBank entry for the complete genome have identical sequence matches with SWISS-PROT/TrEMBL sequences, 13.4% have exact substring matches, and matches for 4.1% can be found using BLAST search; the remaining 2.0% of E.coli protein sequences (most of which are ORFs) have no clear matches to SWISS-PROT/TrEMBL. Although many of these differences can be explained by the complexity of the DB, and by the curation processes used to create it, the scale of the differences is notable.
A simple and fast free energy scoring function (Fresno) has been developed to predict the binding free energy of peptides to class I major histocompatibility (MHC) proteins. It differs from existing scoring functions mainly by the explicit treatment of ligand desolvation and of unfavorable protein-ligand contacts. Thus, it may be particularly useful in predicting binding affinities from three-dimensional models of protein-ligand complexes. The Fresno function was independently calibrated for two different training sets: (a) five HLA-A0201-peptide structures, which had been determined by X-ray crystallography, and (b) three-dimensional models of 37 H-2K(k)-peptide structures, which had been obtained by knowledge-based homology modeling. For both training sets, a good cross-validated fit to experimental binding free energies was obtained with predictive errors of 3-3.5 kJ/mol. As expected, lipophilic interactions were found to contribute the most to HLA-A0201-peptide interactions, whereas H-bonding predominates in H-2K(k) recognition. Both cross-validated models were afterward used to predict the binding affinity of a test set of 26 peptides to HLA-A0204 (an HLA allele closely related to HLA-A0201) and of a series of 16 peptides to H-2K(k). Predictions were more accurate for HLA-A2-binding peptides as the training set had been built from experimentally determined structures. The average error in predicting the binding free energy of the test peptides was 3.1 kJ/mol. For the homology model-derived equation, the average error in predicting the binding free energy of peptides to K(k) was significantly higher (5.4 kJ/mol) but still very acceptable. The present scoring function is thus able to predict with a good accuracy binding free energies from three-dimensional models, at the condition that the backbone coordinates of the MHC-bound peptide have first been determined with an accuracy of about 1-1.5 A. Furthermore, it may be easily recalibrated for any protein-ligand complex.
We predict a structure of the glutamine amidotransferase subunit (hisH) of imidazole glycerol phosphate synthase (IGPS) which catalyzes the fifth step of the histidine biosynthesis in Escherichia coli. The model is constructed using an energy-based threading program augmented by a multiple sequence to structure profile analysis. In developing our model we identified a conserved core region within hisH and a variable domain which is the likely site of interaction with the synthase subunit (hisF) of IGPS. Information available from structural and functional genomics studies was used to improve the structure prediction, to discuss parallels between histidine biosynthesis and other amino acid and nucleotide metabolic pathways, and to better understand the protein-protein interactions between the hisH and hisF domains of IGPS. This work allows us to develop a preliminary model for the structure of the entire IGPS holoenzyme.
The human amyloid precursor-like protein APLP2 is a highly conserved homologue of a sequence-specific DNA-binding mouse protein with a predicted function in the cell cycle. Somatic cell hybrids segregating human chromosomes were used to assign the APLP2 gene to chromosome 11. Fluorescence in situ hybridization confirmed this assignment and further localized the gene to q23-q25.
The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.
The objectives of this paper are to discuss the structure and genetic content of the genome of herpes simplex virus type 1 (HSV-1), the nature of virus DNA replicative processes, and aspects of the evolution of the virus DNA, in particular those bearing on DNA replication. We are in the late stages of determining the complete sequence of the DNA of HSV-1, which contains about 155,000 base pairs, and thus the treatment is primarily from a viewpoint of DNA sequence and organization. The genome possesses around 75 genes, generally densely arranged and without long range ordering. Introns are present in only a few genes. Protein coding sequences have been predicted, and the functions of the proteins are being pursued by various means, including use of existing genetic and biochemical data, computer based analyses, expression of isolated genes and use of oligopeptide antisera. Many proteins are known to be virion structural components, or to have regulatory roles, or to function in synthesis of virus DNA. Many, however, still lack an assigned function. Two classes of genetic entities necessary for virus DNA replication have been characterized: cis-acting sequences, which include origins of replication and packaging signals, and genes encoding proteins involved in replication. Aside from enzymes of nucleotide metabolism, the latter include DNA polymerase, DNA binding proteins, and five species detected by genetic assays, but of presently unknown functions. Complete genome sequences are now known for the related alphaherpesvirus varicella-zoster virus and for the very distinct gammaherpesvirus Epstein-Barr virus. Comparisons between the three sequences show various homologies, and also several types of divergence and rearrangement, and so allow models to be proposed for possible events in the evolution of present day herpesvirus genomes. Another aspect of genome evolution is seen in the wide range of overall base compositions found in present day herpesvirus DNAs. Finally, certain herpesvirus genes are homologous to non-herpesvirus genes, giving a glimpse of more remote relationships.
Here we present the genomic sequence, with analysis, of a pathogenic fowlpox virus (FPV). The 288-kbp FPV genome consists of a central coding region bounded by identical 9.5-kbp inverted terminal repeats and contains 260 open reading frames, of which 101 exhibit similarity to genes of known function. Comparison of the FPV genome with those of other chordopoxviruses (ChPVs) revealed 65 conserved gene homologues, encoding proteins involved in transcription and mRNA biogenesis, nucleotide metabolism, DNA replication and repair, protein processing, and virion structure. Comparison of the FPV genome with those of other ChPVs revealed extensive genome colinearity which is interrupted in FPV by a translocation and a major inversion, the presence of multiple and in some cases large gene families, and novel cellular homologues. Large numbers of cellular homologues together with 10 multigene families largely account for the marked size difference between the FPV genome (260 to 309 kbp) and other known ChPV genomes (178 to 191 kbp). Predicted proteins with putative functions involving immune evasion included eight natural killer cell receptors, four CC chemokines, three G-protein-coupled receptors, two beta nerve growth factors, transforming growth factor beta, interleukin-18-binding protein, semaphorin, and five serine proteinase inhibitors (serpins). Other potential FPV host range proteins included homologues of those involved in apoptosis (e.g., Bcl-2 protein), cell growth (e.g., epidermal growth factor domain protein), tissue tropism (e.g., ankyrin repeat-containing gene family, N1R/p28 gene family, and a T10 homologue), and avian host range (e.g., a protein present in both fowl adenovirus and Marek's disease virus). The presence of homologues of genes encoding proteins involved in steroid biogenesis (e.g., hydroxysteroid dehydrogenase), antioxidant functions (e.g., glutathione peroxidase), vesicle trafficking (e.g., two alpha-type soluble NSF attachment proteins), and other, unknown conserved cellular processes (e.g., Hal3 domain protein and GSN1/SUR4) suggests that significant modification of host cell function occurs upon viral infection. The presence of a cyclobutane pyrimidine dimer photolyase homologue in FPV suggests the presence of a photoreactivation DNA repair pathway. This diverse complement of genes with likely host range functions in FPV suggests significant viral adaptation to the avian host.
An unreported missense mutation of the ribosomal S6 kinase 2 (RSK2) gene has been identified in two male sibs with a mild form of Coffin-Lowry syndrome (CLS) inherited from their healthy mother. They exhibit transient severe hypotonia, macrocephaly, delay in closure of the fontanelles, normal gait, and mild mental retardation, associated in the first sib with transient autistic behaviour. Some dysmorphic features of CLS (in particular forearm fullness and tapering fingers) and many atypical findings (some of which were reminiscent of FG syndrome) were observed as well. The moderate phenotypic expression of this mutation extends the CLS phenotype to include less severe mental retardation and minor, hitherto unreported signs. The missense mutation identified may be less deleterious than those previously described. As this mutation occurs in a protein domain with no predicted function, it could be responsible for a conformational change affecting the protein catalytic function, since a non-polar amino acid is replaced by a charged residue.
The size of protein sequence database is getting larger each day. One common challenge is to predict protein structures or functions of the sequences in databases. It is easy when a sequence shares direct similarity to a well-characterized protein. If there is no direct similarity, we have to rely on a third sequence or a model as intermediate to link two proteins together. We developed a new model based method, called Bayesian search, as a means to connect two distantly related proteins. We compared this Bayesian search model with pairwise and multiple sequence comparison methods on structural databases using structural similarity as the criteria for relationship. The results show that the Bayesian search can link more distantly related sequence pairs than other methods, collectively and consistently over large protein families. If each query made one error on average against SCOP database PDB40D-B, Bayesian search found 36.5% of related pairs, PSI-Blast found 32.6%, and Smith-Waterman method found 25%. Examples are presented to show that the alignments predicted by the Bayesian search agree well with structural alignments. Also false positives found by Bayesian search at low cutoff values are analyzed.
The generation of an antigen-specific T-cell response requires that the T lymphocyte receive two signals from the antigen presenting cell. The specificity of this response is provided by antigen presented to the T lymphocyte and involves stimulation of the T lymphocyte via the T-cell receptor (TCR)/CD3 complex. The second, or costimulatory signal, can be provided by ligation of the B-lymphocyte activation antigens B7-1 (CD80) and B7.2 (CD86) to TCR antigen CD28. The cDNAs for both CD80 and CD86 have been isolated and are predicted to encode type 1 membrane proteins of the immunoglobulin (Ig) superfamily. The predicted protein is composed of a signal peptide followed by two Ig-like extracellular domains, a transmembrane domain, and a cytoplasmic tail. Here we report that the genomic organization of CD86 reflects its functional structure, and is similar to that found for CD80. The gene is composed of eight exons which span more than 22 kilobases. The predicted protein functional domains of signal peptide, extracellular IgV- and IgC-like regions, and transmembrane domain coincide with the genomic structure. Two independent sequences had been reported for CD86 cDNA which differed in their 5'untranslated (UT) regions. We find CD86 exons 1 and 2 correspond to these alternate 5'UT sequences. Splicing of exon 1 or 2 with the signal peptide encoding exon 3 would produce mRNA transcripts complementary to the reported cDNA clones. Exons 4 and 5 correspond to IgV- and IgC-like extracellular domains, respectively. Exon 6 encodes the transmembrane region and beginning of the cytoplasmic tail. Exons 7 and 8 encode the remainder of the cytoplasmic tail and 3'UT sequences.
It was proposed by Bernad et al. [Cell 59 (1989) 219-228] and Blanco et al. [Gene 100 (1991) 27-38] that the 3'----5' exonuclease (Exo) domain of Escherichia coli DNA polymerase I (PolI) is structurally and functionally conserved among prokaryotic and eukaryotic DNA polymerases. The basis for this claim is the presence of three short peptide sequences in many DNA polymerases that resemble PolI sequences that have been shown by x-ray crystallographic and genetic engineering studies to be metal ion binding sites that are essential for PolI 3'----5' Exo activity [Derbyshire et al., Science 240 (1988) 199-201]. This claim is made even though there is little amino acid (aa) sequence similarity between PolI and many eukaryotic and viral DNA polymerases and in spite of significant differences in the amount of 3'----5' Exo activity in the DNA polymerases compared. For at least one DNA polymerase, bacteriophage T4 DNA polymerase, one of the proposed conserved Exo sequences does not appear to be important for 3'----5' Exo activity. This T4 DNA polymerase result provides a reminder that caution must be used when weak aa sequence similarities are used to predict protein structure and function.
Leigh syndrome (LS) associated with cytochrome c oxidase (COX) deficiency is an autosomal recessive neurodegenerative disorder caused by mutations in SURF1. Although SURF1 is ubiquitously expressed, its expression is lower in brain than in other highly aerobic tissues. All reported SURF1 mutations are loss of function, predicting a truncated protein (hSurf1) product. Western blot analysis with anti-hSurf1 antibodies demonstrated a specific 30 kDa protein in control fibroblasts, but no protein in LS patient cells. Steady-state levels of both nuclear- and mitochondrial-encoded COX subunits were also markedly reduced in patient cells, consistent with a failure to assemble or maintain a normal amount of the enzyme complex. An epitope (FLAG)-tagged hSurf1 was targeted to mitochondria in COS7 cells and a mitochondrial import assay showed that the hSurf1 precursor protein (35 kDa) was imported and processed to its mature form (30 kDa) in a membrane potential-dependent fashion. The protein was resistant to alkaline carbonate extraction and susceptible to proteinase K digestion in mitoplasts. Mutant proteins in which the N-terminal transmembrane domain or central loop were deleted, or the C-terminal transmembrane domain disrupted, did not accumulate and could not rescue COX activity in patient cells. Co-expression of the N- and C-terminal transmembrane domains as independent entities also failed to rescue the enzyme deficiency. These data demonstrate that hSurf1 is an integral inner membrane protein with an essential role in the assembly or maintenance of the COX complex and that insertion of both transmembrane domains in the intact protein is necessary for function.