Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “evolutionary conservation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

The oligosaccharyltransferase complex from yeast.

N-Glycosylation of eukaryotic secretory and membrane-bound proteins is an essential and highly conserved protein modification. The key step of this pathway is the en bloc transfer of the high mannose core oligosaccharide Glc3Man9GlcNAc2 from the lipid carrier dolichyl phosphate to selected Asn-X-Ser/Thr sequences of nascent polypeptide chains during their translocation across the endoplasmic reticulum membrane. The reaction is catalysed by the enzyme oligosaccharyltransferase (OST). Recent biochemical and molecular genetic studies in yeast have yielded novel insights into this enzyme with multiple tasks. Nine proteins have been shown to be OST components. These are assembled into a heterooligomeric membrane-bound complex and are required for optimal expression of OST activity in vivo in wild type cells. In accord with the evolutionary conservation of core N-glycosylation, there are significant homologies between the protein sequences of OST subunits from yeast and higher eukaryotes, and OST complexes from different sources show a similar organisation as well.

Binding Sites↗

Original domain for the serum albumin family arose from repeated sequences.

The characteristic three-domain structure has been conserved throughout mammalian evolution by serum albumin and its fetal counterpart, alpha-fetoprotein. Thus, one still detects 35.2% amino acid sequence homology between bovine serum albumin and murine alpha-fetoprotein. Yet, natural selection cannot be invoked as the major factor responsible for the observed conservation of these sequences, for the simple reason of their dispensability. Inherited analbuminemia is apparently a harmless trait in man and the rat. The conservation appears inherent in their repetitious origin. Each protein is made of triplicate copies of the ancestral domain. Furthermore, analysis of the published sequence data suggests that the original coding sequence for the ancestral domain arose as repeats of the 18-base-long primordial building block sequence TTC-ACA-GAG-GAG-CAG-CTG specifying Phe-Thr-Glu-Glu-Gln-Leu and its shorter subsidiary TTC-ATG-GAG-GAG specifying Phe-Met-Glu-Glu. Consequently, the homology between bovine serum albumin and alpha-fetoprotein is mostly confined to small segments still specified by recognizable descendants of these building block sequences. The point to be made here is that evolutionary conservation of coding sequences can be an inherent property; natural selection need not be invoked.

Amino Acid Sequence↗

A functional survey of the enhancer activity of conserved non-coding sequences from vertebrate Iroquois cluster gene deserts.

Recent studies of the genome architecture of vertebrates have uncovered two unforeseen aspects of its organization. First, large regions of the genome, called gene deserts, are devoid of protein-coding sequences and have no obvious biological role. Second, comparative genomics has highlighted the existence of an array of highly conserved non-coding regions (HCNRs) in all vertebrates. Most surprisingly, these structural features are strongly associated with genes that have essential functions during development. Among these, the vertebrate Iroquois (Irx) genes stand out on both fronts. Mammalian Irx genes are organized in two clusters (IrxA and IrxB) that span >1 Mb each with no other genes interspersed. Additionally, a large number of HCNRs exist within Irx clusters. We have systematically examined the enhancer activity of HCNRs from the IrxB cluster using transgenic Xenopus and zebrafish embryos. Most of these HCNRs are active in subdomains of endogenous Irx expression, and some are candidates to contain shared enhancers of neighboring genes, which could explain the evolutionary conservation of Irx clusters. Furthermore, HCNRs present in tetrapod IrxB but not in fish may be responsible for novel Irx expression domains that appeared after their divergence. Finally, we have performed a more detailed analysis on two IrxB ultraconserved non-coding regions (UCRs) duplicated in IrxA clusters in similar relative positions. These four regions share a core region highly conserved among all of them and drive expression in similar domains. However, inter-species conserved sequences surrounding the core, specific for each of these UCRs, are able to modulate their expression.

Animals↗

Novel human HALR (MLL3) gene encodes a protein homologous to ALR and to ALL-1 involved in leukemia, and maps to chromosome 7q36 associated with leukemia and developmental defects.

We have identified and characterized the approximately 12-kb cDNA of a novel human gene (designated HALR for "homologous to ALR" and given the symbol MLL3 by the HUGO Gene Nomenclature Committee) for which open reading frame (ORF) encodes a predicted large hydrophilic nuclear protein comprising 4,025 amino acids with a calculated molecular mass of approximately 443 kD. Within the amino acid sequence of HALR were identified a SUVAR3-9, enhancer of zeste, trithorax (SET) domain, three plant homeodomain (PHD)-type zinc fingers, a high motility group (HMG)-1 box, a leucine-zipper-like pattern, two potential transactivating domains, several nuclear localization signals, and multiple nuclear receptor interaction signature motifs. Especially within the SET domain, PHD fingers and several other regions, the HALR protein exhibits significant similarity to ALR (acute lymphoblastic leukemia [ALL]-1 related), ALL-1/myeloid/lymphoid or mixed-lineage leukemia (ALL-1/MLL), and trithorax, evolutionarily conserved proteins that influence differentiation and development. Northern blot analysis demonstrated transcripts of approximately 11-12 kb, while reverse transcriptase-polymerase chain reaction (RT-PCR) revealed that HALR is expressed in a wide range of human tissues and cancer cell lines. The HALR gene contains 46 exons, is estimated to span >101 kb, and is located on chromosome region 7q36. Terminal 7q deletions are common chromosomal aberrations encountered in hematological neoplasia and in holoprosencephaly 3, a midline embryonic defect involving forebrain development. We have also isolated the partial cDNA of the murine homologue of HALR, which displays high homology to its human counterpart. Taking into consideration its notable protein motifs, ubiquitous expression, evolutionary conservation and chromosomal position, HALR is likely to play a housekeeping role in transcriptional regulation, and may be involved in leukemogenesis and developmental disorders.

Amino Acid Sequence↗

Whole-genome analysis reveals a strong positional bias of conserved dMyc-dependent E-boxes.

Myc is a transcription factor with diverse biological effects ranging from the control of cellular proliferation and growth to the induction of apoptosis. Here we present a comprehensive analysis of the transcriptional targets of the sole Myc ortholog in Drosophila melanogaster, dMyc. We show that the genes that are down-regulated in response to dmyc inhibition are largely identical to those that are up-regulated after dMyc overexpression and that many of them play a role in growth control. The promoter regions of these targets are characterized by the presence of the E-box sequence CACGTG, a known dMyc binding site. Surprisingly, a large subgroup of (functionally related) dMyc targets contains a single E-box located within the first 100 nucleotides after the transcription start site. The relevance of this E-box and its position was confirmed by a mutational analysis of a selected dMyc target and by the observation of its evolutionary conservation in a different Drosophila species, Drosophila pseudoobscura. These observations raise the possibility that a subset of Myc targets share a distinct regulatory mechanism.

Animals↗

A conserved role for the MEK signalling pathway in neural tissue specification and posteriorisation in the invertebrate chordate, the ascidian Ciona intestinalis.

Ascidians are invertebrate chordates with a larval body plan similar to that of vertebrates. The ascidian larval CNS is divided along the anteroposterior axis into sensory vesicle, neck, visceral ganglion and tail nerve cord. The anterior part of the sensory vesicle comes from the a-line animal blastomeres, whereas the remaining CNS is largely derived from the A-line vegetal blastomeres. We have analysed the role of the Ras/MEK/ERK signalling pathway in the formation of the larval CNS in the ascidian, Ciona intestinalis. We show evidence that this pathway is required, during the cleavage stages, for the acquisition of: (1) neural fates in otherwise epidermal cells (in a-line cells); and (2) the posterior identity of tail nerve cord precursors that otherwise adopt a more anterior neural character (in A-line cells). Altogether, the MEK signalling pathway appears to play evolutionary conserved roles in these processes in ascidians and vertebrates, suggesting that this may represent an ancestral chordate strategy.

Animals↗

Mapping and structure of DMXL1, a human homologue of the DmX gene from Drosophila melanogaster coding for a WD repeat protein.

The DmX gene was recently isolated from the X chromosome of Drosophila melanogaster. TBLASTN searches of the dbEST databases revealed sequences with a high level of similarity to DmX in a variety of different species, including insects, nematodes, and mammals showing that DmX is an evolutionarily highly conserved gene. Here we describe the cloning of the cDNA and the chromosomal localization of one of the human homologues of DmX, Dmx-like 1 (DMXL1). The human DMXL1 gene codes for a large mRNA of 11 kb with an open reading frame of 3027 amino acids. The putative protein belongs to the superfamily of WD repeat proteins, which have mostly regulatory functions. The DMXL1 protein contains an exceptionally large number of WD repeat units. The DMXL1 gene is located on chromosome 5q22 as determined by radiation hybrid mapping and fluorescence in situ hybridization. Although the function of the DMXL1 gene and its homologues in other species remains to be discovered, the high level of evolutionary conservation together with the unusual structure suggests that it probably has an important function.

Amino Acid Sequence↗

QuasiMotiFinder: protein annotation by searching for evolutionarily conserved motif-like patterns.

Sequence signature databases such as PROSITE, which include amino acid segments that are indicative of a protein's function, are useful for protein annotation. Lamentably, the annotation is not always accurate. A signature may be falsely detected in a protein that does not carry out the associated function (false positive prediction, FP) or may be overlooked in a protein that does carry out the function (false negative prediction, FN). A new approach has emerged in which a signature is replaced with a sequence profile, calculated based on multiple sequence alignment (MSA) of homologous proteins that share the same function. This approach, which is superior to the simple pattern search, essentially searches with the sequence of the query protein against an MSA library. We suggest here an alternative approach, implemented in the QuasiMotiFinder web server (http://quasimotifinder.tau.ac.il/), which is based on a search with an MSA of homologous query proteins against the original PROSITE signatures. The explicit use of the average evolutionary conservation of the signature in the query proteins significantly reduces the rate of FP prediction compared with the simple pattern search. QuasiMotiFinder also has a reduced rate of FN prediction compared with simple pattern searches, since the traditional search for precise signatures has been replaced by a permissive search for signature-like patterns that are physicochemically similar to known signatures. Overall, QuasiMotiFinder and the profile search are comparable to each other in terms of performance. They are also complementary to each other in that signatures that are falsely detected in (or overlooked by) one may be correctly detected by the other.

Amino Acid Motifs↗

The product of the natural reaction catalyzed by 4-oxalocrotonate tautomerase becomes an affinity label of its mutant.

4-Oxalocrotonate tautomerase (4-OT) catalyzes the isomerization of 4-oxalocrotonate, 1, to 2-oxo-3E-hexenedioate, 3, using a general acid/base mechanism that involves a conserved N-terminal proline residue. The P1A and P1G mutants have been shown to catalyze this isomerization but at reduced rates. Analysis of these mutants by mass spectrometry demonstrated that P1A is susceptible to a 1,4-addition of the N-terminal primary amine across the double bond of enone 3 to form a covalent adduct. Although slower than the isomerization reaction, the addition is fast, with 50% of the active sites being alkylated within 12 min. By contrast, the wt4-OT shows no detectable modification over 24 h. These results support the hypothesis that avoidance of nucleophilic reactions, such as the irreversible Michael addition to the product, could be a contributing factor in the evolutionary conservation of N-terminal proline residues in 4OT.

Affinity Labels↗

Characterization of a monoclonal antibody directed against a sulphoglycolipid that is evolutionarily conserved and developmentally regulated in rat brain.

Monoclonal antibodies (MABs) have been raised against acidic glycolipids extracted from the electric organ of Torpedo marmorata. One of these, designated L9, appears to recognize acidic glycolipids in adult T. marmorata electric organ, electromotor nerves and brain, adult rat sciatic nerve, and in embryonic and neonatal rat brain, starting at embryonic day (ED) 15 and disappearing by the 20th day of post-natal life. The epitope is present in growth cones isolated from 4-day-old rats; its proportion relative to total gangliosides is, however, no higher than that found in whole neonatal brain membranes. Desialidation of the acidic glycolipid fraction modifies neither the immunoreactivity nor the RF value following thin-layer chromatography (TLC) of the antigen; it is concluded that the antigen is not a ganglioside. The MAB, HNK-1, recognizes the L9 antigen. Both HNK-1 and L9 recognize a sulphoglycolipid of the same RF in TLC. The function of the L9 antigen is not known but its evolutionary conservation, presence in growth cones and its developmental regulation in the mammalian central nervous system indicate that it plays an important role in nervous system maturation.

Animals↗

The nucleotide sequence of rabbit embryonic globin gene beta 3.

The nucleotide sequence of a rabbit embryonic globin gene, beta 3, has been determined from 161 base pairs (bp) on the 5' side of the mRNA cap site to 209 base pairs beyond the 3' poly A addition site. The 5' and 3' ends of mRNA from both embryonic globin genes beta 3 and beta 4 have been determined by an S1 protection assay. Sequences that are highly conserved in the 5' flanking region of eukaryotic structural genes, AATAAAA and CCAAT, are located -25 to -31 nucleotides and -81 to -85 nucleotides, respectively, before the cap site. The CCAAT sequence is duplicated at -108 to -112 nucleotides, as it is in the human fetal gamma-globin genes. Small (124 bp) and large (817 bp) intervening sequences are located between codons 30 and 31 and between 104 and 105, respectively. The sequence AATAAA precedes the predominant poly(A) addition site by 19 nucleotides. Although rabbit globin gene beta 3 is transcribed and translated almost exclusively in embryonic erythrocytes, it shares striking homology with the human gamma-globin genes which are expressed in erythrocytes from fetal liver. The evolutionary conservation of rabbit beta 3 and human gamma correlates well with their similar chromosomal positions in the two genes families.

Amino Acid Sequence↗

Isolation of DICE1: a gene frequently affected by LOH and downregulated in lung carcinomas.

In the development and progression of sporadic tumors multiple tumor suppressor genes are inactivated that may be distinct from predisposing cancer genes. Previously, a tumor suppressor locus on human chromosome 13q14 that is distinct from the retinoblastoma predisposing gene 1 (RB1) has been identified in lung, head and neck, breast, ovarian and prostate tumors. By an approach that combines genomic difference cloning and positional cloning we isolated the cDNA of a novel gene (DICE1) located at 13q14.12-14.2. The DICE1 gene is highly conserved in evolution and its mRNA is expressed in a wide variety of fetal and adult tissues. The DICE1 cDNA encodes a predicted protein of 887 amino acids corresponding to an 100 kD protein that shows 92.9% identity to the carboxy-terminal half of the mouse EGF repeat transmembrane protein DBI-1. The DBI-1 protein interferes with the mitogenic response to insulin-like growth factor 1 (IGF-I) and is presumably involved in anchorage-dependent growth. When compared to normal lung tissue expression of the DICE1 mRNA was reduced or undetectable in the majority of non-small cell lung carcinomas analysed. The location of the DICE1 gene in the region of allelic loss, its high evolutionary conservation and the downregulation of expression in carcinoma cells suggests that DICE1 is a candidate tumor suppressor gene in non-small cell lung carcinomas and possibly in other sporadic carcinomas.

3T3 Cells↗

Structural and functional characterization of two mutated R2 proteins of Escherichia coli ribonucleotide reductase.

The R2 protein of ribonucleotide reductase from Escherichia coli is a homodimeric tyrosyl-radical-containing enzyme with two identical dinuclear iron centers. Two randomly generated genomic mutants, nrdB-1 and nrdB-2, that produce R2 enzymes with low enzymatic activity, have been cloned and characterized to identify functionally important residues and areas of the enzyme. The mutations were identified as Pro348 to leucine in nrdB-1 and Leu304 to phenylalanine in nrdB-2. Both mutations are the results of single amino acid replacements of non-conserved residues. The three-dimensional structures of [L348]R2 and [F304]R2 have been determined to 0.26-nm and 0.28-nm resolution, respectively. Compared with wild-type R2, [L348]R2 binds with higher affinity to R1, probably due to increased flexibility of its C-terminus. Since the three-dimensional structure, iron-center properties and radical properties of [L348]R2 are comparable to those of wild-type R2, the low catalytic activity of the holoenzyme is probably caused by a perturbed interaction between R2 and R1. The [F304]R2 enzyme has increased radical sensitivity and low catalytic activity compared with wild-type R2. In [F304]R2 the only significant change in structure is that the evolutionary conserved Ser211 forms a different hydrogen bond to a distorted helix. The results obtained with [F304]R2 indicate that structural changes in E. coli R2 in the vicinity of this helix distortion can influence the catalytic activity of the holoenzyme.

Bacterial Proteins↗

Exploring the interface between the N- and C-terminal helices of cytochrome c by random mutagenesis within the C-terminal helix.

Buried within cytochrome c lies a highly-conserved helix-helix interface formed by the perpendicular packing of the C-terminal helix against the N-terminal helix. This interface involves a peg-in-hole interaction between Gly-6 and Leu-94 and an aromatic-aromatic interaction between Phe-10 and Tyr-97. To gain insight into protein design, we investigated the relationship between the sequence of the interface and the physiological function of yeast iso-1-cytochrome c. A library of mutants at positions 94 and 97 of the C-terminal helix was created to examine the effect of novel amino acid combinations. We isolated 45 of the 400 possible amino acid combinations, 32 of which result in a functional cytochrome c. Contrary to evolutionary conservation of the peg-in-hole and aromatic-aromatic interactions, we find that side-chain volume and conservation of aromatic residues do not play an essential role in determining function. Additionally, we find negatively-charged residues within the interface that result in a functional cytochrome c. Examination of the 45 missense mutants indicates that approximately 120 unique combinations are compatible with function. These results show that the interface is flexible. However, truncation of the C-terminal helix at position 94 abolishes function, suggesting that the interface is essential. The correlation observed between our library of mutants and the mutation matrix compiled by Gonnet et al. [Gonnet, G. H., Cohen, M. A., & Benner, S. A. (1992) Science 256, 1443-1445] demonstrates the potential use of the matrix to predict the effect of sequence changes on natural proteins and to optimize the design of novel proteins.

Amino Acid Sequence↗

Identification of functional SNPs in the 5-prime flanking sequences of human genes.

BACKGROUND: Over 4 million single nucleotide polymorphisms (SNPs) are currently reported to exist within the human genome. Only a small fraction of these SNPs alter gene function or expression, and therefore might be associated with a cell phenotype. These functional SNPs are consequently important in understanding human health. Information related to functional SNPs in candidate disease genes is critical for cost effective genetic association studies, which attempt to understand the genetics of complex diseases like diabetes, Alzheimer's, etc. Robust methods for the identification of functional SNPs are therefore crucial. We report one such experimental approach. RESULTS: Sequence conserved between mouse and human genomes, within 5 kilobases of the 5-prime end of 176 GPCR genes, were screened for SNPs. Sequences flanking these SNPs were scored for transcription factor binding sites. Allelic pairs resulting in a significant score difference were predicted to influence the binding of transcription factors (TFs). Ten such SNPs were selected for mobility shift assays (EMSA), resulting in 7 of them exhibiting a reproducible shift. The full-length promoter regions with 4 of the 7 SNPs were cloned in a Luciferase based plasmid reporter system. Two out of the 4 SNPs exhibited differential promoter activity in several human cell lines. CONCLUSIONS: We propose a method for effective selection of functional, regulatory SNPs that are located in evolutionary conserved 5-prime flanking regions (5'-FR) regions of human genes and influence the activity of the transcriptional regulatory region. Some SNPs behave differently in different cell types.

Algorithms↗

U6 snRNA variants isolated from the posterior silk gland of the silk moth Bombyx mori.

Five U6 small nuclear RNA (snRNA) isoforms were detected and characterized from the posterior silk gland (PSG) of the silk moth Bombyx mori (Nistari strain). Using the currently accepted U6 secondary structure model as a basis for comparison, the variants were analyzed for nucleotide differences across the sequence with a focus on known functional domains. Differences were observed primarily in single-stranded areas of which sixty percent were found in the highly conserved U4-U6 binding sites. In the Nistari strain, the U6A variant was found to be approximately four times more abundant as part of high molecular weight spliceosomal complexes when compared with U6A in the total unfractionated PSG cell lysate. Additionally, the European 703 B. mori strain total cell lysate U6 snRNA was analyzed and only the dominant U6A isoform initially identified in Nistari was found. Due to U6's essential role in pre-mRNA processing, variants may modulate assemblage of the catalytic core and in doing so potentially affect the rate of splicing. Phylogenetic analysis of the U6 snRNA sequences indicate an ancient divergence of U6 from the self-splicing group II intron module and a high degree of evolutionary conservation across species possibly due to functional constraints on the gene. Using in silico analysis, 35 full-length U6 variants were observed in the recently released Whole Genome Shotgun (WGS) database of the p50T strain. The consensus sequence of these U6 genes from p50T is identical to U6A identified in the Nistari strain. Furthermore p50T variant 1, which is represented in 14 genes, is equivalent to Nistari U6A.

Animals↗

Immune-associated nucleotide-1 (IAN-1) is a thymic selection marker and defines a novel gene family conserved in plants.

Positive selection of thymocytes is a complex and crucial event in T cell development that is characterized by cell death rescue, commitment toward the helper or cytotoxic lineage, and functional maturation of thymocytes bearing an appropriate TCR. To search for novel genes involved in this process, we compared gene expression patterns in positively selected thymocytes and their immediate progenitors in mice using the differential display technique. This approach lead to the identification of a novel gene, mIAN-1 (murine immune-associated nucleotide-1), that is switched on upon positive selection and predominantly expressed in the lymphoid system. We show that mIAN-1 encodes a 42-kDa protein sharing sequence homology with the pathogen-induced plant protein aig1 and that it defines a novel family of at least three putative GTP-binding proteins. Analysis of protein expression at various stages of thymocyte development links mIAN-1 to CD3-mediated selection events, suggesting that it represents a key player of thymocyte development and that it participates to peripheral specific immune responses. The evolutionary conservation of the IAN family provides a unique example of a plant pathogen response gene conserved in animals.

Amino Acid Sequence↗

Characterization of human homologs of the Drosophila seven in absentia (sina) gene.

Studies of Drosophila photoreceptor development have illustrated the means by which signal transduction events regulate cell fate decisions in a multicellular organization. Development of the R7 photoreceptor is best understood, and its formation is dependent on the seven in absentia (sina) gene. We have characterized two highly conserved human homologs of sina, termed SIAH1 and SIAH2. SIAH1 maps to chromosome 16q12 and encodes a 282-amino-acid protein with 76% amino acid identity to the Drosophila SINA protein. SIAH2 maps to chromosome 3q25 and encodes a 324-amino-acid protein that shares 68% identity with Drosophila SINA and 77% identity with human SIAH1. SIAH1 and SIAH2 were expressed in many normal and neoplastic tissues, and only subtle differences in their expression were noted. However, one of three murine homologs, Siah1B, was strongly induced in fibroblasts undergoing apoptotic cell death. While a previous study suggested that SINA was a nuclear protein, epitope-tagged SINA and SIAH1 proteins were found in the cytoplasm of Drosophila and mammalian cells. Their substantial evolutionary conservation, role in specifying cell fate, and activation in apoptotic cells suggest the SIAH proteins have important roles in vertebrate development. Furthermore, given the role of sina in Drosophila photoreceptor development, SIAH2 is a candidate for the Usher syndrome type 3 gene at chromosome 3q21-q25.

Adult↗