Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “evolutionary conservation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Usherin expression is highly conserved in mouse and human tissues.

Usher syndrome is an autosomal recessive disease that results in varying degrees of hearing loss and retinitis pigmentosa. Three types of Usher syndrome (I, II, and III) have been identified clinically with Usher type II being the most common of the three types. Usher type II has been localized to three different chromosomes 1q41, 3p, and 5q, corresponding to Usher type 2A, 2B, and 2C respectively. Usherin is a basement membrane protein encoded by the USH2A gene. Expression of usherin has been localized in the basement membrane of several tissues, however it is not ubiquitous. Immunohistochemistry detected usherin in the following human tissues: retina, cochlea, small and large intestine, pancreas, bladder, prostate, esophagus, trachea, thymus, salivary glands, placenta, ovary, fallopian tube, uterus, and testis. Usherin was absent in many other tissues such as heart, lung, liver, kidney, and brain. This distribution is consistent with the usherin distribution seen in the mouse. Conservation of usherin is also seen at the nucleotide and amino acid level when comparing the mouse and human gene sequences. Evolutionary conservation of usherin expression at the molecular level and in tissues unaffected by Usher 2a supports the important structural and functional role this protein plays in the human. In addition, we believe that these results could lead to a diagnostic procedure for the detection of Usher syndrome and those who carry an USH2A mutation.

Animals↗

Structural analysis of the regulatory elements of the type-II procollagen gene. Conservation of promoter and first intron sequences between human and mouse.

Transcription of the type-II procollagen gene (COL2A1) is very specifically restricted to a limited number of tissues, particularly cartilages. In order to identify transcription-control motifs we have sequenced the promoter region and the first intron of the human and mouse COL2A1 genes. With the assumption that these motifs should be well conserved during evolution, we have searched for potential elements important for the tissue-specific transcription of the COL2A1 gene by aligning the two sequences with each other and with the available rat type-II procollagen sequence for the promoter. With this approach we could identify specific evolutionarily well-conserved motifs in the promoter area. On the other hand, several suggested regulatory elements in the promoter region did not show evolutionary conservation. In the middle of the first intron we found a cluster of well-conserved transcription-control elements and we conclude that these conserved motifs most probably possess a significant function in the control of the tissue-specific transcription of the COL2A1 gene. We also describe locations of additional, highly conserved nucleotide stretches, which are good candidate regions in the search for binding sites of yet-uncharacterized cartilage-specific transcription regulators of the COL2A1 gene.

Amino Acid Sequence↗

An evolutionarily conserved enzyme degrades transforming growth factor-alpha as well as insulin.

A single enzyme found in both Drosophila and mammalian cells is able to selectively bind and degrade transforming growth factor (TGF)-alpha and insulin, but not EGF, at physiological concentrations. These growth factors are also able to inhibit binding and degradation of one another by the enzyme. Although there are significant immunological differences between the mammalian and Drosophila enzymes, the substrate specificity has been highly conserved. These results demonstrate the existence of a selective TGF-alpha-degrading enzyme in both Drosophila and mammalian cells. The evolutionary conservation of the ability to degrade both insulin and TGF-alpha suggests that this property is important for the physiological role of the enzyme and its potential for regulating growth factor levels.

Animals↗

Mutation of Gly-444 inactivates the S. pombe malic enzyme.

A mutant malic enzyme gene, mae2-, was cloned from a strain of Schizosaccharomyces pombe that displayed almost no malic enzyme activity. Sequence analysis revealed only one codon-altering mutation, a guanine to adenine at nucleotide 1331, changing the glycine residue at position 444 to an aspartate residue. Gly-444 is located in Region H, previously identified as one of eight highly conserved regions in malic enzymes. We found that Gly-444 is absolutely conserved in 27 malic enzymes from various prokaryotic and eukaryotic sources, as well as in three bacterial malolactic enzymes investigated. The evolutionary conservation of Gly-444 suggests that this residue is important for enzymatic function.

Amino Acid Sequence↗

The human chromosome 3 gene cluster ACY1-CACNA1D-ZNF64-ATP2B2 is evolutionarily conserved in Ateles paniscus chamek (Platyrrhini, Primates).

Comparative mapping of Ateles paniscus chamek and man indicated that four human 3p markers are syntenic in this karyotypically rearranged neotropical primate. The evolutionary conservation of this gene cluster includes three adjacent human shortest regions of overlap (SROs): 3p21.1 (ACY1), 3p21.3-->p21.2 (CACNA1D), and 3p21.3 (ZNF64). A fourth syntenic marker (ATP2B2), at a more distal human SRO (3p26-->p25), indicated that human 3pter-->p14 is evolutionarily conserved in Ateles chromosome 3 (APC 3). Conversely, allocations of two human 3q markers (AGTR1 and IL12A) clearly excluded APC 3. Finally, allocation of the major histocompatibility complex class I genes further confirmed human 6p-6q dissociations in Ateles.

Animals↗

Mutation, selection, and evolution of the Crohn disease susceptibility gene CARD15.

Three common mutations in the CARD15 (NOD2) gene are known to be associated with susceptibility to Crohn disease (CD), and genetic data suggest a gene dosage model with an increased risk of 2-4-fold in heterozygotes and 20-40-fold in homozygotes. However, the discovery of numerous rare variants of CARD15 indicates that some heterozygotes for the common mutations have a rare mutation on the other CARD15 allele, which would support a recessive model for CD. We addressed this issue by screening CARD15 for mutations in 100 CD patients who were heterozygous for one of the three common mutations. We also developed a strategy for evaluating potential disease susceptibility alleles (DSAs) that involves assessing the degree of evolutionary conservation of involved residues, predicted effects on protein structure and function, and genotyping in a large sample of cases and controls. The evolutionary analysis was aided by sequencing the entire coding region of CARD15 in three primates (chimp, gibbon, and tamarin) and aligning the human sequence with these and orthologs from other species. We found that 11 of the 100 CD patients screened had a second potential pathogenic mutation within the exonic and periexonic sequences examined. Assuming that there are no additional pathogenic mutations in noncoding regions, our study suggests that most carriers of the common DSAs are true heterozygotes, and supports previous evidence for a gene dosage model. Four novel nonsynonymous mutations were detected, one of which would produce premature termination of translation c.2686C>T (p.Arg896X). Two potential DSAs--c.2107C>T (p.Arg703Cys) and g.2238T>A (c.74-7T>A)--were significantly associated with CD in the case control sample. Analysis of the evolution of CARD15 revealed strong conservation of the encoded protein, with identity to the human sequence ranging from 99.1% in the chimp to 44.5% in fugu. Higher primates possess an open reading frame (ORF) upstream of the putative initiation site in other species that encodes a further 27 N-terminal amino acids, while four regions of high conservation are observed outside of the known domains of CARD15, indicative of additional residues of functional importance. The strategy developed here may have general application to the assessment of mutation pathogenicity and genetic models in other complex disorders.

Alleles↗

Genome-wide characterization of the bZIP gene family in Rattus norvegicus and expression profiling analysis during brain development.

BACKGROUND: The brown rat (Rattus norvegicus) serves as a cornerstone model organism in biomedical research, particularly for understanding physiological homeostasis and stress responses. The basic leucine zipper (bZIP) transcription factor family is a pivotal regulatory network involved in growth, organogenesis, and neurodevelopment. Despite its importance, a systematic characterization of the bZIP gene family in rats has remained elusive. RESULTS: In this study, we performed a genome-wide identification of 61 RnbZIP genes, which were categorized into 10 distinct subfamilies based on phylogenetic relationships and chromosomal localization. Structural analysis revealed conserved motif arrangements within subfamilies, while collinearity analysis identified significant gene duplication events-predominantly tandem and segmental duplications-that have driven the evolutionary expansion of the RnbZIP family. Quantitative analysis showed that members within the same subfamily shared 45%-92% sequence similarity (calculated using the BLOSUM62 scoring matrix), and all duplicated gene pairs underwent strong purifying selection (Ka/Ks&#x2009;<&#x2009;1). Comparative genomics across seven rodent species further underscored the evolutionary conservation and divergence of these factors. Expression profiling across diverse organs and brain developmental stages indicated that RnbZIP genes exhibit high tissue specificity. Notably, 10 candidate genes, including RnbZIP01, RnbZIP02, and RnbZIP08, demonstrated dynamic expression patterns during brain maturation, suggesting their essential roles in neurodevelopmental processes. CONCLUSIONS: Our findings provide a comprehensive structural and evolutionary framework for the RnbZIP gene family, highlighting their potential regulatory functions in rat organogenesis and brain development. This study establishes a valuable resource for further functional characterization of specific bZIP members in mammalian neurological systems.

Animals↗

Gene expression evolves faster in narrowly than in broadly expressed mammalian genes.

Despite much recent interest, it remains unclear what determines the rate of evolution of gene expression. To study this issue we develop a new measure, called "Expression Conservation Index" (ECI), to quantify the degree of tissue-expression conservation between two homologous genes. Applying this measure to a large set of gene expression data from human and mouse, we show that tissue expression tends to evolve rapidly for genes that are expressed in only a limited number of tissues, whereas tissue expression can be conserved for a long time for genes expressed in a large number of tissues. Therefore, expression breadth is an important determinant for evolutionary conservation of tissue expression. In addition, we find a rapid decrease in ECI with the synonymous divergence between duplicate genes, suggesting fast divergence in tissue expression between duplicate genes.

Animals↗

Granulocyte colony-stimulating factor.

Granulocyte colony-stimulating factor was discovered during attempts to define the normal regulators present in cell supernatants that could induce terminal differentiation of the murine myeloid leukemic cell line WEHI-3B D+. The purification and subsequent cloning of both murine and human G-CSF allowed the normal functions of this molecule to be elucidated and indicated that it was a relatively lineage-specific stimulator of the survival, proliferation, and differentiation of precursor cells of the neutrophilic granulocyte cell lineage as well as an activator of mature neutrophil function. Murine and human G-CSFs as well as the ligand-binding domains of G-CSF receptors have been strongly conserved so that the biological activities and receptor-binding characteristics of G-CSFs are completely species cross-reactive. The evolutionary conservation of G-CSFs is suggestive of an important role for the molecule in maintaining neutrophil levels and activity in vivo, a role amply supported by a series of animal experiments and clinical trials that have been performed recently. These experiments and trials have suggested that G-CSF will be clinically useful in augmenting patients' resistance to certain infections, in enhancing the neutrophil responses of some patients with reduced granulocyte counts or defective granulocytes, and in reducing the decline in granulocytes that occurs during chemotherapy or irradiation therapy and bone marrow transplantation. The special capacity of G-CSF among the CSFs to induce terminal differentiation in some myeloid leukemias is not understood in molecular terms. However, the capacity of G-CSF to stimulate proliferation in some myeloid leukemias as well argues that caution needs to be exercised in the timing of G-CSF administration and patient selection before considering this form of therapeutic approach.

Amino Acid Sequence↗

Predicting gene function by conserved co-expression.

We show that gene co-expression, which generally provides only a very weak signal for the prediction of functional interactions, can provide a reliable signal by exploiting evolutionary conservation. The encoded proteins of conserved co-expressed gene pairs are highly likely to be part of the same pathway not only after speciation (98%), but also after parallel gene duplication (97%). Conserved co-expression combined with homology data enables us to predict specific gene functions. The use of conservation between parallel duplicated gene pairs to predict function is especially promising given that gene duplication is common in eukaryotes, and that data from only a single organism can be used.

Animals↗

Detecting conserved interaction patterns in biological networks.

Molecular interaction data plays an important role in understanding biological processes at a modular level by providing a framework for understanding cellular organization, functional hierarchy, and evolutionary conservation. As the quality and quantity of network and interaction data increases rapidly, the problem of effectively analyzing this data becomes significant. Graph theoretic formalisms, commonly used for these analysis tasks, often lead to computationally hard problems due to their relation to subgraph isomorphism. This paper presents an innovative new algorithm, MULE, for detecting frequently occurring patterns and modules in biological networks. Using an innovative graph simplification technique based on ortholog contraction, which is ideally suited to biological networks, our algorithm renders these problems computationally tractable and scalable to large numbers of networks. We show, experimentally, that our algorithm can extract frequently occurring patterns in metabolic pathways and protein interaction networks from the KEGG, DIP, and BIND databases within seconds. When compared to existing approaches, our graph simplification technique can be viewed either as a pruning heuristic, or a closely related, but computationally simpler task. When used as a pruning heuristic, we show that our technique reduces effective graph sizes significantly, accelerating existing techniques by several orders of magnitude! Indeed, for most of the test cases, existing techniques could not even be applied without our pruning step. When used as a stand-alone analysis technique, MULE is shown to convey significant biological insights at near-interactive rates. The software, sample input graphs, and detailed results for comprehensive analysis of nine eukaryotic PPI networks are available at www.cs.purdue.edu/homes/koyuturk/mule.

Algorithms↗

The intracellular distribution and function of the high mobility group chromosomal proteins.

This brief review provides a framework for discussing current approaches being used to determine the cellular localization and function of the high mobility group chromosomal (HMG) proteins. The four main constituents of this group (HMG 1, 2, 14, 17) are present in all four eukaryotic kingdoms, have a relatively well conserved primary sequence and contain several functional domains which enable them to interact with DNA, histones and other components of the genome. The evolutionary conservation in the primary and tertiary structure as well as the observed correlations between cell phenotype and quantitative changes in protein levels and in post-synthesis modifications suggests that these proteins are components obligatory for proper cellular function. Proteins HMG 1, 2 are DNA-binding proteins which can distinguish between various types of single-stranded regions of the genome. Proteins HMG 14, 17 may be involved in maintaining specific chromatin regions in particular conformations. The data available presently suggests that these proteins are important structural elements of chromatin and chromosomes.

Amino Acid Sequence↗

Conservation of sequence and function of the pag-3 genes from C. elegans and C. briggsae.

The Caenorhabditis briggsae homologue of the Caenorhabditis elegans pag-3 gene was cloned and sequenced. When transformed into a C. elegans pag-3 mutant, the C. briggsae pag-3 gene rescued the pag-3 reverse kinker and lethargic phenotypes. The C. elegans pag-3 gene fused to lacZ was expressed in the same pattern in C. elegans and C. briggsae. Unlike many gene homologues compared between C. elegans and C. briggsae, extensive sequence conservation was found in the non-coding regions upstream of the pag-3 exons, in several of the introns and in the downstream non-coding region. Furthermore, the splice acceptor and splice donor sites were conserved, and the size of the introns and exons was surprisingly similar. The predicted protein sequence of C. briggsae PAG-3 was 85% identical to the protein sequence of C. elegans PAG-3. Because so much of the non-coding region of pag-3 was conserved, the control of pag-3 may be quite complex, involving the binding of many trans-acting factors. These results suggest the evolutionary conservation of the pag-3 gene sequence, its expression and function.

Amino Acid Sequence↗

HAT3.1, a novel Arabidopsis homeodomain protein containing a conserved cysteine-rich region.

Homeodomain proteins have been shown to play a major role in the development of various organisms. A novel Arabidopsis homeodomain protein has been isolated based on its capability to interact with a DNA motif derived from the light-induced cab-E promoter of Nicotiana plumbaginifolia. The homeodomain of this protein, designated HAT3.1, differs substantially from those in other plant homeobox proteins identified so far. Furthermore, HAT3.1 is unique among other Arabidopsis proteins in that it does not contain a leucine zipper motif following the homeodomain. HAT3.1 is further characterized by an N-terminal region that shares substantial sequence similarity with the maize homeodomain protein Zmhox1a. Within this conserved region, the presence of eight regularly spaced cysteine/histidine residues was observed reminiscent of other metal-binding domains. Based on the strong evolutionary conservation of this domain, it is proposed that this region represents a novel protein-motif which is denoted PHD-finger (plant homeodomain-finger). In vitro DNA binding studies demonstrated that HAT3.1 is capable of interacting with any DNA fragment larger than 100 bp. Interestingly, a deletion of the N-terminal PHD-finger domain completely abolished DNA binding, suggesting that this region may play an important functional role in protein-protein or protein-DNA interaction. HAT3.1 mRNA was primarily detected in root tissue, implying a regulatory function of this protein in root development.

Amino Acid Sequence↗

Patterns of conservation and divergence at the even-skipped locus of Drosophila.

The even-skipped (eve) gene of Drosophila melanogaster has been intensively studied as a model for spatial and temporal control of gene expression, using in vitro and transgenic techniques. Here, the study of eve is extended, using evolutionary conservation of DNA sequences. Conservation of much of the protein, and of known regulatory elements, supports models for eve function and regulation that have previously been advanced, and extensive conservation found in noncoding sequences predicts that functional elements exist that have yet to be defined. In contrast, a part of the protein implicated in transcriptional repression has diverged extensively while preserving overall amino acid composition, highlighting potentially essential features of this domain. Also, the basal promoter has diverged extensively, indicating evolutionary flexibility of promoter function.

Amino Acid Sequence↗

C. elegans DAF-12, Nuclear Hormone Receptors and human longevity and disease at old age.

In Caenorhabditis elegans, DAF-12 appears to be a decisive checkpoint for many life history traits including longevity. The daf-12 gene encodes a Nuclear Hormone Receptor (NHR) and is member of a superfamily that is abundantly represented throughout the animal kingdom, including humans. It is, however, unclear which of the human receptor representatives are most similar to DAF-12, and what their role is in determining human longevity and disease at old age. Using a sequence similarity search, we identified human NHRs similar to C. elegans DAF-12 and found that, based on sequence similarity, Liver X Receptor A and B are most similar to C. elegans DAF-12, followed by the Pregnane X Receptor, Vitamin D Receptor, Constitutive Andosteron Receptor and the Farnesoid X Receptor. Their biological functions include, amongst others, detoxification and immunomodulation. Both are processes that are involved in protecting the body from harmful environmental influences. Furthermore, the DAF-12 signalling systems seem to be functionally conserved and all six human NHRs have cholesterol derived compounds as their ligands. We conclude that the DAF-12 signalling system seems to be evolutionary conserved and that NHRs in man are critical for body homeostasis and survival. Genomic variations in these NHRs or their target genes are prime candidates for the regulation of human lifespan and disease at old age.

Aging↗

Molecular biology of the human pyruvate dehydrogenase complex: structural aspects of the E2 and E3 components.

The availability of the primary amino acid sequences of the E2 of PDC, alpha-KGDC and BCKADC from several prokaryotic and eukaryotic species has allowed us to compare the structural aspects of human PDC-E2 with those of the E2 components from the other complexes. The PDC-E2 components from all the species examined so far contain three structurally identifiable regions: the lipoyl-bearing domain, the E3-binding site, and the catalytic domain. The primary structure of the lipoyl-bearing domain shows considerable variation in its size, ranging from one to three repeating units of approximately 110 amino acids, but essentially preserving its function in the E2 components. In contrast, the sizes of the E3-binding site and the catalytic domain of PDC-E2 from several species are essentially similar and show considerable conservation of specific amino acid residues. Obviously, additional studies are warranted to better understand the structure-function relationships of these domains and the evolutionary conservation of PDC-E2 in different species. Similarly, the availability of the primary amino acid sequences of E3 from several prokaryotes and eukaryotes has also permitted comparison of the structural domains of these proteins with that of the known structure of human GR, a flavoprotein member of the pyridine nucleotide-disulfide oxidoreductase family. Four structural domains (FAD, NAD+, central, and interface domains) have been identified in the E3 components. On the basis of the comparison of the secondary structural elements of GR and E3, the core structure of these two proteins are shown to be similar. It is hoped that further analysis of E3 using site-directed mutagenesis and determination of its crystal structure will provide better insight into its structure-function relationships.

Amino Acid Sequence↗

Protein-protein interactions more conserved within species than across species.

Experimental high-throughput studies of protein-protein interactions are beginning to provide enough data for comprehensive computational studies. Today, about ten large data sets, each with thousands of interacting pairs, coarsely sample the interactions in fly, human, worm, and yeast. Another about 55,000 pairs of interacting proteins have been identified by more careful, detailed biochemical experiments. Most interactions are experimentally observed in prokaryotes and simple eukaryotes; very few interactions are observed in higher eukaryotes such as mammals. It is commonly assumed that pathways in mammals can be inferred through homology to model organisms, e.g. the experimental observation that two yeast proteins interact is transferred to infer that the two corresponding proteins in human also interact. Two pairs for which the interaction is conserved are often described as interologs. The goal of this investigation was a large-scale comprehensive analysis of such inferences, i.e. of the evolutionary conservation of interologs. Here, we introduced a novel score for measuring the overlap between protein-protein interaction data sets. This measure appeared to reflect the overall quality of the data and was the basis for our two surprising results from our large-scale analysis. Firstly, homology-based inferences of physical protein-protein interactions appeared far less successful than expected. In fact, such inferences were accurate only for extremely high levels of sequence similarity. Secondly, and most surprisingly, the identification of interacting partners through sequence similarity was significantly more reliable for protein pairs within the same organism than for pairs between species. Our analysis underlined that the discrepancies between different datasets are large, even when using the same type of experiment on the same organism. This reality considerably constrains the power of homology-based transfer of interactions. In particular, the experimental probing of interactions in distant model organisms has to be undertaken with some caution. More comprehensive images of protein-protein networks will require the combination of many high-throughput methods, including in silico inferences and predictions. http://www.rostlab.org/results/2006/ppi_homology/

Animals↗