Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Conservation analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

The 58,000-dalton cellular inhibitor of the interferon-induced double-stranded RNA-activated protein kinase (PKR) is a member of the tetratricopeptide repeat family of proteins.

PKR is a serine/threonine protein kinase induced by interferon treatment and activated by double-stranded RNAs. As a result of activation, PKR becomes autophosphorylated and catalyzes phosphorylation of the alpha subunit of protein synthesis eukaryotic initiation factor 2 (eIF-2). While studying the regulation of PKR in virus-infected cells, we found that a cellular 58-kDa protein (P58) was recruited by influenza virus to downregulate PKR and thus avoid the kinase's deleterious effects on viral protein synthesis and replication. We now report on the cloning, sequencing, expression, and structural analysis of the P58 PKR inhibitor, a 504-amino-acid hydrophilic protein. P58, expressed as a histidine fusion protein in Escherichia coli, blocked both the autophosphorylation of PKR and phosphorylation of the alpha subunit of eIF-2. Western blot (immunoblot) analysis showed that P58 is present not only in bovine cells but also in human, monkey, and mouse cells, suggesting the protein is highly conserved. Computer analysis revealed that P58 contains regions of homology to the DnaJ family of proteins and a much lesser degree of similarity to the PKR natural substrate, eIF-2 alpha. Finally, P58 contains nine tandemly arranged 34-amino-acid repeats, demonstrating that the PKR inhibitor is a member of the tetratricopeptide repeat family of proteins, the only member identified thus far with a known biochemical function.

Amino Acid Sequence↗

An analysis of the periodicity of conserved residues in sequence alignments of G-protein coupled receptors. Implications for the three-dimensional structure.

Twenty-three sequences from the family of G-protein coupled receptors have been aligned according to the 'historical alignment' procedure of Feng and Doolittle. Fourier transform analysis of this reveals that parts of five of the seven putative membrane-spanning regions exhibit a periodicity of conserved/nonconserved residues which is compatible with the periodicity of the alpha-helix. This would place the conserved residues on one side of the helix, which may face the inside of the proposed seven membered helical bundle.

Amino Acid Sequence↗

Recruitment to the cytoplasm of a cellular lamin-like protein from the nucleus during a poxvirus infection.

Monoclonal antibodies (Mabs) directed against core proteins of rabbit poxvirus (RPV) have proven effective in the identification of host cell proteins such as RNA polymerase II (Pol II) that may play a role in the infectious process (D. K. Morrison and R. W. Moyer, 1986, Cell 44, 587-596). In this article we describe a Mab that has allowed the detection and characterization of a lamin-like protein derived from the nucleus of the infected cell, which like Pol II is recruited to the cytoplasm following RPV infection. A portion of the gene encoding this protein has been isolated through the screening of a lambda gt11 expression vector library. Sequence analysis of the gene shows it to be derived from a member of the HindIII 1.9-kb repetitive element, a family of mammalian repetitive sequences that are highly conserved. Immunoblot analysis and sequence analysis of the open reading frame show divergent relatedness to certain nuclear lamins. The protein is not, however, one of the three principal lamins characterized to date, but instead appears to be a perinuclear protein related to the highly conserved nuclear lamins that is recruited to the cytoplasm during the infectious process.

Amino Acid Sequence↗

An approach based on spatial multicriteria analysis to map the nature conservation value of agricultural land.

Knowledge of the nature conservation value of agricultural land provides a useful input to land-use planning. However, the scarcity of suitable data causes this component to rarely play a role. The paper proposes a methodology based on commonly available data to assess the nature conservation value of agricultural landscapes, and to generate cartographic results to be used as decision variables in planning. The approach relies on landscape ecological indicators and on the application of multicriteria analysis in a Geographical Information System (GIS) context. Four criteria were selected: the agricultural landscape type, the cover of vegetation remnants and marginal features, the length of forest-agriculture ecotones, and the proximity to nature reserves. These criteria were assessed directly or by means of specific indicators, generating maps that were subsequently aggregated through spatial multicriteria analysis. The approach was tested in an alpine area located in Trentino (northern Italy).

Agriculture↗

Mutational analysis of CvaA in the highly conserved domain of the membrane fusion protein family.

The antibacterial peptide toxin colicin V (ColV) uses a dedicated signal sequence-independent export system for its secretion in Escherichia coli that involves the products of three genes, cvaA, cvaB, and tolC in this process. As a member of the membrane fusion protein (MFP) family, the CvaA protein has been proposed to interact with an outer membrane protein TolC via its C-terminal hydrophobic domain. The importance of this domain, which is highly conserved throughout the members of MFP family, was analyzed by use of site-directed mutagenesis of missense or nonsense mutations with suppressors. All the nonsense mutations tested resulted in the loss of ColV secretion, indicating the importance of the C-terminus of CvaA, including the last 100 residue-hydrophilic domain. The missense mutations of several conserved amino acids have no drastic effects. On the other hand, when Glu-248, Ala-262, Thr-274, Leu-285, Gly-313, Ala-322, or Val-335 of CvaA protein was mutated, the secretion of ColV was greatly reduced in certain mutants. While some mutations resulted in structural instability, Glu-248 to Lys and Ala-322 to Gly proteins were relatively stable, but were not functional in ColV secretion. The results indicate that these conserved amino acids are important for the structure and functions of CvaA in the secretion of ColV.

ATP-Binding Cassette Transporters↗

Mutational analysis of the functional role of conserved arginine and lysine residues in transmembrane domains of the murine reduced folate carrier.

The reduced folate carrier (RFC1) plays a major role in the delivery of folates into mammalian cells. RFC1 is an anion exchanger with seven conserved positively charged amino acid residues within 12 predicted transmembrane domains. This article explores the role of these residues in transport function by the development of cell lines in which arginines and lysines in RFC1 were replaced with leucine by site-directed mutagenesis. Three cell lines transfected with R131L, R155L, or R366L all lacked activity, despite high levels of protein expression in the plasma membrane, suggesting the crucial role of these amino acid residues in RFC1 function. In several mutant carriers, R26L, R42L, and K332L, there was little or no change in the influx K(t) value for MTX or influx K(i) value for folic acid. However, the R26L, R42L, and K332L carriers had decreased affinity for reduced folates. This was most prominent for K404L, which had 11- and 4-fold increases in influx K(i) for 5-methyl-THF and 5-formyl-THF, respectively, compared with L1210 cells. The marked influx stimulation observed with wild-type carrier when extracellular chloride was decreased was significantly diminished when influx was mediated by the K404L carrier, but was only slightly decreased with the R26L, R42L, and K332L mutants. This suggested that the K404 residue may be a major site of inhibition by chloride in the wild-type carrier. These studies indicate the important role that some positively charged residues within transmembrane domains of RFC1 play in RFC1 function.

Animals↗

Application of string kernels in protein sequence classification.

INTRODUCTION: The production of biological information has become much greater than its consumption. The key issue now is how to organise and manage the huge amount of novel information to facilitate access to this useful and important biological information. One core problem in classifying biological information is the annotation of new protein sequences with structural and functional features. METHOD: This article introduces the application of string kernels in classifying protein sequences into homogeneous families. A string kernel approach used in conjunction with support vector machines has been shown to achieve good performance in text categorisation tasks. We evaluated and analysed the performance of this approach, and we present experimental results on three selected families from the SCOP (Structural Classification of Proteins) database. We then compared the overall performance of this method with the existing protein classification methods on benchmark SCOP datasets. RESULTS: According to the F1 performance measure and the rate of false positive (RFP) measure, the string kernel method performs well in classifying protein sequences. The method outperformed all the generative-based methods and is comparable with the SVM-Fisher method. DISCUSSION: Although the string kernel approach makes no use of prior biological knowledge, it still captures sufficient biological information to enable it to outperform some of the state-of-the-art methods.

Algorithms↗

Cloning, sequencing and expression of stem cell factor (c-kit ligand) cDNA of brushtail possum (Trichosurus vulpecula).

By means of reverse transcription polymerase chain reaction (RT-PCR), three stem cell factor (SCF) cDNAs (822-738 bp in size) were amplified from brushtail possum ovarian poly (A)+ RNA. The largest and smallest of these cDNAs were cloned and sequenced. Characterization of these cDNAs has revealed that possum SCF has approximately 75% and 66% homology to SCF of eutherian mammals at the nucleotide level and the predicted amino acid level respectively. Nucleotide sequencing shows that the 738-bp cDNA represents an mRNA splice variant, equivalent to that found in eutherian mammals, in which an exon (84 bp) encoding a potential proteolytic cleavage site is removed. Comparison of the predicted possum SCF amino acid sequence with the predicted SCF amino acid sequences from eutherian mammals reveals conservation of all cysteine residues and 3 of 4 potential N-linked glycosylation sites. In addition, the hydropathicity profile of the possum SCF protein is similar to that of eutherian SCF suggesting that protein conformation is conserved. Northern analysis was used to characterize possum SCF gene expression in adult ovary and testis. A major transcript of 9 kb was observed in both ovarian and testicular tissue. The conservation of the SCF gene and its predicted protein, suggests that SCF in the possum has similar biological activities to SCF in eutherian mammals.

Alternative Splicing↗

Sequencing and analysis of the Mmethylococcus capsulatus (Bath) solublemethane monooxygenase genes.

The soluble methane monooxygenase (sMMO) hydroxylase is a prototypical member of the class of proteins with non-heme carboxylate-bridged diiron sites. The sMMO subclass of enzyme systems has several distinguishing characteristics, including the ability to catalyze hydroxylation or epoxidation chemistry, a multisubunit hydroxylase containing diiron centers in its alpha subunits, and the requirement of a coupling protein for optimal activity. Sequence homology alignment of known members of the sMMO family was performed in an effort to identify protein regions giving rise to these unique features. DNA sequencing of the Methylococcus capsulatus (Bath) sMMO genes confirmed previously identified sequencing errors and corrected two additional errors, each of which was confirmed by at least one independent method. Alignments of homologous proteins from sMMO, phenol hydroxylase, toluene 2-, 3-, and 4-monooxygenases, and alkene monooxygenase systems revealed an interesting set of absolutely conserved amino-acid residues, including previously unidentified residues located outside the diiron active site of the hydroxylase. By mapping these residues on to the M. capsulatus (Bath) sMMO hydroxylase crystal structure, functional and structural roles were proposed for the conserved regions. Analysis of the active site showed a highly conserved hydrogen-bonding network on one side of the diiron cluster but little homology on the opposite side, where substrates are presumed to bind. It is suggested that conserved residues on the hydroxylase surface may be important for protein-protein interactions with the reductase and coupling ancillary proteins and/or serve as part of an electron-transfer pathway. A possible way by which binding of the coupling protein at the surface of the hydroxylase might transfer information to the diiron active site at the interior is proposed.

Amino Acid Sequence↗

Developmental relations between notational counting and number conservation.

The 2 studies of this report sought to determine the developmental relationship between the child's use of counting as a notational symbol system to extract, compare, and reproduce numerical information and the development of number conservation. In study 1, children between 4 and 6 years of age were administered notational-counting and number-conservation tasks. Analysis of children's profiles across tasks indicated that children develop quantitative counting strategies (but do not necessarily count accurately) before they develop number-conservation concepts. In study 2, the generality of this sequence was tested. A population of 7- to 9-year-old "learning-disabled" children who reportedly were developing atypical counting skills were administered notational-counting and number-conservation tasks. All of these children who conserved number also used quantitative counting strategies, although some of these children frequently counted arrays inaccurately. The significance of these findings is discussed with respect to existing models of counting/number-conservation relations, and an alternative formulation is suggested based upon the new findings.

Child↗

Mitochondrial DNA variation of the common hippopotamus: evidence for a recent population expansion.

Mitochondrial DNA control region sequence variation was obtained and the population history of the common hippopotamus was inferred from 109 individuals from 13 localities covering six populations in sub-Saharan Africa. In all, 100 haplotypes were defined, of which 98 were locality specific. A relatively low overall nucleotide diversity was observed (pi = 1.9%), as compared to other large mammals so far studied from the same region. Within populations, nucleotide diversity varied from 1.52% in Zambia to 1.92% in Queen Elizabeth and Masai Mara. Overall, low but significant genetic differentiation was observed in the total data set (F(ST) = 0.138; P = 0.001), and at the population level, patterns of differentiation support previously suggested hippopotamus subspecies designations (F(CT) = 0.103; P = 0.015). Evidence that the common hippopotamus recently expanded were revealed by: (i) lack of clear geographical structure among haplotypes, (ii) mismatch distributions of pairwise differences (r = 0.0053; P = 0.012) and site-frequency spectra, (iii) Fu's neutrality statistics (F(S) = -155.409; P < 0.00001) and (iv) Fu and Li's statistical tests (D* = -3.191; P < 0.01, F* = -2.668; P = 0.01). Mismatch distributions, site-frequency spectra and neutrality statistics performed at subspecies level also supported expansion of Hippopotamus amphibius across Africa. We interpret observed common hippopotamus population history in terms of Pleistocene drainage overflow and suggest recognising the three subspecies that were sampled in this study as separate management units in future conservation planning.

Africa South of the Sahara↗

Arabidopsis microarrays identify conserved and differentially expressed genes involved in shoot growth and development from distantly related plant species.

Expressed sequence tags (EST)-based microarrays are powerful tools for gene discovery and signal transduction studies in a small number of well-characterized species. To explore the usefulness of this technique for poorly characterized species, we have hybridized the 11,522-element Arabidopsis microarrays with labeled cDNAs from mature leaf and shoot apices from several different species. Expression of 23 to 47% of the genes on the array was detected, demonstrating that a large number of genes from distantly related species can be surveyed on Arabidopsis arrays. Differential expression of genes with known functions was indicative of the physiological state of the tissues tested. Genes involved in cell division, stress responses, and development were conserved and expressed preferentially in growing shoots.

Arabidopsis↗

Automatic clustering of orthologs and inparalogs shared by multiple proteomes.

MOTIVATION: The complete sequencing of many genomes has made it possible to identify orthologous genes descending from a common ancestor. However, reconstruction of evolutionary history over long time periods faces many challenges due to gene duplications and losses. Identification of orthologous groups shared by multiple proteomes therefore becomes a clustering problem in which an optimal compromise between conflicting evidences needs to be found. RESULTS: Here we present a new proteome-scale analysis program called MultiParanoid that can automatically find orthology relationships between proteins in multiple proteomes. The software is an extension of the InParanoid program that identifies orthologs and inparalogs in pairwise proteome comparisons. MultiParanoid applies a clustering algorithm to merge multiple pairwise ortholog groups from InParanoid into multi-species ortholog groups. To avoid outparalogs in the same cluster, MultiParanoid only combines species that share the same last ancestor. To validate the clustering technique, we compared the results to a reference set obtained by manual phylogenetic analysis. We further compared the results to ortholog groups in KOGs and OrthoMCL, which revealed that MultiParanoid produces substantially fewer outparalogs than these resources. AVAILABILITY: MultiParanoid is a freely available standalone program that enables efficient orthology analysis much needed in the post-genomic era. A web-based service providing access to the original datasets, the resulting groups of orthologs, and the source code of the program can be found at http://multiparanoid.cgb.ki.se.

Algorithms↗

Multi-species sequence comparison: the next frontier in genome annotation.

Multi-species comparisons of DNA sequences are more powerful for discovering functional sequences than pairwise DNA sequence comparisons. Most current computational tools have been designed for pairwise comparisons, and efficient extension of these tools to multiple species will require knowledge of the ideal evolutionary distance to choose and the development of new algorithms for alignment, analysis of conservation, and visualization of results.

Animals↗

Efficient estimation of emission probabilities in profile hidden Markov models.

MOTIVATION: Profile hidden Markov models provide a sensitive method for performing sequence database search and aligning multiple sequences. One of the drawbacks of the hidden Markov model is that the conserved amino acids are not emphasized, but signal and noise are treated equally. For this reason, the number of estimated emission parameters is often enormous. Focusing the analysis on conserved residues only should increase the accuracy of sequence database search. RESULTS: We address this issue with a new method for efficient emission probability (EEP) estimation, in which amino acids are divided into effective and ineffective residues at each conserved alignment position. A practical study with 20 protein families demonstrated that the EEP method is capable of detecting family members from other proteins with sensitivity of 98% and specificity of 99% on the average, even if the number of free emission parameters was decreased to 15% of the original. In the database search for TIM barrel sequences, EEP recognizes the family members nearly as accurately as HMMER or Blast, but the number of false positive sequences was significantly less than that obtained with the other methods. AVAILABILITY: The algorithms written in C language are available on request from the authors.

Algorithms↗

The nucleobase-ascorbate transporter (NAT) signature motif in UapA defines the function of the purine translocation pathway.

UapA, a member of the NAT/NCS2 family, is a high affinity, high capacity, uric acid-xanthine/H+ symporter of Aspergillus nidulans. We have previously presented evidence showing that a highly conserved signature motif ([Q/E/P]408-N-X-G-X-X-X-X-T-[R/K/G])417 is involved in UapA function. Here, we present a systematic mutational analysis of conserved residues in or close to the signature motif of UapA. We show that even the most conservative substitutions of residues Q408, N409 and G411 modify the kinetics and specificity of UapA, without affecting targeting in the plasma membrane. Q408 substitutions show that this residue determines both substrate binding and transport catalysis, possibly via interactions with position N9 of the imidazole ring of purines. Residue N409 is an irreplaceable residue necessary for transport catalysis, but is not involved in substrate binding. Residue G411 determines, indirectly, both the kinetics (K(m), V) and specificity of UapA, probably due to its particular property to confer local flexibility in the binding site of UapA. In silico predictions and a search in structural databases strongly suggest that the first part of the NAT signature motif of UapA (Q(408)NNG(411)) should form a loop, the structure of which is mostly affected by mutations in G411. Finally, substitutions of residues T416 and R417, despite being much better tolerated, can also affect the kinetics or the specificity of UapA. Our results show that the NAT signature motif defines the function of the UapA purine translocation pathway and strongly suggest that this might occur by determining the interactions of UapA with the imidazole part of purines.

Amino Acid Motifs↗

MECP2 gene mutation analysis in the British and Italian Rett Syndrome patients: hot spot map of the most recurrent mutations and bioinformatic analysis of a new MECP2 conserved region.

Rett syndrome (RTT) is an X-linked dominant neurological disorder, which appears to be the most common genetic cause of profound combined intellectual and physical disability in Caucasian females. This syndrome has been associated with mutations of the MECP2 gene, a transcriptional repressor of unknown target genes. We report a detailed mutational analysis of a large cohort of RTT patients from the UK and Italy. This study has permitted us to produce a hot spot map of the mutations identified. Bioinformatic analysis of the mutations, taking advantage of structural and evolutionary data, leads us to postulate the existence of a new functional domain in the MeCP2 protein, conserved among brain-specific regulatory factors.

Adolescent↗

Different substitutions at conserved amino acids in domains II and III in the Sendai L RNA polymerase protein inactivate viral RNA synthesis.

The Sendai virus RNA polymerase is a complex of two virus-encoded proteins, the phosphoprotein (P) and the large (L) protein, where L is believed to possess all the enzymatic activities necessary for viral transcription and replication. The alignment of amino acid sequences of L proteins from negative-sense RNA viruses shows six regions, designated domains I-VI, of good conservation which have been proposed to be important for the various enzymatic activities of the polymerase. To directly address the role(s) of domains II and III, site-directed mutations were constructed by the substitution of multiple amino acids at 13 highly or mostly conserved residues. Analysis of in vitro viral transcription and replication showed that the majority of the mutations completely inactivated the L protein for all aspects of RNA synthesis, thus conservation correlated with the essential nature of the amino acid. At some positions different phenotypes, from inactivation to partial activities, were observed which depended on the nature of the amino acid that was substituted. Two mutants, K543R and K666V, could synthesize some leader RNA, but were defective in mRNA synthesis and replication. K666R and G737E had significantly reduced replication compared to transcription in vitro, but replicated genome RNA much more efficiently in vivo. K666A gave transcription, but no replication. Representative inactive L mutants, however, were still able to bind P protein and the polymerase complex was capable of binding nucleocapsids, so the defect appeared to be in the initiation of RNA synthesis.

Amino Acid Sequence↗