Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Conservation analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

HNH family subclassification leads to identification of commonality in the His-Me endonuclease superfamily.

The HNHc (SMART ID: SM00507) domain (SCOP nomenclature: HNH family) can be subclassified into at least eight subsets by iterative refinement of HMM profiles. An initial clustering of 323 proteins containing the HNHc domain helped identify the subsets. The subsets could be differentiated on the basis of the pattern of occurrence of seven defining features. Domain association is also different between the subsets. The subsets show organism as well as domain-based clustering, suggestive of propagation by both duplication and horizontal transfer events. Structure-based sequence analysis of the subsets led to the identification of common structural and sequence motifs in the HNH family with the other three families under the His-Me endonuclease superfamily.

Amino Acid Sequence↗

Protein-coding regions prediction combining similarity searches and conservative evolutionary properties of protein-coding sequences.

The gene identification procedure in a completely new gene with no good homology with protein sequences can be a very complex task. In order to identify the protein-coding region, a new method, 'SYNCOD', based on the analysis of conservative evolutionary properties of coding regions, has been realized. This program is able to identify and use the coding region homologies of the non-annotated (unknown) protein-coding sequences already present in the nucleotide sequence databases by using the alignment produced by BLASTN. The ratio of number mismatches resulting in synonymous codons to the number of mismatches resulting in non-synonymous codons is estimated for each open reading frame. Monte Carlo simulations are then used to estimate the significance of the ratio deviation from random behavior. The SYNCOD program has been tested on generated random sequences and on different control sets. The high accuracy of predicting protein-coding regions (the correlation coefficient, CC, varies from 0.67 to 0.79) and the high specificity (the portion of wrong exons, WE, varies from 0.06 to 0.07) have proved to be important features of the suggested approach. The SYNCOD program is resident on the ITBA-CNR Web Server and can be used via the Internet (URL: www.itba.mi.cnr.it/webgene).

Algorithms↗

Molecular analysis of muskelin identifies a conserved discoidin-like domain that contributes to protein self-association.

Muskelin is an intracellular protein with a C-terminal kelch-repeat domain that was initially characterized as having functional involvement in cell spreading on the extracellular matrix glycoprotein thrombospondin-1. As one approach to understanding the functional properties of muskelin, we have combined bioinformatic and biochemical studies. Through analysis of a new dataset of eight animal muskelins, we showed that the N-terminal region of the polypeptide corresponds to a predicted discoidin-like domain. This domain architecture is conserved in fungal muskelins and reveals a structural parallel between the muskelins and certain extracellular fungal galactose oxidases, although the phylogeny of the two groups appears distinct. In view of the fact that a number of kelch-repeat proteins have been shown to self-associate, co-immunoprecipitation, protein pull-down assays and studies of cellular localization were carried out with wild-type, deletion mutant and point mutant muskelins to investigate the roles of the discoidin-like and kelch-repeat domains. We obtained evidence for cis- and trans-interactions between the two domains. These studies provide evidence that muskelin self-associates through a head-to-tail mechanism involving the discoidin-like domain.

Amino Acid Sequence↗

Characterization of the ATPase and unwinding activities of the yeast DEAD-box protein Has1p and the analysis of the roles of the conserved motifs.

The yeast DEAD-box protein Has1p is required for the maturation of 18S rRNA, the biogenesis of 40S r-subunits and for the processing of 27S pre-rRNAs during 60S r-subunit biogenesis. We purified recombinant Has1p and characterized its biochemical activities. We show that Has1p is an RNA-dependent ATPase in vitro and that it is able to unwind RNA/DNA duplexes in an ATP-dependent manner. We also report a mutational analysis of the conserved residues in motif I (86AKTGSGKT93), motif III (228SAT230) and motif VI (375HRVGRTARG383). The in vivo lethal K92A substitution in motif I abolishes ATPase activity in vitro. The mutations S228A and T230A partially dissociate ATPase and helicase activities, and they have cold-sensitive and lethal growth phenotypes, respectively. The H375E substitution in motif VI significantly decreased helicase but not ATPase activity and was lethal in vivo. These results suggest that both ATPase and unwinding activities are required in vivo. Has1p possesses a Walker A-like motif downstream of motif VI (383GTKGKGKS390). K389A substitution in this motif significantly increases the Has1p activity in vitro, which indicates it potentially plays a role as a negative regulator. Finally, rRNAs and poly(A) RNA serve as the best stimulators of the ATPase activity of Has1p among the tested RNAs.

Adenosine Triphosphatases↗

The Fanconi anemia gene network is conserved from zebrafish to human.

Fanconi anemia (FA) is a complex disease involving nine identified and two unidentified loci that define a network essential for maintaining genomic stability. To test the hypothesis that the FA network is conserved in vertebrate genomes, we cloned and sequenced zebrafish (Danio rerio) cDNAs and/or genomic BAC clones orthologous to all nine cloned FA genes (FANCA, FANCB, FANCC, FANCD1, FANCD2, FANCE, FANCF, FANCG, and FANCL), and identified orthologs in the genome database for the pufferfish Tetraodon nigroviridis. Genomic organization of exons and introns was nearly identical between zebrafish and human for all genes examined. Hydrophobicity plots revealed conservation of FA protein structure. Evolutionarily conserved regions identified functionally important domains, since many amino acid residues mutated in human disease alleles or shown to be critical in targeted mutagenesis studies are identical in zebrafish and human. Comparative genomic analysis demonstrated conserved syntenies for all FA genes. We conclude that the FA gene network has remained intact since the last common ancestor of zebrafish and human lineages. The application of powerful genetic, cellular, and embryological methodologies make zebrafish a useful model for discovering FA gene functions, identifying new genes in the network, and identifying therapeutic compounds.

Amino Acid Sequence↗

Crystal structure of Escherichia coli PurE, an unusual mutase in the purine biosynthetic pathway.

BACKGROUND: Conversion of 5-aminoimidazole ribonucleotide (AIR) to 4-carboxyaminoimidazole ribonucleotide (CAIR) in Escherichia coli requires two proteins - PurK and PurE. PurE has recently been shown to be a mutase that catalyzes the unusual rearrangement of N(5)-carboxyaminoimidazole ribonucleotide (N(5)-CAIR), the PurK reaction product, to CAIR. PurEs from higher eukaryotes are homologous to E. coli PurE, but use AIR and CO(2) as substrates to produce CAIR directly. RESULTS: The 1.50 A crystal structure of PurE reveals an octameric structure with 422 symmetry. A central three-layer (alphabetaalpha) sandwich domain and a kinked C-terminal helix form the folded structure of the monomeric unit. The structure reveals a cleft at the interface of two subunits and near the C-terminal helix of a third subunit. Co-crystallization experiments with CAIR confirm this to be the mononucleotide-binding site. The nucleotide is bound predominantly to one subunit, with conserved residues from a second subunit making up one wall of the cleft. CONCLUSIONS: The crystal structure of PurE reveals a unique quaternary structure that confirms the octameric nature of the enzyme. An analysis of the native crystal structure, in conjunction with sequence alignments and studies of co-crystals of PurE with CAIR, reveals the location of the active site. The environment of the active site and the analysis of conserved residues between the two classes of PurEs suggests a model for the differences in their substrate specificities and the relationship between their mechanisms.

Amino Acid Sequence↗

Genetic consequence of restricted habitat and population decline in endangered Isoetes sinensis (Isoetaceae).

BACKGROUND AND AIMS: Isoetes sinensis (Isoeteaceae) is a critically endangered aquatic quillwort in eastern China. Rapid decline of extant population size and local population extinction have occurred in recent years and have raised great concerns among conservationists. METHODS: Amplified fragment length polymorphisms (AFLPs) were used to investigate the genetic variation and population structure of seven extant populations of the species. KEY RESULTS: Eight primer combinations produced a total of 343 unambiguous bands of which 210 (61.2 %) were polymorphic. Isoetes sinensis exhibited a high level of intra-population genetic diversity (H(E) = 0.118; hs = 0.147; I = 0.192; P = 35.2 %). The genetic variation within each of the populations was not positively correlated with their size, suggesting recent population decline, which is well in accordance with field data of demographic surveys. Moreover, a high degree of genetic differentiation (F(ST) = 0.535; G(ST) = 0.608; theta(B) = 0.607) was detected among populations and no correlation was found between geographical and genetic distance, suggesting that populations were in disequilibrium of migration-drift. Genetic drift played a more important role than gene flow in the current population genetic structure of I. sinensis because migration of I. sinensis is predominantly water-mediated and habitat range was highly influenced by environment changes. CONCLUSIONS: Genetic information obtained in the present study provides useful baseline data for formulating conservation strategies. Conservation management, including both reinforcement for in situ populations and ex situ conservation programmes should be carefully designed to avoid the potential risk of outbreeding depression by admixture of individuals from different regions. However, translocation within the same regional population should be considered as a measure of genetic enhancement to rehabilitate local populations. An ex situ conservation strategy for conserving all extant populations to maximize genomic representation of the species is also recommended.

China↗

Unexpected flexibility in an evolutionarily conserved protein-RNA interaction: genetic analysis of the Sm binding site.

Human autoantibodies of the Sm specificity recognize a conserved set of proteins found in the U class small nuclear ribonucleoproteins (U snRNPs), key trans-acting factors involved in the splicing of mRNA precursors. The Sm protein binding site in U snRNAs is unusual because of its single-stranded nature and its simple sequence motif (AU5-6GPu). Here we use genetics to probe this specific protein-RNA interaction by saturation mutagenesis of the Sm binding site of the Saccharomyces cerevisiae U5 snRNA. The assay system used to analyze these mutations takes advantage of a conditionally expressed U5 gene which does not support growth under non-permissive conditions; U5 genes containing Sm site mutations were tested for their ability to complement this lethal phenotype. Our results indicate that the Sm binding site is remarkably tolerant to mutation despite its high degree of conservation, suggesting that relatively few or redundant specific contacts can determine recognition of single-stranded RNA by protein. A complementary biochemical analysis of these mutants demonstrates that integrity of the Sm site is necessary for snRNP stability in vivo and in vitro.

Autoantigens↗

Analysis of the recourse to conservative surgery in the treatment of breast tumors.

AIMS AND BACKGROUND: Conservative surgery is the treatment of choice for malignant tumors, at least up to stage II. The aim of this study was to analyze the recourse to conservative surgery for breast tumors and its determinants (ie, characteristics of hospitals and patients). METHODS: The study was conducted in Italy's Lazio region and was based on administrative data of the regional Hospital Information System, a database on hospitalizations. We selected all regional hospitalizations for therapeutic breast surgery over 1997, classifying them as either "conservative" or "non-conservative". The other variables considered were type of hospital, number of beds, volume of activity (average annual number of hospitalizations for breast cancer surgery), specific diagnosis, severity of cancer, and patient's age, place of residence, and socioeconomic level. A logistic model was used for multivariate analysis. RESULTS: A total of 7235 hospitalizations were analyzed, 3570 (49%) for malignant tumors and 3665 (51%) for benign disease. The logistic model showed that the factors most closely correlated with conservative surgery were age (OR = 2.2; 95% Cl: 1.8-2.6, for the age group <50 years compared to >70 years); severity of cancer (OR = 0.6; 95% Cl: 0.5-0.8, for non-localized compared to localized tumors), and volume of activity of the hospital (OR = 1.3; 95% CI: 1.0-1.6, for hospitals with >70 operations/year compared to those with <20 operations/year). The study also revealed that surgery for malignant tumors was performed by both high-volume and low-volume hospitals throughout the region. CONCLUSIONS: The association between conservative surgery and younger age, even after controlling for the severity of cancer, points to the need to encourage adherence to the existing guidelines. The association between conservative surgery and high-volume hospitals and the finding that a high proportion of breast operations is performed in low-volume facilities suggest that further efforts should be made to promote admission to high-volume hospitals.

Adult↗

phi 29 DNA polymerase residue Leu384, highly conserved in motif B of eukaryotic type DNA replicases, is involved in nucleotide insertion fidelity.

Replicative DNA polymerases achieve insertion fidelity by geometric selection of a complementary nucleotide followed by induced fit: movement of the fingers subdomain toward the active site to enclose the incoming and templating nucleotides generating a binding pocket for the nascent base pair. Several residues of motif B of DNA polymerases from families A and B, localized in the fingers subdomain, have been described to be involved in template/primer binding and dNTP selection. Here we complete the analysis of this motif, which has the consensus "KLX2NSXYG" in DNA polymerases from family B, characterized by mutational analysis of conserved leucine, Leu384 of phi 29 DNA polymerase. Mutation of Leu384 into Arg resulted in a phi 29 DNA polymerase with reduced nucleotide insertion fidelity during DNA-primed polymerization and protein-primed initiation reactions. However, the mutation did not alter the intrinsic affinity for the different dNTPs, as shown in the template-independent terminal protein-deoxynucleotidylation reaction. We conclude that Leu384 of phi 29 DNA polymerase plays an important role in positioning the templating nucleotide at the polymerization active site and in controlling nucleotide insertion fidelity. This agrees with the localization of the corresponding residue in the closed ternary complexes of family A and family B DNA polymerases, contributing to form the binding pocket for the nascent base pair. As an additional effect, mutant polymerase L384R was strongly reduced in DNA binding, resulting in reduced processivity during polymerization.

Amino Acid Motifs↗

An efficient algorithm for optimizing whole genome alignment with noise.

MOTIVATION: This paper is concerned with algorithms for aligning two whole genomes so as to identify regions that possibly contain conserved genes. Motivated by existing heuristic-based software tools, we initiate the study of an optimization problem that attempts to uncover conserved genes with a global concern. Another interesting feature in our formulation is the tolerance of noise, which also complicates the optimization problem. A brute-force approach takes time exponential in the noise level. RESULTS: We show how an insight into the optimization structure can lead to a drastic improvement in the time and space requirement [precisely, to O(k2n2) and O(k2n), respectively, where n is the size of the input and k is the noise level]. The reduced space requirement allows us to implement the new algorithm, called MaxMinCluster, on a PC. It is exciting to see that when tested with different real data sets, MaxMinCluster consistently uncovers a high percentage of conserved genes that have been published by GenBank. Its performance is indeed favorably compared to MUMmer (perhaps the most popular software tool for uncovering conserved genes in a whole-genome scale). AVAILABILITY: The source code is available from the website http://www.csis.hku.hk/~colly/maxmincluster/ detailed proof of the propositions can also be found there.

Algorithms↗

The demographic history of the New Zealand short-tailed bat Mystacina tuberculata inferred from modified control region sequences.

Short-tailed bats Mystacina tuberculata were widespread throughout the forest that dominated prehuman New Zealand, but extensive deforestation has restricted them to scattered populations in forest fragments. In a previous study, the species' intraspecific phylogeny was investigated using multiple mitochondrial gene sequences. Six phylogroups were identified with estimated divergences of 0.93-0.68 Ma. In the current study, the phylogeographical structure and demographic history of the phylogroups were investigated using control region sequences modified by removing homoplasic sites. Phylogeographical structure in the North Island was generally consistent with an isolation-by-distance dispersal model. Coalescent-based analyses (i.e. mismatch distributions, skyline plots, lineage dispersal analysis and nested clade analysis) indicated that the three phylogroups found in central and southern North Island expanded before the last glacial maximum, presumably during interstadials when Nothofagus forest was most extensive. Genetic structure within a central North Island hybrid zone was consistent with range expansion from separate refugia following reforestation after catastrophic volcanic eruptions. Phylogeographical structure in the South Island was consistent with southern populations originating during rapid southward range expansion from refugia in northern South Island following postglacial reforestation of the South Island 10-9 kya.

Animals↗

Sequence diversity in S1 genes and S1 translation products of 11 serotype 3 reovirus strains.

The S1 gene nucleotide sequences of 10 type 3 (T3) reovirus strains were determined and compared with the T3 prototype Dearing strain in order to study sequence diversity in strains of a single reovirus serotype and to learn more about structure-function relationships of the two S1 translation products, sigma 1 and sigma 1s. Analysis of phylogenetic trees constructed from variation in the sigma 1-encoding S1 nucleotide sequences indicated that there is no pattern of S1 gene relatedness in these strains based on host species, geographic site, or date of isolation. This suggests that reovirus strains are transmitted rapidly between host species and that T3 strains with markedly different S1 sequences circulate simultaneously. Comparison of the deduced sigma 1 amino acid sequences of the 11 T3 strains was notable for the identification of conserved and variable regions of sequence that correlate with the proposed domain organization of sigma 1 (M.L. Nibert, T.S. Dermody, and B. N. Fields, J. Virol. 64:2976-2989, 1990). Repeat patterns of apolar residues thought to be important for sigma 1 structure were conserved in all strains examined. The deduced sigma 1s amino acid sequences of the strains were more heterogeneous than the sigma 1 sequences; however, a cluster of basic residues near the amino terminus of sigma 1s was conserved. This analysis has allowed us to investigate molecular epidemiology of T3 reovirus strains and to identify conserved and variable sequence motifs in the S1 translation products, sigma 1 or sigma 1s.

Amino Acid Sequence↗

HPV as a Molecular Hacker: Computational Exploration of HPV-Driven Changes in Host Regulatory Networks.

Human Papillomavirus (HPV), particularly high-risk strains such as HPV16 and HPV18, is a leading cause of cervical cancer and a significant risk factor for several other epithelial malignancies. While the oncogenic mechanisms of viral proteins E6 and E7 are well characterized, the broader effects of HPV infection on host transcriptional regulation remain less clearly defined. This study explores the hypothesis that conserved genomic motifs within the HPV genome may act as molecular decoys, sequestering human transcription factors (TFs) and thereby disrupting normal gene regulation in host cells. Such interactions could contribute to oncogenesis by altering the transcriptional landscape and promoting malignant transformation.We conducted a computational analysis of the genomes of high-risk HPV types using MEME-ChIP for de novo motif discovery, followed by Tomtom for identifying matching human TFs. Protein-protein interactions among the predicted TFs were examined using STRING, and biological pathway enrichment was performed with Enrichr. The analysis identified conserved viral motifs with the potential to interact with host transcription factors (TFs), notably those from the FOX, HOX, and NFAT families, as well as various zinc finger proteins. Among these, SMARCA1, DUX4, and CDX1 were not previously associated with HPV-driven cell transformation. Pathway enrichment analysis revealed involvement in several key biological processes, including modulation of Wnt signaling pathways, transcriptional misregulation associated with cancer, and chromatin remodeling. These findings highlight the multifaceted strategies by which HPV may influence host cellular functions and contribute to pathogenesis. In this context, the study underscores the power of in silico approaches for elucidating viral-host interactions and reveals promising therapeutic targets in computationally predicted regulatory network changes.

Humans↗

The power and perils of 'molecular taxonomy': a case study of eyeless and endangered Cicurina (Araneae: Dictynidae) from Texas caves.

Rapid development in karst-rich regions of the US state of Texas has prompted the listing of four Cicurina species (Araneae, Dictynidae) as US Federally Endangered. A major constraint in the management of these taxa is the extreme rarity of adult specimens, which are required for accurate species identification. We report a first attempt at using mitochondrial DNA (mtDNA) sequences to accurately identify immature Cicurina specimens. This identification is founded on a phylogenetic framework that is anchored by identified adult and/or topotypic specimens. Analysis of approximately 1 kb of cytochrome oxidase subunit I (CO1) mtDNA data for over 100 samples results in a phylogenetic tree that includes a large number of distinctive, easily recognizable, tip clades. These tip clades almost always correspond to a priori species hypotheses, and show nonoverlapping patterns of sequence divergence, making it possible to place species names on a number of immature specimens. Three cases of inconsistency between recovered tip clades and a priori species hypotheses suggest possible introgression between cave-dwelling Cicurina, or alternatively, species synonymy. Although species determination is not possible in these instances, the inconsistencies point to areas of taxonomic ambiguity that require further study. Our molecular phylogenetic sample is largest for the Federally Endangered C. madla. These data suggest that C. madla occurs in more than twice the number of caves as previously reported, and indicate the possible synonymy of C. madla with C. vespera, which is also Federally Endangered. Network analyses reveal considerable genetic divergence and structuring across caves in this species. Although the use of DNA sequences to identify previously 'unidentifiable' specimens illustrates the potential power of molecular data in taxonomy, many other aspects of the same dataset speak to the necessity of a balanced taxonomic approach.

Animals↗

Development of an efficient method for the isolation of factors involved in gene transcription during rice embryo development.

Summary An efficient yeast-based system was developed for the isolation of plant cDNAs encoding transcription factors (TFs) and proteins with transcription activation functions (co-activators). The system consists of two vectors: (i) a reporter vector (pG221) harboring the iso-1-cytochrome c (CYC1) core promoter and the beta-galactosidase (lacZ) gene; and (ii) a cDNA library construction vector (pYF503), which yields a library of plant peptides fused to the GAL4-binding domain (GAL4-BD). Expression of a peptide harboring the characteristics of a transcriptional activator leads to expression of lacZ, allowing for selection of relevant colonies. TFs during rice embryo development were isolated through this system. Approximately 200 confirmed positive colonies were obtained from screening 10(6) yeast colonies, and sequence analysis of conserved domains identified 75 independent cDNAs, 20 of which encoded plant TFs or co-activators, including members of the APETALA2 (AP2)/ethylene-responsive element-binding protein (EREBP), MYB and growth-regulating factor (GRF) families. Peptides encoded by 13 of the isolated cDNAs were classified as potential TFs or co-activators because of the presence of conserved TF-like domains. Additionally, 2, 11, and 13 clones encoded kinases, chromosome-related proteins, and unknown proteins, respectively, while the remaining 16 cDNAs were associated with specific functions seemingly unrelated to TFs. Expression pattern analysis of selected TF-encoding genes via RT-PCR revealed that these genes were expressed during seed development, with differential transcription observed during various stages. This work provides informative hints for further study of the regulatory mechanism of rice seed development and illustrates an identification strategy that will be of practical value for the isolation of TFs and co-activators associated with specific plant developmental processes.

Amino Acid Sequence↗

Targeted mutagenesis of the human papillomavirus type 16 E2 transactivation domain reveals separable transcriptional activation and DNA replication functions.

The E2 gene products of papillomavirus play key roles in viral replication, both as regulators of viral transcription and as auxiliary factors that act with E1 in viral DNA replication. We have carried out a detailed structure-function analysis of conserved amino acids within the N-terminal domain of the human papillomavirus type 16 (HPV16) E2 protein. These mutants were tested for their transcriptional activation activities as well as transient DNA replication and E1 binding activities. Analysis of the stably expressed mutants revealed that the transcriptional activation and replication activities of HPV16 E2 could be dissociated. The 173A mutant was defective for the transcriptional activation function but retained wild-type DNA replication activity, whereas the E39A mutant wild-type transcriptional activation function but was defective in transient DNA replication assays. The E39A mutant was also defective for HPV16 E1 binding in vitro, suggesting that the ability of E2 protein to form a complex with E1 appears to be essential for its function as an auxiliary replication factor.

Amino Acid Sequence↗

Sequence-related human proteins cluster by degree of evolutionary conservation.

Gene duplication followed by adaptive evolution is thought to be a central mechanism for the emergence of novel genes. To illuminate the contribution of duplicated protein-coding sequences to the complexity of the human genome, we study the connectivity of pairwise sequence-related human proteins and construct a network (N) of linked protein sequences with shared similarities. We find that (i) the connectivity distribution P (k) for k sequence-related proteins decays as a power law P (k) approximately k(-gamma) with gamma approximately 1.2 , (ii) the top rank of N consists of a single large cluster of proteins ( approximately 70%) , while bottom ranks consist of multiple isolated clusters, and (iii) structural characteristics of N show both a high degree of clustering and an intermediate connectivity ("small-world" features). We gain further insight into structural properties of N by studying the relationship between the connectivity distribution and the phylogenetic conservation of proteins in bacteria, plants, invertebrates, and vertebrates. We find that (iv) the proportion of sequence-related proteins increases with increasing extent of evolutionary conservation. Our results support that small-world network properties constitute a footprint of an evolutionary mechanism and extend the traditional interpretation of protein families.

Chromosome Mapping↗