Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

A rat brain cDNA encoding the neurotransmitter transporter with an unusual structure.

A rat cDNA clone encoding the novel membrane protein of the neurotransmitter transporters family was cloned and sequenced. The cDNA was identified as a transcript of the gene NTT4 of which a partial genomic clone was previously sequenced. Alignment of the amino acid sequence of NTT4 with other members of the neurotransmitter transporter family revealed a marked deviation from the conserved structure of all other members of the family. The largest extracellular loop with a potential glycosylation site was identified between membrane segments 7 and 8. The protein retains the common glycosylated loop between transmembrane helices 3 and 4 in all members of the family. The transcript of NTT4 was found exclusively in the central nervous system and is more abundant in the cerebellum and the cerebral cortex.

Amino Acid Sequence↗

Nucleotide sequence of the triosephosphate isomerase gene from Aspergillus nidulans: implications for a differential loss of introns.

A functional cDNA from Aspergillus nidulans encoding triosephosphate isomerase (TPI) was isolated by its ability to complement a tpi1 mutation in Saccharomyces cerevisiae. This cDNA was used to obtain the corresponding gene, tpiA. Alignment of the cDNA and genomic DNA nucleotide sequences indicated that tpiA contains five introns. The intron positions in the tpiA gene were compared with those in the TPI genes of human, chicken, and maize. One intron is present at an identical position in all four organisms, two other introns are located in similar positions in A. nidulans and maize, and the remaining two introns are unique to A. nidulans. These Aspergillus-specific introns are located in regions of the protein that were predicted to be interrupted by introns based on analysis of a Go plot of chicken TPI. These comparisons are discussed in relation to the evolution of introns within TPI genes.

Amino Acid Sequence↗

Entamoeba histolytica ribosomal RNA genes are carried on palindromic circular DNA molecules.

Highly abundant DNA fragments obtained after restriction enzyme digests of nuclear DNA of Entamoeba histolytica strain HM-1:IMSS have been cloned and characterized. Northern blot hybridization to E. histolytica rRNA and sequence analysis identified the abundant DNAs as ribosomal DNA containing species. Several overlapping clones containing these abundant DNAs were isolated from 4 different genomic libraries of E. histolytica. Alignment of the restriction maps was consistent with a circular molecule, about 24.6 kilobase pairs (kb) in size. Nuclease BA131 digestion provided additional evidence for the circular nature of this DNA. The ribosomal DNA molecule contains two large inverted repeat-regions, each at least 5.2 kb in length. Sequence analysis of clone R715 revealed homology to the large rRNA units of various eukaryotic organisms. This clone was located in both inverted repeats, suggesting two rRNA cistrons per molecule. The inverted repeats are flanked by stretches of DNA which contain tandemly reiterated sequences. Southern blot analysis of E. histolytica nuclear DNA revealed the presence of two populations of molecules. These molecules have identical arrangements of restriction sites, but differ in size (0.7 kb) in a fragment containing tandemly reiterated sequences. Analysis of E. histolytica nuclear DNA by electron microscopy also revealed circular molecules. These molecules are about 26.6 kb +/- 0.5 kb in size and contain structural features predicted by the restriction map of the extrachromosomal ribosomal DNA of E. histolytica.

Animals↗

Characterization of the transcription start site of the ACTH receptor gene: presence of an intronic sequence in the 5'-flanking region.

Corticotropin (ACTH) regulates glucocorticoid production through specific receptors on the adrenal cortex. Analysis of the ACTH receptor mRNA in human adrenal has revealed the presence of five transcripts ranging from 1.8 to 11 kilobases (kb). Characterization of the 5'-untranslated regions (UTRs) of the ACTH receptor mRNA demonstrated the presence of one major initiation site of transcription 177 bp away from the ATG codon. Analysis of this 5' sequence showed a perfect alignment with the previously described genomic sequence until position -128 bp from the ATG. The upstream 49-bp sequence was divergent, suggesting the occurrence of a splicing and indicating the presence of an intronic sequence in the UTRs, as well as the presence of an upstream exon containing this 49-bp sequence and located at least 1.8 kb away from the exon encoding the protein.

Adrenal Glands↗

Enzymatic description of the anhydrofructose pathway of glycogen degradation II. Gene identification and characterization of the reactions catalyzed by aldos-2-ulose dehydratase that converts 1,5-anhydro-D-fructose to microthecin with ascopyrone M as the intermediate.

The anhydrofructose pathway describes the degradation of glycogen and starch to metabolites via 1,5-anhydro-D-fructose (1,5AnFru). Enzymes that form 1,5AnFru, ascopyrone P (APP), and ascopyrone M (APM) have been reported from our laboratory earlier. In the present study, APM formed from 1,5AnFru was found to be the intermediate to the antimicrobial microthecin. The microthecin forming enzyme from the fungus Phanerochaete chrysosporium proved to be aldos-2-ulose dehydratase (AUDH, EC 4.2.1.-), which was purified and characterized for its enzymatic and catalytic properties. The purified AUDH showing a molecular mass of 97.4 kDa on SDS-PAGE was partially sequenced. Total 332 amino acid residues in length were obtained, representing some 37% of the AUDH protein. The obtained amino acid sequences showed no homology to known proteins but to an unannotated DNA sequence in Scaffold 62 of the published genome of the fungus. The alignment revealed three introns of the identified AUDH gene (Audh; ph.chr), thus the first gene coding for a neutral sugar dehydratase is identified. AUDH was found to be a bi-functional enzyme, being able to dehydrate 1,5AnFru to APM and further isomerizing the APM formed to microthecin. The optimal pH for the formation of APM and microthecin was pH 5.8 and 6.8, respectively. AUDH showed 5 fold higher activity toward 1,5AnFru than toward its analogue glucosone, when tested at concentrations from 0.6 mM to 0.2 M. Based on the characteristic UV absorbance of microthecin (230 nm) and APM (262 nm) assay methods were developed for the microthecin forming enzymes.

Amino Acid Sequence↗

Rapid identification of potato virus Y strains by one-step triplex RT-PCR.

A one-step triplex RT-PCR method was characterised that allows rapid, strain-specific detection of potato virus Y (PVY) occurring on potato: PVY(N), PVY(O), PVY(NTN) (recombinant isolates), PVY(N)Wi and PVY(C). Three specific primer pairs were designed on aligned PVY sequences available from genomic data banks. The specificity of the selected primers was first examined by simplex RT-PCR with a large number of PVY reference isolates. Two fragments of 0.44 and 1.11kb were amplified for PVY(N) and non-recombinant PVY(NTN) isolates, two fragments of 0.53 and 0.66kb for PVY(O) isolates, a single fragment of 0.44kb for recombinant PVY(NTN) isolates, a 0.66kb fragment for PVY(C) isolates and a 0.53kb fragment for PVY(N)Wi isolate. The primers were then combined in a one-step triplex RT-PCR reaction, optimised stepwise and validated with the reference isolates. The great similarity between the genomes of PVY(N) and non-recombinant PVY(NTN) prevented their differentiation using this method. No fragments were amplified with samples infected by non-related potato viruses, as well as with samples from healthy tobacco and potato plants. The one-step triplex RT-PCR described here fastens specific detection of PVY strains that are otherwise only distinguishable by combined serological and biological assays.

DNA Primers↗

Comparative analysis of the Band 4.1/ezrin-related protein tyrosine phosphatase Pez from two Drosophila species: implications for structure and function.

The FERM-PTPs are a group of proteins that have FERM (Band 4.1, ezrin, radixin, moesin homology) domains at or near their N-termini, and PTP (protein tyrosine phosphatase) domains at their C-termini. Their central regions contain either PSD-95, Dlg, ZO-1 homology domains or putative Src homology 3 domain binding sites. The known FERM-PTPs fall into three distinct classes, which we name BAS, MEG, and PEZ, after representative human PTPs. Here we analyze Pez, a novel gene encoding the single PEZ-class protein present in Drosophila. Pez cDNAs were sequenced from the distantly related flies Drosophila melanogaster and Drosophila silvestris, and found to be highly conserved except in the central region, which contains at least 21 insertions and deletions. Comparison of fly and human Pez reveals several short conserved motifs in the central region that are likely protein binding sites and/or phosphorylation sites. We also identified novel invertebrate members of the BAS and MEG classes using genome data, and generated an alignment of vertebrate and invertebrate FERM domains of each class. 'Specialized' residues were identified that are conserved only within a given class of PTPs. These residues highlight surface regions that may bind class-specific ligands; for PEZ, these residues cluster on and near FERM subdomain F1. Finally, the PTP domain of fly Pez was modeled based on known PTP tertiary structures, and we conclude that Pez is likely a functional phosphatase despite some unusual features of the active site cleft sequences. Biochemical confirmation of this hypothesis and genetic analysis of Pez are currently underway.

Amino Acid Sequence↗

Complete sequence and characterization of the channel catfish mitochondrial genome.

In order to support analysis of channel catfish populations and genetic improvement programs, the channel catfish, Ictalurus punctatus, mitochondrial genome was completely sequenced and revealed gene structure and gene order common to vertebrates. Nucleotide sequence comparisons of cytochrome b (Cytb) and cytochrome c oxidase subunit 1 (COI) demonstrated genetic separation of the genera Ictalurus, Pylodictis and Ameiurus consistent with the taxonomic classification within Ictaluridae. The ictalurid Cytb nucleotide sequences were significantly different from a putative channel catfish Cytb sequence in GenBank. Genetic relationships based on mitochondrial DNA sequences indicated the value of channel catfish in genomic comparisons between teleosts. Pairwise alignment of DNA sequences revealed conservation of regulatory sequences in the D-loop region with other vertebrates. Analysis of D-loop sequences in commercial populations and a research strain revealed 28 polymorphic sites and 33 D-loop haplotypes. Sequence analysis revealed clustering of haplotypes within commercial farms and the USDA103 research line, but D-loop haplotypes were not sufficient to discriminate the USDA103 fish from commercial catfish.

Analysis of Variance↗

ASD: the Alternative Splicing Database.

Alternative splicing is widespread in mammalian gene expression, and variant splice patterns are often specific to different stages of development, particular tissues or a disease state. There is a need to systematically collect data on alternatively spliced exons, introns and splice isoforms, and to annotate this data. The Alternative Splicing Database consortium has been addressing this need, and is committed to maintaining and developing a value-added database of alternative splice events, and of experimentally verified regulatory mechanisms that mediate splice variants. In this paper we present two of the products from this project: namely, a database of computationally delineated alternative splice events as seen in alignments of EST/cDNA sequences with genome sequences, and a database of alternatively spliced exons collected from literature. The reported splice events are from nine different organisms and are annotated for various biological features including expression states and cross-species conservation. The data are presented on our ASD web pages (http://www.ebi.ac.uk/asd).

Alternative Splicing↗

Mosaic structure and retropositional dynamics during evolution of subfamilies of short interspersed elements in African cichlids.

The African cichlid (AFC) family of short interspersed elements (SINEs) is found in the genomes of cichlid fish. The alignment of the sequences of 70 members of this family, isolated from such fish in Africa, revealed the presence of correlated changes in specific nucleotides (diagnostic nucleotides) that allowed us to categorize the various members into six subfamilies, which were designated Af1 through Af6. Dividing the SINE consensus sequence into a 5'-head and 3'-tail region, these subfamilies were defined by various combinations of four types of head region (A-D) and three types of tail region [X, Y, and (YX)], with each region of each type including unique diagnostic nucleotides. The observed structures of the subfamilies Af1 through Af6 were AX, AY, CY, A(YX), BY, and DX, respectively. The formation of such structures might have involved the shuffling of head or tail regions among preexisting and existing (or both) subfamilies of the AFC family (and, probably, even another SINE family or a pseudogene for a tRNA in the case of the Af6 subfamily) by recombination at the so-called core region during the course of evolution. By plotting the timing of the retroposition of individual members of each subfamily on a phylogenetic tree of AFCs, we found that the Af3 and Af6 subfamilies became active only recently in the evolutionary history of these fish. The integrity of the 3'-tails of SINEs, which are, apparently, recognized by reverse transcriptase, has been reported to be indispensable for retention of retropositional activity. Therefore, we postulate that recombination might have been involved in the apparent recent activation of the retroposition of the Af3 and Af6 subfamilies via introduction of active tails (types Y and X, respectively) into potential ancestral sequences that might have had inactive tails. If this hypothesis is correct, shuffling of tail regions among subfamilies by recombination at the core region might have played a role in the recycling of dead copies of AFC SINEs.

Animals↗

Disruption of GAD1 protein architecture by a novel missense variant in a consanguineous family with autosomal recessive intellectual disability.

BACKGROUND: Intellectual disability represents a heterogeneous group of neurodevelopmental disorders marked by significant impairments in intellectual functioning and adaptive behavior. Among the various causes, genetic factors play a major role, with autosomal recessive intellectual disability (ARID) constituting a genetically diverse subgroup. ARID is prevalent in consanguineous families and arises from homozygous mutations that disrupt critical genes involved in brain development and function. OBJECTIVE: This study aimed to identify disease-causing genetic variants responsible for ARID in a consanguineous Pakistani family and to evaluate the structural and functional impact of a novel variant identified in GAD1 through protein modeling. METHODS: A consanguineous family affected with intellectual disability was enrolled. Whole-exome sequencing was performed on an affected individual, followed by bioinformatics analysis including alignment to the GRCh38 reference genome, variant calling, and annotation. Variants were filtered based on rarity, predicted functional impact, and autosomal recessive inheritance pattern. Candidate variants were validated and assessed by Sanger sequencing and segregation analysis. Protein modeling was performed to evaluate the structural impact of the identified variant. RESULTS: A novel homozygous missense variant NM_000817:c.1700G>A;p.Arg567Gln in GAD1 was identified. Segregation analysis confirmed co-segregation of the variant with the affected phenotype. Protein modeling suggested that the variant may disrupt GAD1 enzymatic function involved in gamma-aminobutyric acid synthesis. CONCLUSION: This study emphasizes the significance of genetic investigation in familial cases and the crucial role that GAD1 mutations play in neurodevelopmental disorders with intellectual disability. The results advance the knowledge of molecular causes of ARID and broaden the mutational range.

Pakistani↗

Cloning and disruption of the ornithine decarboxylase gene of Ustilago maydis: evidence for a role of polyamines in its dimorphic transition.

The gene encoding ornithine decarboxylase (ODC) from Ustilago maydis was cloned. A conserved PCR product amplified from U. maydis DNA was synthesized and used to screen a genomic library of the fungus. Alignment of its deduced protein sequence with those of other cloned ODCs showed a high degree of homology. Gene replacement was obtained by removal of a central part of the gene and insertion of the hygromycin resistance cassette. The null mutant thus obtained displayed no ODC activity and behaved as a polyamine auxotroph. This result is evidence that a single ODC gene exists in the fungus, and that U. maydis utilizes the ODC pathway as the only mechanism for polyamine biosynthesis. When grown in polyamine-containing media, the null mutant accumulated a polyamine pool which further sustained its normal rate of growth in polyamine-free media for approximately 12-16 h. When putrescine concentrations lower than 0.5 mM were employed, the mutant grew at a normal rate but was unable to engage in the dimorphic transition. Under conditions favourable for mycelial growth, the mutant grew with a yeast-like morphology in liquid media, and formed smooth colonies consisting of yeast cells on solid media. Reversion to normal dimorphic phenotype required high concentrations of putrescine or spermidine. These results are evidence that concentrations of polyamines higher than those necessary to sustain vegetative growth are required for the dimorphic transition in U. maydis.

Amino Acid Sequence↗

Gene lineages and eastern North American palaeodrainage basins: phylogeography and speciation in salamanders of the Eurycea bislineata species complex.

Contemporary North American drainage basins are composites of formerly isolated drainages, suggesting that fragmentation and fusion of palaeodrainage systems may have been an important factor generating current patterns of genetic and species diversity in stream-associated organisms. Here, we combine traditional molecular-phylogenetic, multiple-regression, nested clade, and molecular-demographic analyses to investigate the relationship between phylogeographic variation and the hydrogeological history of eastern North American drainage basins in semiaquatic plethodontid salamanders of the Eurycea bislineata species complex. Four hundred forty-two sequences representing 1108 aligned bases from the mitochondrial genome are reported for the five formally recognized species of the E. bislineata complex and three outgroup taxa. Within the in-group, 270 haplotypes are recovered from 144 sampling locations. Geographic patterns of mtDNA-haplotype coalescence identify 13 putatively independent population-level lineages, suggesting that the current taxonomy of the group underestimates species-level diversity. Spatial and temporal patterns of phylogeographic divergence are strongly associated with historical rather than modern drainage connections, indicating that shifts in major drainage patterns played a pivotal role in the allopatric fragmentation of populations and build-up of lineage diversity in these stream-associated salamanders. More generally, our molecular genetic results corroborate geological and faunistic evidence suggesting that palaeodrainage connections altered by glacial advances and headwater erosion occurring between the mid-Miocene and Pleistocene epochs explain regional patterns of biodiversity in eastern North American streams.

Animals↗

Multivariate Effects of SNPs on Environmental Streptococcal Mastitis Evaluated With an NGS-Based Association Study Using Targeted Resequencing in the Bovine MHC Region.

Mastitis is an inflammatory reaction caused by bacterial infection of the teat, and a relationship between its onset and cattle major histocompatibility complex (BoLA) region has been reported. However, no comprehensive genetic analysis of mastitis caused by environmental streptococci has been reported. Here, we resequenced the BoLA region using a hybridisation capture target next-generation sequencing (NGS) method to identify disease susceptibility markers mapped to the BoLA region in environmental streptococcal mastitis. This study examined 75 cows with mastitis caused by environmental streptococci selected from 1641 cows with mastitis and 222 healthy cows without mastitis in Japan. Targeted sequences obtained from MiSeq NGS were aligned to the bovine reference genome (ARS-UCD1.2/bosTau9), and 2,920,355 variants were detected within the BoLA region of the 297 Holstein cattle. In an association study using 2264 variants after quality control, the top 20 variants with the lowest P values were selected and assigned to the 18 surrounding candidate genes, and a gene network analysis of these genes resulted in the narrowing down of five candidate genes POU5F1, IER3, GNL1, ABCF1, and PRR3. Multivariate effect analysis of all 6 SNPs associated with these 5 genes revealed that they were significantly correlated with mastitis, indicating that they were useful for classification of mastitis-resistant and mastitis-susceptible cattle. This is the first report to identify SNPs associated with environmental streptococcal mastitis with an NGS-based association study using targeted resequencing in the BoLA region, and understanding host factors may provide important clues for mastitis control.

Animals↗

Lineage-associated differences in adenine methylation patterns of mammalian-associated Campylobacter fetus isolates: a possible role for epigenetic factors in host tropism and pathogenesis.

Mammalian Campylobacter fetus (CF) is divided into two subspecies, C. fetus fetus (CFF) and C. fetus venerealis (CFV), the latter being bovine-adapted and responsible for the notifiable disease bovine genital campylobacteriosis (BGC). Differentiation between CF subspecies has traditionally been undertaken by a few biochemical tests, but these are complicated by the existence of a biotype, C. fetus venerealis intermedius (CFVi), which shares attributes of both CFF and CFV. Molecular methods targeting specific genes have gained acceptance for more accurate subtype identification and align well with whole-genome analysis. However, limited genomic diversity between subtypes has confounded efforts to understand the genetic basis for differential host tropism and pathogenesis of these organisms. A previous study of a small cohort of C. fetus isolates suggested that dam gene coding variations might correlate with CF subtype. Accordingly, this study examines a cohort of 331 C. fetus genomes, representative of all seven phylogenetic groups for their complement of adenine methylases and the genomic motifs they target in representative isolates. All CF isolates retained a cfeM1 gene, the presence of which correlates with RAATTY methylation, while seven other adenine methylase genes exhibited distinct cladal distributions. Notably, a cjeM1 gene appears to target the CCAN7TAG/CTAN7TGG motif in CFV and CFVi isolates only. Given the increasing recognition of the impact of adenine methylation on bacterial-host interactions, further exploration of the role of adenine methylation in C. fetus pathogenesis could reveal mechanisms contributing to BGC and thus aid in its eradication.IMPORTANCECampylobacter fetus remains an important zoonotic pathogen, for which a better understanding of its host tropism and pathogenesis is sought. However, the limited genomic variation observed between subtypes has to date confounded efforts in this regard. This study suggests that an alternative approach that examines epigenetic differences between subtypes, specifically adenine methylation patterns, may reveal mechanisms critical to the pathologies of these organisms.

Animals↗

Chromosomal mapping and expression of the human cyr61 gene in tumour cells from the nervous system.

AIMS: To characterise the human cyr61 gene (cyr61H) and determine its chromosomal locality. To compare expression of cyr61H in human tumour cell lines with that of two other structurally related genes, novH (nephroblastoma overexpressed gene) and CTGF (connective tissue growth factor), that are likely to play a role in the control of cell proliferation and differentiation. METHODS: To isolate the human cyr61 gene, placental genomic and HeLa cDNA libraries were screened with murine cyr61 cDNA. The nucleotide sequence of the complete cyr61H cDNA was established. Both Southern blotting of a panel of somatic cell hybrids and in situ hybridisation on chromosomes were performed to map the cyr61H gene. Expression of cyr61H, novH, CTGF, and novH was analysed by northern blotting in both human neuroblastomas and glioblastoma cell lines. RESULTS: Genomic and cDNA clones encompassing the cyr61H gene were isolated and characterised. Comparison of mouse and human cyr61 sequences indicated that their genomic organisation is highly conserved. Alignment of coding sequences highlighted the conservation of cyr61 regions that might be critical for its biological function. The data showed that the cyr61H gene is assigned to chromosome 1p22.3 and that different levels of cyr61H, CTGF, and novH mRNA have been detected in several human tumour cell lines derived from the nervous system. CONCLUSIONS: The human cyr61 gene belongs to an emerging family of genes including CTGF/fisp12 and nov. The murine cyr61 encodes an extracellular cysteine rich protein that exhibits chemotactic activity, promotes attachment and spreading of cells, and potentiates the mitogenic effect of growth factors. Assignment of the cyr61H gene to chromosome 1p22.3 will allow studies to determine whether human pathologies derived from the nervous system or from other tissues are associated with chromosomal abnormalities involving this region. Although the coding regions of cyr61H, CTGF, and novH are highly homologous, a growing body of evidence suggests that expression of these genes is regulated differentially, and that a balance between expression of these genes might represent a key element in determining the stage of differentiation and/or the malignant potential of tumour cells.

Animals↗

The genes encoding granule-bound starch synthases at the waxy loci of the A, B, and D progenitors of common wheat.

Three genes encoding granule-bound starch synthase (wx-TmA, wx-TsB, and wx-TtD) have been isolated from Triticum monococcum (AA), and Triticum speltoides (BB), by the polymerase chain reaction (PCR) approach, and from Triticum tauschii (DD), by screening a genomic DNA library. Multiple sequence alignment indicated that the wx-TmA, wx-TsB, and wx-TtD genes had the same extron and (or) intron structure as the previously reported waxy gene from barley. The lengths of the three wx-TmA, wx-TsB, and wx-TtD genes were 2834 bp, 2826 bp, and 2893 bp, respectively, each covering 31 bp in the untranslated leader and the entire coding region consisting of 11 exons and 10 introns. The three genes had identical lengths of exons, except exonl, and shared over 95% identity with each other within the exon regions. The majority of introns were significantly variable in length and sequence, differing mainly in length (1-57 bp) as a result of insertion and (or) deletion events. The deduced amino acid sequence from these three genes indicated that the mature WX-TMA, -TSB, and -TTD proteins contained the same number of amino acids, but differed in predicted molecular weight and isoelectric point (pI) due to amino acid substitutions (13-18). The predicted physical characteristics of the WX proteins matched the respective proteins in wheat very closely, but the match was not perfect. Furthermore the exon5 sequences of the wx-TmA, wx-TsB, and wx-TtD genes were different from a cDNA encoding a waxy gene of common wheat previously reported. The striking difference was that an insertion of 11 amino acids occurred in the cDNA sequence that could not be observed in the exons of the A, B, and D genes. It was noted, however, that the 3' end of intron4 of these genes could account for the additional 11 amino acids. The sequence information from the available waxy genes identified the intron4-exon5-intron5 region as being diagnostic for sequence variation in waxy. The sequence variation in the waxy genes provides the basis for primer design to distinguish the respective genes in common wheat, and its progenitors, using PCR.

Amino Acid Sequence↗

Identification of SmtB/ArsR cis elements and proteins in archaea using the Prokaryotic InterGenic Exploration Database (PIGED).

Microbial genome sequencing projects have revealed an apparently wide distribution of SmtB/ArsR metal-responsive transcriptional regulators among prokaryotes. Using a position-dependent weight matrix approach, prokaryotic genome sequences were screened for SmtB/ArsR DNA binding sites using data derived from intergenic sequences upstream of orthologous genes encoding these regulators. Sixty SmtB/ArsR operators linked to metal detoxification genes, including nine among various archaeal species, are predicted among 230 annotated and draft prokaryotic genome sequences. Independent multiple sequence alignments of putative operator sites and corresponding winged helix-turn-helix motifs define sequence signatures for the DNA binding activity of this SmtB/ArsR subfamily. Prediction of an archaeal SmtB/ArsR based upon these signature sequences is confirmed using purified Methanosarcina acetivorans C2A protein and electrophoretic mobility shift assays. Tools used in this study have been incorporated into a web application, the Prokaryotic InterGenic Exploration Database (PIGED; http://bioinformatics.uwp.edu/~PIGED/home.htm), facilitating comparable studies. Use of this tool and establishment of orthology based on DNA binding signatures holds promise for deciphering potential cellular roles of various archaeal winged helix-turn-helix transcriptional regulators.

Archaea↗