Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

An operator associated with autoregulation of the repressor gene in actinophage phiC31 is found in highly conserved copies in intergenic regions in the phage genome.

Previous reports have suggested that the repressor gene, c, of phiC31 is autoregulated and that likely operators are conserved inverted repeat sequences (CIRs1&2) located just upstream of the promoters, cp1 and cp2. Evidence is now presented that the CIRs 1&2 are indeed binding sites for one of the three inframe, N-terminally different protein isoforms of 42, 54 and 74 kDa produced by the c gene. A cp1-aphII fusion was repressed in a Streptomyces coelicolor A3(2) phiC31 lysogen and characterisation of an operator-constitutive (Oc) mutant showed a single mutation in CIR-1. CIR-1 containing fragments were retarded in electrophoresis gels by the 42 kDa repressor protein isoform and this retardation was inhibited by the addition of competing DNA fragments containing either CIR-1 or CIR-2. Using a combination of Southern blotting and analysis of available DNA sequence we also show that at least 18 copies of the CIRs are present throughout the phiC31 genome. Alignment of 9 CIR sequences showed that 8 contained a perfectly conserved 17 bp core whilst the exception had a single mismatch. The core includes a 16 bp inverted repeat (IR), and is usually part of a more extensive and less highly conserved palindrome. When superimposed on a previously derived transcription map of the early region, the CIRs lie in intergenic regions associated with transcription initiation and/or termination.

Bacteriophages↗

Genomic epidemiology of enteropathogenic Escherichia coli in southwestern Nigeria.

BACKGROUND: Enteropathogenic Escherichia coli (EPEC) are etiological agents of diarrhea. We studied the genetic diversity and virulence factors of EPEC in southwestern Nigeria, where this pathotype is rarely characterized. METHODOLOGY/PRINCIPAL FINDINGS: EPEC isolates (n&#x2009;=&#x2009;96) recovered from recent southwestern Nigeria diarrhea case-control studies were whole genome-sequenced using Illumina technology. Genomes were assembled using SPAdes and quality was evaluated using QUAST. Virulencefinder, Ectyper, and ResFinder were used to identify virulence genes, serotypes, and resistance genes. Multilocus sequence typing was done by STtyping. Single nucleotide polymorphisms (SNPs) were called out of whole genome alignment using SNP-sites and a phylogenetic tree was constructed using IQtree. Thirty-nine of the 96(40.6%) EPEC isolates were from diarrhea cases diarrhea. Nine isolates from diarrhea patients and four from healthy controls were typical EPEC, harboring bundle-forming pilus (bfp) genes whilst the rest were atypical EPEC. There were 15 EPEC-EAEC hybrids. Atypical serotypes O71:H19 (16, 16.6%), O108:H21 (6, 6.3%), O157:H39 (5, 5.2%), and O165:H9 (4, 4.2%) were the most prevalent; only 8 (8.3%) isolates belonged to classical EPEC serovars. The largest, ST517 clade harbored multiple siderophore and serine protease autotransporter genes and included an O71:H19 subclade <10 SNPs apart, representing a likely outbreak involving 15 children, four with diarrhea. Likely outbreaks, of typical O119:H6(ST28) and atypical O127:H29(ST7798) were additionally identified. CONCLUSION/SIGNIFICANCE: EPEC circulating in southwestern Nigeria are diverse and differ substantially from well-characterized lineages seen previously elsewhere. EPEC carriage and outbreaks could be commonplace but are largely undetected, hence, unreported, and require genomic surveillance for identification.

Nigeria↗

Comparative genomics of transcriptional regulation in yeasts and its application to identification of a candidate alpha-isopropylmalate transporter.

Conservation rates in non-protein-coding regions of five yeast genomes of the genus Saccharomyces were analyzed using multiple whole-genome alignments. This analysis confirmed previously shown decrease in conservation rates observed immediately upstream of the translation start point and downstream of the stop-codon. Further, there was a sharp conservation peak in the upstream regions likely related to the core promoter (-35 bp to +35 bp around TSS) and a conservation peak downstream of the stop-codon whose function is not yet clear. Regulation of leucine and methionine biosynthesis controlled by the global regulator Gcn4p and pathway-specific regulators was analyzed in detail. A candidate alpha-isopropylmalate carrier, YOR271cp, was identified based on conservation of Leu3p binding sites, analysis of ChIP-chip data, protein localization and sequence similarity.

Chromosome Mapping↗

A cereal centromeric sequence.

We report the identification of a family of sequences located by in situ hybridisation to the centromeres of all the Triticeae chromosomes studied, including the supernumerary and midget chromosomes, the centromeres of all maize chromosomes and the heterochromatic regions of rice chromosomes. This family of sequences (CCS1), together with the cereal genome alignments, will allow the evolution of the cereal centromeres and their sites to be studied. The family of sequences also shows homology to the CENP-B box. The centromeres of the cereal species and the proteins that interact with them can now be characterised.

Autoantigens↗

Mouse USF1 gene cloning: comparative organization within the c-myc gene family.

Upstream stimulatory factors (USF/MLTF) belong to the c-myc family of transcription factors. Through binding to target DNA as dimers, the ubiquitous USF proteins regulate a variety of genes. USF proteins are encoded by two genes, USF1 and USF2. Protein sequences of USF1 and 2 are highly homologous across species, suggesting functional conservation. To determine whether the genomic organization was conserved between USF1 and USF2, we isolated the murine USF1 gene and characterized its genomic structure. Both genes are similarly organized in 10 exons spanning over 10 kbp. By the 5'-rapid amplification of cDNA ends and S1 nuclease mapping methods, exon 1 was defined and the transcription initiation sites were mapped. The sequence of 8 kb of the gene, including 1.75 kb of 5'-flanking DNA, was determined. The promoter region is GC rich and lacks a typical TATA or CCAAT element. Strikingly, a comparison of the murine and human untranslated sequences reveals regions that exhibit greater than 73% sequence identity. A genomic alignment of the dimerization and DNA binding domains is presented for five genes of the c-myc family, suggesting a hypothetical common ancestor gene.

Amino Acid Sequence↗

Exon-based mapping of microarray probes: recovering differential gene expression signal in underpowered hypoxia experiment.

There is an immense collection of underpowered Affymetrix gene array experiments. Although a majority of these experiments generated biologically feasible results, the considerable fraction of assays failed to identify expected transcriptional changes. There is an unused potential of Affymetrix probe-set redundancy for common exonic and UTR regions. We hypothesized that group analysis of multiple probe-sets which hybridize to the same exon or UTR will increase array discriminating power of transcriptional changes. To test this hypothesis, we analyzed Affymetrix mouse probe-sets that share the same exon using blocking feature of the Significance Analysis of Microarrays (SAM). Two-thousand two-hundred one exon-sharing probe-sets targeting 1011 transcripts were identified by mapping 36701 MG-U74v2 probe-sets to genomic alignments of 3,971,086 known mouse transcripts. Using the blocking feature of SAM with an underpowered (two microarrays per experimental condition) mouse hypoxia-induced pulmonary hypertension model, we identified 24 genes that were significantly (FDR<5%) affected by hypoxia but were not detected by regular SAM. The relevance of the four newly identified genes (Mig6, F3, Bmp6, and Ndrg1) to known hypoxia-associated responses was confirmed by PubMatrix; and hypoxia-induced up-regulation of Mig6 expression was validated by real-time RT-PCR. We demonstrated that analysis of exon-sharing probe-sets allowed discovery of additional hypoxia-affected genes in an underpowered array experiment. This method will facilitate re-evaluation of existing underpowered Affymetrix gene expression profiles.

5' Untranslated Regions↗

Cloning, sequencing, structural and molecular biological characterization of placental protein 20 (PP20)/human thiamin pyrophosphokinase (hTPK).

Full-length cDNAs of placental protein 20 (PP20) were cloned by screening a human placental cDNA library, which encode a 243 amino acid protein, identical to human thiamin pyrophosphokinase (hTPK) as confirmed by protein sequence analysis. Genomic alignment showed that the PP20/hTPK gene contains 9 exons. It is abundantly expressed in placenta, as numerous EST clones were identified. As thiamine metabolism deficiencies have been seen in placental infarcts previously, these indicate that PP20/hTPK may have a role in placental diseases. Analysis of the 1kb promoter region showed numerous putative transcription factor binding sites, which might be responsible for the ubiquitous PP20/hTPK expression. This may also be in accordance with the presence of the protein in tissues responsible for the regulation of the exquisite balance between cell division, differentiation and survival. TPK activity of the purified and recombinant protein was proved by mass spectrometry with electrospray ionization. By Western blot, PP20/hTPK was found in all human normal and tumorous adult and fetal tissues in nearly equal amounts, but not in sera. By immunohistochemical and immunofluorescent confocal imaging methods, diffuse labelling in the cytoplasm of the syncytiotrophoblasts and weak staining of the trophoblasts were observed, and the amount of PP20/hTPK decreased from the first trimester to the end of gestation. A 3D model of PP20/hTPK was computed (PDB No.: 1OLY) by homology modelling. A high degree of structural homology showed that the thiamin binding site was highly similar to that of the mouse enzyme, but highly different from the bacterial ones. Comparison of the catalytic centre sequences revealed differences, raising the possibility of designing new drugs which specifically inhibit bacterial and fungal enzymes without affecting PP20/hTPK and offering the possibility for safe antimicrobial therapy during pregnancy.

Adult↗

Variola and camelpox virus-specific sequences are part of a single large open reading frame identified in two German cowpox virus strains.

A large open reading frame (ORF) has been identified in two German cowpox virus strains. The ORFs (5676 and 5679 nt, respectively) differ in 10 nucleotides, resulting in an amino acid homology of 99.8%. In searching GenBank nucleotide sequences (>90% identity) were present in several small ORFs in variola major, variola minor and camelpox virus genomes. Alignments revealed that these small ORFs are fragments of a large ORF. However, sequences of the ORF described here are entirely absent in the two cowpox virus reference strains. Databank analysis revealed amino acid identities (ranging from 25 to 39%) with so-called B22R-like poxviral proteins with unknown function encoded by several chordopoxviruses. Further sequencing of one cowpox virus strain under study identified an ORF (5790 nt) which displays high levels of nucleotide identity to ORFs present in several orthopoxvirus species. Taken together, the two cowpox viruses analyzed here contain one large ORF which is conserved within the genus Orthopoxvirus and a unique, more distantly related ORF of similar size, which is conserved in the subfamily Chordopoxvirinae.

Chordopoxvirinae↗

A significant fraction of conserved noncoding DNA in human and mouse consists of predicted matrix attachment regions.

Noncoding DNA in the human-mouse orthologous intergenic regions contains "islands" of conserved sequences, the functions of which remain largely unknown. We hypothesized that some of these regions might be matrix-scaffold attachment regions, MARs (or S/MARs). MARs comprise one of the few classes of eukaryotic noncoding DNA with an experimentally characterized function, being involved in the attachment of chromatin to the nuclear matrix, chromatin remodeling and transcription regulation. To test our hypothesis, we analyzed the co-occurrence of predicted MARs with highly conserved noncoding DNA regions in human-mouse genomic alignments. We found that 11% of the conserved noncoding DNA consists of predicted MARs. Conversely, more than half of the predicted MARs co-occur with one or more independently identified conserved sequence blocks. An excess of conserved predicted MARs is seen in intergenic regions preceding 5' ends of genes, suggesting that these MARs are primarily involved in transcriptional control.

Animals↗

The finished DNA sequence of human chromosome 12.

Human chromosome 12 contains more than 1,400 coding genes and 487 loci that have been directly implicated in human disease. The q arm of chromosome 12 contains one of the largest blocks of linkage disequilibrium found in the human genome. Here we present the finished sequence of human chromosome 12, which has been finished to high quality and spans approximately 132 megabases, representing approximately 4.5% of the human genome. Alignment of the human chromosome 12 sequence across vertebrates reveals the origin of individual segments in chicken, and a unique history of rearrangement through rodent and primate lineages. The rate of base substitutions in recent evolutionary history shows an overall slowing in hominids compared with primates and rodents.

Animals↗

Cloning and nucleotide sequence of the ispA gene responsible for farnesyl diphosphate synthase activity in Escherichia coli.

The molecular cloning and the determination of the nucleotide sequence of the ispA gene responsible for farnesyl diphosphate (FPP) synthase [EC 2.5.1.1] activity in Escherichia coli are described. E. coli ispA strains have temperature-sensitive FPP synthase, and the defective gene is located at about min 10 on the chromosome. The wild-type ispA gene was subcloned from a lambda phage clone containing the chromosomal fragment around min 10, picked up from the aligned genomic library of Kohara et al. [Kohara, Y., Akiyama, K., & Isono, K. (1987) Cell 50, 495-508]. The cloned gene was identified as the ispA gene by the recovery and amplification of FPP synthase activity in an ispA strain. A 1,452-nucleotide sequence of the cloned fragment was determined. This sequence specifies two open reading frames, ORF-1 and ORF-2, encoding proteins with the expected molecular weights of 8,951 and 32,158, respectively. A part of the deduced amino acid sequence of ORF-2 showed similarity to the sequences of eucaryotic FPP synthases and of crtE product of a photosynthetic bacterium. The plasmid carrying ORF-2 downstream of the lac promoter complemented the defect of FPP synthase activity of the ispA mutant, showing that the product encoded by ORF-2 is the ispA product. The maxicell analysis indicated that a protein of molecular weight 36,000, approximately consistent with the molecular weight of the deduced ORF-2-encoded protein, is the gene product.

Alkyl and Aryl Transferases↗

Primary structure of ribosomal proteins S3 and S7 from Manduca sexta.

We have isolated from Manduca sexta full-length cDNAs encoding proteins homologous to human ribosomal proteins S3 and S7. These are the first ribosomal protein sequences obtained from non-Dipteran insects. M. sexta ribosomal protein S3 has a molecular mass of 26,715 Da. Ribosomal protein S7 has a mass 21,870 Da. Both are basic proteins, with abundant Lys and Arg residues that may interact with ribosomal RNA in the ribosome. Southern blot hybridization suggests the presence of single genes for both ribosomal proteins in the M. sexta genome. Alignments with other S3 and S7 sequences available In the database indicate regions of the ribosomal proteins that have been the most highly conserved in evolution and may point to important functional regions in the proteins. Ribosomal protein S3 appears to be more highly conserved then ribosomal protein S7. This may be due to greater constraints on the structure of S3 because of its dual functions in translation as a ribosomal protein and in DNA repair in the nucleus.

Amino Acid Sequence↗

Characterization of Staphylococcus aureus SarA binding sites.

The staphylococcal accessory regulator locus (sarA) encodes a DNA-binding protein (SarA) that modulates expression of over 100 genes. Whether this occurs via a direct interaction between SarA and cis elements associated with its target genes is unclear, partly because the definitive characteristics of a SarA binding site have not been identified. In this work, electrophoretic mobility shift assays (EMSAs) were used to identify a SarA binding site(s) upstream of the SarA-regulated gene cna. The results suggest the existence of multiple high-affinity binding sites within the cna promoter region. Using a SELEX (systematic evolution of ligands by exponential enrichment) procedure and purified, recombinant SarA, we also selected DNA targets that contain a high-affinity SarA binding site from a random pool of DNA fragments. These fragments were subsequently cloned and sequenced. Randomly chosen clones were also examined by EMSA. These DNA fragments bound SarA with affinities comparable to those of recognized SarA-regulated genes, including cna, fnbA, and sspA. The composition of SarA-selected DNAs was AT rich, which is consistent with the nucleotide composition of the Staphylococcus aureus genome. Alignment of selected DNAs revealed a 7-bp consensus (ATTTTAT) that was present with no more than one mismatch in 46 of 56 sequenced clones. By using the same criteria, consensus binding sites were also identified upstream of the S. aureus genes spa, fnbA, sspA, agr, hla, and cna. With the exception of cna, which has not been previously examined, this 7-bp motif was within the putative SarA binding site previously associated with each gene.

Adhesins, Bacterial↗

Sulfur and nitrogen limitation in Escherichia coli K-12: specific homeostatic responses.

We determined global transcriptional responses of Escherichia coli K-12 to sulfur (S)- or nitrogen (N)-limited growth in adapted batch cultures and cultures subjected to nutrient shifts. Using two limitations helped to distinguish between nutrient-specific changes in mRNA levels and common changes related to the growth rate. Both homeostatic and slow growth responses were amplified upon shifts. This made detection of these responses more reliable and increased the number of genes that were differentially expressed. We analyzed microarray data in several ways: by determining expression changes after use of a statistical normalization algorithm, by hierarchical and k-means clustering, and by visual inspection of aligned genome images. Using these tools, we confirmed known homeostatic responses to global S limitation, which are controlled by the activators CysB and Cbl, and found that S limitation propagated into methionine metabolism, synthesis of FeS clusters, and oxidative stress. In addition, we identified several open reading frames likely to respond specifically to S availability. As predicted from the fact that the ddp operon is activated by NtrC, synthesis of cross-links between diaminopimelate residues in the murein layer was increased under N-limiting conditions, as was the proportion of tripeptides. Both of these effects may allow increased scavenging of N from the dipeptide D-alanine-D-alanine, the substrate of the Ddp system.

Cluster Analysis↗

Chromosome evolution in the Thermotogales: large-scale inversions and strain diversification of CRISPR sequences.

In the present study, the chromosomes of two members of the Thermotogales were compared. A whole-genome alignment of Thermotoga maritima MSB8 and Thermotoga neapolitana NS-E has revealed numerous large-scale DNA rearrangements, most of which are associated with CRISPR DNA repeats and/or tRNA genes. These DNA rearrangements do not include the putative origin of DNA replication but move within the same replichore, i.e., the same replicating half of the chromosome (delimited by the replication origin and terminus). Based on cumulative GC skew analysis, both the T. maritima and T. neapolitana lineages contain one or two major inverted DNA segments. Also, based on PCR amplification and sequence analysis of the DNA joints that are associated with the major rearrangements, the overall chromosome architecture was found to be conserved at most DNA joints for other strains of T. neapolitana. Taken together, the results from this analysis suggest that the observed chromosomal rearrangements in the Thermotogales likely occurred by successive inversions after their divergence from a common ancestor and before strain diversification. Finally, sequence analysis shows that size polymorphisms in the DNA joints associated with CRISPRs can be explained by expansion and possibly contraction of the DNA repeat and spacer unit, providing a tool for discerning the relatedness of strains from different geographic locations.

Base Sequence↗

A multispecies comparison of the metazoan 3'-processing downstream elements and the CstF-64 RNA recognition motif.

BACKGROUND: The Cleavage Stimulation Factor (CstF) is a required protein complex for eukaryotic mRNA 3'-processing. CstF interacts with 3'-processing downstream elements (DSEs) through its 64-kDa subunit, CstF-64; however, the exact nature of this interaction has remained unclear. We used EST-to-genome alignments to identify and extract large sets of putative 3'-processing sites for mRNA from ten metazoan species, including Homo sapiens, Canis familiaris, Rattus norvegicus, Mus musculus, Gallus gallus, Danio rerio, Takifugu rubripes, Drosophila melanogaster, Anopheles gambiae, and Caenorhabditis elegans. In order to further delineate the details of the mRNA-protein interaction, we obtained and multiply aligned CstF-64 protein sequences from the same species. RESULTS: We characterized the sequence content and specific positioning of putative DSEs across the range of organisms studied. Our analysis characterized the downstream element (DSE) as two distinct parts - a proximal UG-rich element and a distal U-rich element. We find that while the U-rich element is largely conserved in all of the organisms studied, the UG-rich element is not. Multiple alignment of the CstF-64 RNA recognition motif revealed that, while it is highly conserved throughout metazoans, we can identify amino acid changes that correlate with observed variation in the sequence content and positioning of the DSEs. CONCLUSION: Our analysis confirms the early reports of separate U- and UG-rich DSEs. The correlated variations in protein sequence and mRNA binding sequences provide novel insights into the interactions between the precursor mRNA and the 3'-processing machinery.

Amino Acid Motifs↗

Variation in alternative splicing across human tissues.

BACKGROUND: Alternative pre-mRNA splicing (AS) is widely used by higher eukaryotes to generate different protein isoforms in specific cell or tissue types. To compare AS events across human tissues, we analyzed the splicing patterns of genomically aligned expressed sequence tags (ESTs) derived from libraries of cDNAs from different tissues. RESULTS: Controlling for differences in EST coverage among tissues, we found that the brain and testis had the highest levels of exon skipping. The most pronounced differences between tissues were seen for the frequencies of alternative 3' splice site and alternative 5' splice site usage, which were about 50 to 100% higher in the liver than in any other human tissue studied. Quantifying differences in splice junction usage, the brain, pancreas, liver and the peripheral nervous system had the most distinctive patterns of AS. Analysis of available microarray expression data showed that the liver had the most divergent pattern of expression of serine-arginine protein and heterogeneous ribonucleoprotein genes compared to the other human tissues studied, possibly contributing to the unusually high frequency of alternative splice site usage seen in liver. Sequence motifs enriched in alternative exons in genes expressed in the brain, testis and liver suggest specific splicing factors that may be important in AS regulation in these tissues. CONCLUSIONS: This study distinguishes the human brain, testis and liver as having unusually high levels of AS, highlights differences in the types of AS occurring commonly in different tissues, and identifies candidate cis-regulatory elements and trans-acting factors likely to have important roles in tissue-specific AS in human cells.

Alternative Splicing↗

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC↗