Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

A set of conserved PCR primers for the analysis of simple sequence repeat polymorphisms in chloroplast genomes of dicotyledonous angiosperms.

Short runs of mononucleotide repeats are present in chloroplast genomes of higher plants. In soybean, rice, and pine, PCR (polymerase chain reaction) with flanking primers has shown that the numbers of A or T residues in such repeats are variable among closely related taxa. Here we describe a set of primers for studying mononucleotide repeat variation in chloroplast DNA of angiosperms where database information is limited. A total of 39 (A)n and (T)n repeats (n > or = 10) were identified in the tobacco chloroplast genome, and DNA sequences encompassing these 39 regions were aligned with orthologous DNA sequences in the databases. Consensus primer pairs were constructed and used to amplify total genomic DNA from a hierarchical set of angiosperms. All 10 primer pairs generated PCR products from members of the Solanaceae, and 8 of the 10 were also functional in most other angiosperm species. Levels of interspecific polymorphism within the genera Nicotiana, Lycopersicon (both Solanaceae), and Actinidia (Actinidiaceae) proved to be high, while intraspecific variation in Nicotiana tabacum, Lycopersicon esculentum, and Actinidia chinensis was limited. Sequence analysis of PCR products from three primer pairs revealed variable numbers of A, G, and T residues in mononucleotide arrays as the major cause of polymorphism in Actinidia. Our results suggest that universal primers targeted to mononucleotide repeats may serve as general tools to study chloroplast variation in angiosperms.

Base Sequence↗

Comparative genomics of the Archaea (Euryarchaeota): evolution of conserved protein families, the stable core, and the variable shell.

Comparative analysis of the protein sequences encoded in the four euryarchaeal species whose genomes have been sequenced completely (Methanococcus jannaschii, Methanobacterium thermoautotrophicum, Archaeoglobus fulgidus, and Pyrococcus horikoshii) revealed 1326 orthologous sets, of which 543 are represented in all four species. The proteins that belong to these conserved euryarchaeal families comprise 31%-35% of the gene complement and may be considered the evolutionarily stable core of the archaeal genomes. The core gene set includes the great majority of genes coding for proteins involved in genome replication and expression, but only a relatively small subset of metabolic functions. For many gene families that are conserved in all euryarchaea, previously undetected orthologs in bacteria and eukaryotes were identified. A number of euryarchaeal synapomorphies (unique shared characters) were identified; these are protein families that possess sequence signatures or domain architectures that are conserved in all euryarchaea but are not found in bacteria or eukaryotes. In addition, euryarchaea-specific expansions of several protein and domain families were detected. In terms of their apparent phylogenetic affinities, the archaeal protein families split into bacterial and eukaryotic families. The majority of the proteins that have only eukaryotic orthologs or show the greatest similarity to their eukaryotic counterparts belong to the core set. The families of euryarchaeal genes that are conserved in only two or three species constitute a relatively mobile component of the genomes whose evolution should have involved multiple events of lineage-specific gene loss and horizontal gene transfer. Frequently these proteins have detectable orthologs only in bacteria or show the greatest similarity to the bacterial homologs, which might suggest a significant role of horizontal gene transfer from bacteria in the evolution of the euryarchaeota.

Amino Acid Sequence↗

Bioinformatics issues for automating the annotation of genomic sequences.

The rapid explosion in the amount of biological data being generated worldwide is surpassing efforts to manage analysis of the data. As part of an ongoing project to automate and manage bioinformatics analysis, the authors have designed and implemented a simple automated annotation system, which is described in this paper. The system is applied to existing GenBank/DDBJ/EMBL entries and compared with existing annotations to illustrate not only potential errors but also that they are generally not up-to-date, as a result of new versions of analysis tools and updates of genomic repositories. We highlight the important Bioinformatics issues of storage and management of information to ensure data and results are kept up-to-date in light of new information becoming available. Surprisingly, from just four database entries, a significant number of new features were found. We describe the results as well as identify important issues that need to be addressed in order to automate the re-analysis/re-annotation of genomic sequences within a reasonable timeframe.

Computational Biology↗

Mining the mammalian genome for artiodactyl systematics.

A total of 7,806 nucleotide positions derived from one mitochondrial and eight nuclear DNA segments were used to provide a robust phylogeny for members of the order Artiodactyla. Twenty-four artiodactyl and two cetacean species were included, and the horse (order Perissodactyla) was used as the outgroup. Limited rate heterogeneity was observed among the nuclear genes. The partition homogeneity tests indicated no conflicting signal among the nuclear genes fragments, so the sequence data were analyzed together and as separate loci. Analyses based on the individual nuclear DNA fragments and on 34 unique indels all produced phylogenies largely congruent with the topology from the combined data set. In sharp contrast to the nuclear DNA data, the mtDNA cytochrome b sequence data showed high levels of homoplasy, failed to produce a robust phylogeny, and were remarkably sensitive to taxon sampling. The nuclear DNA data clearly support the paraphyletic nature of the Artiodactyla. Additionally, the family Suidae is diphyletic, and the nonruminating pigs and peccaries (Suiformes) were the most basal cetartiodactyl group. The morphologically derived Ruminantia was always monophyletic; within this group, all taxa with paired bony structures on their skulls clustered together. The nuclear DNA data suggest that the Antilocaprinae account for a unique evolutionary lineage, the Cervidae and Bovidae are sister taxa, and the Giraffidae are more primitive.

Animals↗

[Is alternative splicing of mammalian genes conservative?].

The conservation of alternative splicing in orthologous genes from the human and mouse genomes was analyzed. Alternatively spliced mouse genes from the AsMamDB database were used to scan the draft human genome. The mouse protein isoforms were aligned with respect to orthologous human genes, and thus the exon-intron structure of the latter was established. Proteins isoforms that could not be aligned throughout their length were analyzed in detail using the human EST alignment.

Alternative Splicing↗

Insulin-like growth factor I differentially regulates the expression of HIRF1/hCAF1 and BTG1 genes in human MCF-7 breast cancer cells.

Differential display PCR analysis (DD-PCR) was used to identify novel genes that respond to IGF-I treatment in human MCF-7 breast cancer cells. Fifty-three cDNAs showed alterations in their mRNA levels in IGF-I treated cells. One of these genes showed a significant increase in the mRNA level in IGF-I treated cells in comparison to non-treated cells. We named this gene HIRF1 (human IGF-I regulated factor 1). Nucleotide blast analysis revealed that this gene has a 100% sequence identity with the sequence for BTG1 (B-cell translocation gene) binding factor 1 (human CCR4-associated factor 1 gene, hCAF1). By alignment of cloned HIRF1 cDNA and genomic DNA 8p21.3-p22 sequence, we were able to determine the exon-intron structure of the cloned HIRF1 gene on chromosome 8. Northern blot and real-time PCR analysis showed that BTG1 and c-fos reached their maximal expression fairly early within 10 min to 1 h, and decreased to basal levels after 3 h of IGF-I treatment. HIRF1/hCAF1 expression reached maximal stimulation after 3 h of IGF-I treatment and then gradually decreased to basal level. HIRF1 and BTG1 mRNA was inhibited by inhibitors of the cell signaling pathways, PI3/Akt kinase and MAPK kinases (ERK1/2 and p38). In summary, cloned HIRF1/hCAF1 is coregulated with BTG1 in response to IGF-I. The regulation of these genes as early response genes may have an important role in differentiation, growth and proliferation of breast cancer cells.

Amino Acid Sequence↗

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article↗

Characterization and targeting of the murine alpha2-antiplasmin gene.

Alpha2-Antiplasmin (alpha2-AP) is the main physiological plasmin inhibitor in mammalian plasma. As a first step toward the generation of alpha2-AP deficient mice, the murine alpha2-AP gene was characterized and a targeting vector for homologous recombination in embryonic stem (ES) cells constructed. Alignment of nucleotide sequences obtained from genomic subclones allowed location of exons 2 through 10 of the alpha2-AP gene, but failed to identify the 5' boundary of exon 1. Compared to the human gene, exons 2 through 9 in the murine gene have identical size and intron-exon boundaries obeying the GT/AG rule. The 5' boundary of exon 10 is identical in both genes while the 3' non-coding region is 64 bp longer in the human gene. Introns 2, 3, 6 and 8 have similar sizes in the mouse and human genes; intron 1 is 6-fold smaller, introns 5, 7 and 9 are 2- to 3-fold smaller, whereas intron 4 is about 2-fold larger in the mouse gene. Compared to the human 5' flanking sequence, an insertion of a simple repeat region with sequence (TGG)n has occurred. The open reading frame of the mouse alpha2-AP gene encodes a 491-amino-acid protein comprising the experimentally determined NH2-terminus of the mature protein Val-Asp-Leu-Pro-Gly-. A targeting vector, pPNT.alpha2-AP, was constructed by introducing a homologous sequence of 8.3 kb in total in the parental pPNT vector. In pPNT.alpha2-AP, the neomycin resistance expression cassette replaces a 7 kb genomic fragment comprising exon 2 through part of exon 10 (including the stop codon), which represents the entire sequence encoding the mature protein, including the fibrin-binding domain, the reactive site peptide bond and the plasmin(ogen)-binding region. Electroporation of 129R1 embryonic stem (ES) cells with the linearized vector pPNT.alpha2-AP yielded three targeted clones with correct homologous recombination at the 5'- and 3'-ends, as confirmed by Southern blot analysis of purified genomic DNA with appropriate restriction enzymes and probes. These targeted clones will be used to generate alpha2-AP deficient mice.

Amino Acid Sequence↗

Functional annotation of putative aminoglycoside antibiotic modifying proteins in Mycobacterium tuberculosis H37Rv.

The growing availability of sequences of bacterial genomes has revealed a number of open reading frames predicted by sequence alignment to encode antibiotic resistance proteins. The presence of these putative resistance genes within bacterial genomes raises important questions regarding potential reservoirs of resistance elements and their evolution. Here we examine four gene products encoding predicted aminoglycoside-aminocyclitol antibiotic modifying enzymes, two phosphotransferases and two acetyltransferases, derived from analysis of the genome sequence of Mycobacterium tuberculosis strain H37Rv with the goal of assigning biochemical function by purification of each protein and characterization of their ability to modify aminoglycoside antibiotics. Only one of these enzymes, the previously characterized aminoglycoside acetyltransferase AAC(2')-Ic, displayed compelling aminoglycoside modifying activity. While the putative phosphotransferase encoded by the Rv3225c gene did display low levels of aminoglycoside kinase activity, the predicted kinase encoded by the Rv3817 gene lacked any such activity. A potential aminoglycoside 6'-acetyltransferase, encoded by the Rv1347c gene, did not show antibiotic acylation activity but did demonstrate selective thioesterase activity with numerous acyl-CoAs. This activity, together with the genomic environment of the Rv1347c gene in a likely polyketide synthesis cluster, suggests a role for this protein in secondary metabolism and not in antibiotic modification. It was thus shown that only one of four putative aminoglycosides modifying enzymes derived from the whole genome sequencing of M. tuberculosis H37Rv showed sufficient predicted enzyme activity to be annotated as an aminoglycoside resistance element. This study demonstrates the necessity of biochemical annotation methods as a follow up to in silico sequence alignment-based methods of assigning gene product function.

Acetyltransferases↗

A new contribution to the integration of human and porcine genome maps: 623 new points of homology.

In this study we examined homologies between 1,735 porcine microsatellites and human sequence. For 1,710 microsatellites we directly used the sequence flanking the repeat available in GenBank. For a set of 305 microsatellites, a BAC library was screened and end-sequencing provided 461 additional sequences. Altogether 2,171 porcine sequences were tentatively aligned with the sequence of the human genome using the fasta program. Human homologies were observed for 652 microsatellite loci and porcine chromosome assignments available for 623 microsatellites provide useful links in the human and pig comparative map. Moreover for 92 STS, a significant sequence similarity was detected using at least two sequences and in all cases corresponding human locations were consistent. The present study allowed the integration of anonymous markers and the porcine linkage map into the framework of the comparative data between human and porcine genomes (http://w3.toulouse.inra.fr/lgc/pig/msat/). Moreover all conserved syntenic segments were defined on human chromosomes.

Animals↗

The mosaic structure of variation in the laboratory mouse genome.

Most inbred laboratory mouse strains are known to have originated from a mixed but limited founder population in a few laboratories. However, the effect of this breeding history on patterns of genetic variation among these strains and the implications for their use are not well understood. Here we present an analysis of the fine structure of variation in the mouse genome, using single nucleotide polymorphisms (SNPs). When the recently assembled genome sequence from the C57BL/6J strain is aligned with sample sequence from other strains, we observe long segments of either extremely high (approximately 40 SNPs per 10 kb) or extremely low (approximately 0.5 SNPs per 10 kb) polymorphism rates. In all strain-to-strain comparisons examined, only one-third of the genome falls into long regions (averaging >1 Mb) of a high SNP rate, consistent with estimated divergence rates between Mus musculus domesticus and either M. m. musculus or M. m. castaneus. These data suggest that the genomes of these inbred strains are mosaics with the vast majority of segments derived from domesticus and musculus sources. These observations have important implications for the design and interpretation of positional cloning experiments.

Albinism↗

Approximate matching of structured motifs in DNA sequences.

Several methods have been developed for identifying more or less complex RNA structures in a genome. All these methods are based on the search for conserved primary and secondary sub-structures. In this paper, we present a simple formal representation of a helix, which is a combination of sequence and folding constraints, as a constrained regular expression. This representation allows us to develop a well-founded algorithm that searches for all approximate matches of a helix in a genome. The algorithm is based on an alignment graph constructed from several copies of a pushdown automaton, arranged one on top of another. This is a first attempt to take advantage of the possibilities of pushdown automata in the context of approximate matching. The worst time complexity is O(krpn), where k is the error threshold, n the size of the genome, p the size of the secondary expression, and r its number of union symbols. We then extend the algorithm to search for pseudo-knots and secondary structures containing an arbitrary number of helices.

Algorithms↗

Profile-based detection of microRNA precursors in animal genomes.

MOTIVATION: MicroRNAs (miRNA) are essential 21-22 nt regulatory RNAs produced from larger hairpin-like precursors. Local sequence alignment tools such as BLAST are able to identify new members of known miRNA families, but not all of them. We set out to estimate how many new miRNAs could be recovered using a profile-based strategy such as that implemented in the ERPIN program. RESULTS: We constructed alignments for 18 miRNA families and performed ERPIN searches on animal genomes. Results were compared to those of a WU-BLAST search at the same E-value cutoff. The two combined approaches produced 265 new miRNA candidates that were not found in miRNA databases. About 17% of hits were ERPIN specific. They showed better structural characteristics than BLAST-specific hits and included interesting candidates such as members of the miR-17 cluster in Tetraodon. Profile-based RNA detection will be an important complement of similarity search programs in the completion of miRNA collections.

Algorithms↗

SLAM web server for comparative gene finding and alignment.

SLAM is a program that simultaneously aligns and annotates pairs of homologous sequences. The SLAM web server integrates SLAM with repeat masking tools and the AVID alignment program to allow for rapid alignment and gene prediction in user submitted sequences. Along with annotations and alignments for the submitted sequences, users obtain a list of predicted conserved non-coding sequences (and their associated alignments). The web site also links to whole genome annotations of the human, mouse and rat genomes produced with the SLAM program. The server can be accessed at http://bio.math.berkeley.edu/slam.

Algorithms↗

A comprehensive view on proteasomal sequences: implications for the evolution of the proteasome.

Proteasomes are large multimeric self-compartmentizing proteases, which play a crucial role in the clearance of misfolded proteins, breakdown of regulatory proteins, processing of proteins by specific partial proteolysis, cell cycle control as well as preparation of peptides for immune presentation. Two main types can be distinguished by their different tertiary structure: the 20S proteasome and the proteasome-like heat shock protein encoded by heat shock locus V, hslV. Usually, each biological kingdom is characterized by its specific type of proteasome. The 20S proteasomes occur in eukarya and archaea whereas hslV protease is prevalent in bacteria. To verify this rule we applied a genome-wide sequence search to identify proteasomal sequences in data of finished and yet unfinished genome projects. We found several exceptions to this paradigm: (1) Protista: in addition to the 20S proteasome, Leishmania, Trypanosoma and Plasmodium contained hslV, which may have been acquired from an alpha-proteobacterial progenitor of mitochondria. (2) Bacteria: for Magnetospirillum magnetotacticum and Enterococcus faecium we found that each contained two distinct hslVs due to gene duplication or horizontal transfer. Including unassembled data into the analyses we confirmed that a number of bacterial genomes do not contain any proteasomal sequence due to gene loss. (3) High G+C Gram-positives: we confirmed that high G+C Gram-positives possess 20S proteasomes rather than hslV proteases. The core of the 20S proteasome consists of two distinct main types of homologous monomers, alpha and beta, which differentiated into seven subtypes by further gene duplications. By looking at the genome of the intracellular pathogen Encephalitozoon cuniculi we were able to show that differentiation of beta-type subunits into different subtypes occurred earlier than that of alpha-subunits. Additionally, our search strategy had an important methodological consequence: a comprehensive sequence search for a particular protein should also include the raw sequence data when possible because proteins might be missed in the completed assembled genome. The structure-based multiple proteasomal alignment of 433 sequences from 143 organisms can be downloaded from the URL dagger and will be updated regularly.

Amino Acid Sequence↗