Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Personalized medicine strategy for MPNSTs: using precision oncology on PDOX models to inform tumor boards.

BACKGROUND: Malignant peripheral nerve sheath tumors (MPNSTs) are a heterogeneous group of aggressive soft tissue sarcomas with poor prognosis. Currently there is a lack of effective treatments for MPNSTs. Here, we propose a personalized medicine approach that integrates a precision oncology strategy guided by MPNST genomic analysis, with a functional validation of treatment response in an orthotopic xenograft model (PDOX) derived from the same MPNST. METHODS: Comprehensive whole genome sequencing analysis was performed in primary MPNSTs, relapses and (in one case) metastases, following disease progression in two independent individuals. Matched MPNST PDOX models were generated by orthotopically implanting tumor fragments near the sciatic nerve of immunodeficient mice. Candidate targeted combination therapies were prioritized based on genomic alterations and tested in vivo in the PDOX models. RESULTS: The feasibility of the developed strategy is illustrated for two MPNST patients, one Neurofibromatosis type 1 (NF1) individual that developed two independent MPNSTs and another sporadic MPNST case with multiple metastatic relapses. Genomic analysis revealed a remarkable degree of genomic stability across primary MPNSTs and their successive relapses in each patient, and even metastases in one individual. While based on a small number of cases requiring additional analyses, this finding aligns with previous evidence suggesting a fair genomic conservation throughout tumor evolution. This stability supports the identification of consistent therapeutic vulnerabilities throughout disease progression. Among the therapies tested, co-treatment of MEK inhibitor (MEKi) plus bromodomain inhibitor (BETi) elicited the highest antitumor activity, resulting in approximately 60% tumor volume reduction in the sporadic MPNST PDX model, whose patient has been receiving this therapy for eight months with sustained remission. CONCLUSIONS: This study demonstrates the feasibility and clinical utility of integrating genomic-driven precision oncology with PDOX-based functional testing for MPNSTs. This strategy may support molecular tumor boards (MTBs) in their treatment decisions. The observed genomic stability supports the use of longitudinal tumor profiling to guide treatment, and the success of MEKi+BETi highlights its potential as a combination therapy for MPNSTs.

Precision Medicine↗

Towards rice genome scanning by map-based AFLP fingerprinting.

Map-based DNA fingerprinting with AFLP markers provides a fast method for scanning the rice genome. Three hundred AFLP markers identified with ten primer combinations were mapped in two rice populations. The genetic maps were aligned and almost full coverage of the rice genome was obtained. The transferability of AFLP markers between indica x japonica and indica x indica crosses was tested. The chromosomes were divided into DNA Fingerprint Linkage Blocks (DFLBs) defined by specific AFLP markers. Using these blocks, the degree of similarity or divergence within specific chromosome regions was calculated for nine varieties. Applications of map-based fingerprinting for biodiversity studies and maker-assisted selection are discussed.

Chromosome Mapping↗

Orthologs of the vaccinia A13L and A36R virion membrane protein genes display diversity in species of the genus Orthopoxvirus.

Alignment of vaccinia and variola virus genomes has highlighted some targets that display diversity. We have investigated the sequence diversity of two viral membrane protein genes from 36 different orthopoxvirus (OPV) strains to evaluate the suitability of these loci to differentiate between OPV species. Orthologs of the vaccinia virus Copenhagen A13L gene were all predicted to have functional genes that ranged between 201-213 bps in length. Whereas the N- and C-termini of each protein were relatively well conserved within the genus, a central proline-rich domain displayed characteristic species-specific amino acid motifs. Orthologs of the A36R gene displayed considerable sequence variation between species and strains. The majority of variation was localised to the last 100 bps of the gene. Multiple-alignment of these sequences identified the presence of gaps, insertions or frame-shift mutations among the samples examined. Nearly all strains of cowpox virus contained different nucleotide sequences at this locus. Phylogenetic analysis of the aligned sequences showed that variola and camelpox viruses shared a common ancestry with cowpox virus, whereas ectromelia viruses were divergent from all the other OPVs examined. Phylogeny generated with A13L sequences distributed the OPV species in a manner that correlated to their known properties.

Acinonyx↗

In vitro assembly properties of purified bacterially expressed capsid proteins of human immunodeficiency virus.

The Gag polyprotein of retroviruses is sufficient for assembly and budding of virus-like particles from the host cell. In the case of human immunodeficiency virus (HIV), Gag contains the domains matrix, capsid (CA), nucleocapsid (NC) and p6 which are separated by the viral proteinase inside the nascent virion, leading to morphological maturation to yield an infectious virus. In the mature virus, CA forms a capsid shell surrounding the ribonucleoprotein core consisting of NC and the genomic RNA. To define requirements for particle assembly and functional contributions of individual domains, we expressed domains of HIV Gag in Escherichia coli and purified the products to near homogeneity. In vitro assembly of CA, with or without the C-terminally adjacent spacer peptide, yielded tubular structures with a diameter of approximately 55 nm and heterogeneous length. Efficient particle formation required high protein concentration, high salt and neutral to alkaline pH. In contrast, in vitro assembly of CA-NC occurred at a 20-fold lower protein concentration and in low salt, but required addition of RNA. These results suggest that hydrophobic interactions of capsid proteins are sufficient for particle formation while the RNA-binding nucleocapsid domain may concentrate and align structural proteins on the viral genome.

Base Sequence↗

Complexity: an internet resource for analysis of DNA sequence complexity.

The search for DNA regions with low complexity is one of the pivotal tasks of modern structural analysis of complete genomes. The low complexity may be preconditioned by strong inequality in nucleotide content (biased composition), by tandem or dispersed repeats or by palindrome-hairpin structures, as well as by a combination of all these factors. Several numerical measures of textual complexity, including combinatorial and linguistic ones, together with complexity estimation using a modified Lempel-Ziv algorithm, have been implemented in a software tool called 'Complexity' (http://wwwmgs.bionet.nsc.ru/mgs/programs/low_complexity/). The software enables a user to search for low-complexity regions in long sequences, e.g. complete bacterial genomes or eukaryotic chromosomes. In addition, it estimates the complexity of groups of aligned sequences.

Algorithms↗

Universal trees based on large combined protein sequence data sets.

Universal trees of life based on small-subunit (SSU) ribosomal RNA (rRNA) support the separate mono/holophyly of the domains Archaea (archaebacteria), Bacteria (eubacteria) and Eucarya (eukaryotes) and the placement of extreme thermophiles at the base of the Bacteria. The concept of universal tree reconstruction recently has been upset by protein trees that show intermixing of species from different domains. Such tree topologies have been attributed to either extensive horizontal gene transfer or degradation of phylogenetic signals because of saturation for amino acid substitutions. Here we use large combined alignments of 23 orthologous proteins conserved across 45 species from all domains to construct highly robust universal trees. Although individual protein trees are variable in their support of domain integrity, trees based on combined protein data sets strongly support separate monophyletic domains. Within the Bacteria, we placed spirochaetes as the earliest derived bacterial group. However, elimination from the combined protein alignment of nine protein data sets, which were likely candidates for horizontal gene transfer, resulted in trees showing thermophiles as the earliest evolved bacterial lineage. Thus, combined protein universal trees are highly congruent with SSU rRNA trees in their strong support for the separate monophyly of domains as well as the early evolution of thermophilic Bacteria.

Amino Acid Sequence↗

ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.

Information about the genomic coordinates and the sequence of experimentally identified transcription factor binding sites is found scattered under a variety of diverse formats. The availability of standard collections of such high-quality data is important to design, evaluate and improve novel computational approaches to identify binding motifs on promoter sequences from related genes. ABS (http://genome.imim.es/datasets/abs2005/index.html) is a public database of known binding sites identified in promoters of orthologous vertebrate genes that have been manually curated from bibliography. We have annotated 650 experimental binding sites from 68 transcription factors and 100 orthologous target genes in human, mouse, rat or chicken genome sequences. Computational predictions and promoter alignment information are also provided for each entry. A simple and easy-to-use web interface facilitates data retrieval allowing different views of the information. In addition, the release 1.0 of ABS includes a customizable generator of artificial datasets based on the known sites contained in the collection and an evaluation tool to aid during the training and the assessment of motif-finding programs.

Animals↗

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses, often with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors…. These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac . We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC . Alevin-fry-atac is highly scalable and, when using 32 threads, is approximately 2.8 times faster than Chromap (the second fastest approach) while using approximately one third of the memory and mapping slightly more reads. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual colorenhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry ) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger . Furthermore, our modified pseudoalignment approach should be easily applicable and extendable to other genome-centric mapping-based tasks and modalities such as standard DNA-seq, DNase-seq, Chip-seq and Hi-C.

Journal Article↗

Alignment of a 1.2-Mb chromosomal region from three strains of Rhodobacter capsulatus reveals a significantly mosaic structure.

High-resolution physical maps of the genomes of three Rhodobacter capsulatus strains, derived from ordered cosmid libraries, were aligned. The 1.2-Mb segment of the SB1003 genome studied here is adjacent to a 1-Mb region analyzed previously [Fonstein, M., Nikolskaya, T. & Haselkorn, H. (1995) J. Bacteriol. 177, 2368-2372]. Probes derived from the ordered cosmid set of R. capsulatus SB1003 were used to link cosmids from the St. Louis and 2.3.1 strain libraries. Cosmids selected this way did not merge into a single contig but formed several unlinked groups. EcoRV restriction maps of the ordered cosmids were then constructed using lambda terminase and fused to derive fragments of the chromosomal map. In order to link these fragments, their ends were transcribed to produce secondary probes for hybridization to gridded cosmid libraries of the same strains. This linking reduced the number of subcontigs to three for the St. Louis strain and one for the 2.3.1 strain. Hybridization of the same probes back to the ordered cosmid set of SB1003 positioned the subcontigs on the high-resolution physical map of SB1003. The final alignment of the restriction maps shows numerous large and small translocations in this 1.2-Mb chromosomal region of the three Rhodobacter strains. In addition, the chromosomes of the three strains, whose fine-structure maps can now be compared over 2.2 Mb, are seen to contain regions of 15-80 kb in which restriction sites are highly polymorphic, interspersed among regions in which the positions of restriction sites are highly conserved.

Chromosomes, Bacterial↗

Lineage-specific tandem repeats riding on a transposable element of MITE in Xenopus evolution: a new mechanism for creating simple sequence repeats.

Xstir is a repetitive DNA sequence element that is extremely amplified as a common component of two different structures: a tandem repeat (Xstir array) and a MITE (miniature inverted-repeat transposable element) in the genome of Xenopus laevis. To elucidate the origin and evolutionary history of Xstir-related sequences, we investigated their species specificity among three Xenopus species (X. laevis, X. borealis, and X. tropicalis). Analyses by sequence alignment and digestion with restriction enzymes of genomic Xstir-related sequences revealed that the MITE (Xmix MITE) was well conserved among the three Xenopus species, with small lineage-specific differences. On the other hand, the tandem repeat element (tropXstir) in X. tropicalis was different from the Xstir that X. laevis and X. borealis have in common. Both sequences of Xstir and tropXstir were, however, different segments of the Xmix MITE. The results suggest that these tandem repeats were formed by partial tandem duplication of the MITE internal sequence in each lineage of X. tropicalis and of X. borealis/X. laevis after their branching. A molecular mechanism for creating and elongating the tandem repeats from the MITE is proposed.

Animals↗

Genomic organization and expression of the human mono-ADP-ribosyltransferase ART3 gene.

Here we describe an RT-PCR analysis of mono-ADP-ribosyltransferase 3 (ART3) mRNA expression in macrophages, testis, semen, tonsil, heart and skeletal muscle and the complete gene structure as obtained by sequence alignment of PCR products with a human genomic clone (GenBank accession no. AC112719). Twelve exons (ex1-12) were found to make up the coding region of the gene (one more than previously published). Two prominent classes of ART3 splice variants could be distinguished by the presence or absence of ex2 which encodes most of ART3 protein. Among the ex2-containing mRNA species, the most frequently amplified variant did not include exons 9 to 11, except in skeletal muscle, in which the major splice variant lacked ex10 only. Two different, previously not reported 5' non-translated regions (5' UTRs) were identified, demonstrating the presence of two alternative promoters that we termed palpha and pbeta. Whereas the 5'UTR originating from palpha, was split up into three exons, a single exon represented the 5' UTR of pbeta transcripts. Strikingly, in heart, skeletal muscle and tonsils the upstream promoter palpha was totally inactive and ART3 transcription appears to be driven solely by pbeta. In all other cell types tested, transcription started mainly (if not exclusively) at palpha. Thus, ART3 expression in human cells appears to be governed by a combination of differential splicing and tissue-preferential use of two alternative promoters. This specific use is evolutionary conserved as shown by analysis of the 5' UTR of the mouse ART3 mRNA.

5' Untranslated Regions↗

DiffTool: building, visualizing and querying protein clusters.

UNLABELLED: DiffTool is a resource to build and visualize protein clusters computed from a sequence database. The package provides a clustering tool to construct protein families according to sequence similarities and a web interface to query the corresponding clusters. A subtractive genome analysis tool selects protein families specific for a genome or a group of genomes. For each protein cluster, DiffTool includes access to sequences, coloured multiple alignments and phylogenetic trees. AVAILABILITY: A cluster database built from yeast and complete prokaryotic genomes is queryable at http://bioweb.pasteur.fr/seqanal/difftool. All the Perl sources are freely available to non-profit organizations upon request.

Cluster Analysis↗

Whole-genome shotgun optical mapping of Rhodospirillum rubrum.

Rhodospirillum rubrum is a phototrophic purple nonsulfur bacterium known for its unique and well-studied nitrogen fixation and carbon monoxide oxidation systems and as a source of hydrogen and biodegradable plastic production. To better understand this organism and to facilitate assembly of its sequence, three whole-genome restriction endonuclease maps (XbaI, NheI, and HindIII) of R. rubrum strain ATCC 11170 were created by optical mapping. Optical mapping is a system for creating whole-genome ordered restriction endonuclease maps from randomly sheared genomic DNA molecules extracted from cells. During the sequence finishing process, all three optical maps confirmed a putative error in sequence assembly, while the HindIII map acted as a scaffold for high-resolution alignment with sequence contigs spanning the whole genome. In addition to highlighting optical mapping's role in the assembly and confirmation of genome sequence, this work underscores the unique niche in resolution occupied by the optical mapping system. With a resolution ranging from 6.5 kb (previously published) to 45 kb (reported here), optical mapping advances a "molecular cytogenetics" approach to solving problems in genomic analysis.

Contig Mapping↗

Immunological identification of rat tissue kallikrein cDNA and characterization of the kallikrein gene family.

A tissue kallikrein cDNA was identified by direct immunological screening with affinity-purified anti-rat tissue kallikrein antibody from a rat submandibular cDNA library constructed with the expression vector pUC8. Sequence analysis of the kallikrein cDNA revealed an encoded protein 97% homologous to the partial amino acid sequence of rat submandibular kallikrein. This cDNA was used to hybrid-select kallikrein-specific RNA from submandibular gland. Translation of the hybrid-selected RNA in a cell-free assay system resulted in the production of a 37 kDa peptide representing the preproenzyme. In addition, hybrid-selection of RNA under less stringent conditions showed cross-hybridization with other submandibular gland mRNA species. In correlation with these results, analysis of rat genomic DNA showed extensive hybridization, suggesting a family of closely related kallikrein-like genes. Consequently, a Charon 4A rat genomic library was screened for kallikrein genes by hybridization with rat tissue kallikrein cDNA. Thirty-four clones were isolated and found to be highly homologous by hybridization and restriction enzymes analyses. Fourteen unique clones were identified by restriction enzyme site polymorphisms within DNA segments which hybridized to the kallikrein cDNA probe and it was estimated that at least 17 different kallikrein-like genes are present in the rat. Sequence and structural analysis of one of the genomic clones revealed a gene structure similar to that of other serine proteinases. Comparison of the partially sequenced exon regions of the gene with the sequence of rat tissue kallikrein cDNA reveals 89% identity when aligned for the greatest homology. However, the genomic sequence predicts termination codons in all three translational reading frames, implying that this gene is nonfunctional, i.e., a pseudogene. Comparison of the rat genomic sequence to a kallikrein-like gene from the mouse reveals extensive preservation of exons, less identity within introns and no significant homology between extragenic regions.

Animals↗

Partial structure of the phylloxin gene from the giant monkey frog, Phyllomedusa bicolor: parallel cloning of precursor cDNA and genomic DNA from lyophilized skin secretion.

Phylloxin is a novel prototype antimicrobial peptide from the skin of Phyllomedusa bicolor. Here, we describe parallel identification and sequencing of phylloxin precursor transcript (mRNA) and partial gene structure (genomic DNA) from the same sample of lyophilized skin secretion using our recently-described cloning technique. The open-reading frame of the phylloxin precursor was identical in nucleotide sequence to that previously reported and alignment with the nucleotide sequence derived from genomic DNA indicated the presence of a 175 bp intron located in a near identical position to that found in the dermaseptins. The highly-conserved structural organization of skin secretion peptide genes in P. bicolor can thus be extended to include that encoding phylloxin (plx). These data further reinforce our assertion that application of the described methodology can provide robust genomic/transcriptomic/peptidomic data without the need for specimen sacrifice.

Amphibian Proteins↗

Kalign--an accurate and fast multiple sequence alignment algorithm.

BACKGROUND: The alignment of multiple protein sequences is a fundamental step in the analysis of biological data. It has traditionally been applied to analyzing protein families for conserved motifs, phylogeny, structural properties, and to improve sensitivity in homology searching. The availability of complete genome sequences has increased the demands on multiple sequence alignment (MSA) programs. Current MSA methods suffer from being either too inaccurate or too computationally expensive to be applied effectively in large-scale comparative genomics. RESULTS: We developed Kalign, a method employing the Wu-Manber string-matching algorithm, to improve both the accuracy and speed of multiple sequence alignment. We compared the speed and accuracy of Kalign to other popular methods using Balibase, Prefab, and a new large test set. Kalign was as accurate as the best other methods on small alignments, but significantly more accurate when aligning large and distantly related sets of sequences. In our comparisons, Kalign was about 10 times faster than ClustalW and, depending on the alignment size, up to 50 times faster than popular iterative methods. CONCLUSION: Kalign is a fast and robust alignment method. It is especially well suited for the increasingly important task of aligning large numbers of sequences.

Algorithms↗

A shotgun optical map of the entire Plasmodium falciparum genome.

The unicellular parasite Plasmodium falciparum is the cause of human malaria, resulting in 1.7-2.5 million deaths each year. To develop new means to treat or prevent malaria, the Malaria Genome Consortium was formed to sequence and annotate the entire 24.6-Mb genome. The plan, already underway, is to sequence libraries created from chromosomal DNA separated by pulsed-field gel electrophoresis (PFGE). The AT-rich genome of P. falciparum presents problems in terms of reliable library construction and the relative paucity of dense physical markers or extensive genetic resources. To deal with these problems, we reasoned that a high-resolution, ordered restriction map covering the entire genome could serve as a scaffold for the alignment and verification of sequence contigs developed by members of the consortium. Thus optical mapping was advanced to use simply extracted, unfractionated genomic DNA as its principal substrate. Ordered restriction maps (BamHI and NheI) derived from single molecules were assembled into 14 deep contigs corresponding to the molecular karyotype determined by PFGE (ref. 3).

Animals↗

In silico identification of new secretory peptide genes in Drosophila melanogaster.

Bioactive peptides play critical roles in regulating most biological processes in animals. The elucidation of the amino acid sequence of these regulatory peptides is crucial for our understanding of animal physiology. Most of the (neuro)peptides currently known were identified by purification and subsequent amino acid sequencing. With the entire genome sequence of some animals now available, it has become possible to predict novel putative peptides. In this way, BLAST (Basic Local Alignment Searching Tool) analysis of the Drosophila melanogaster genome has allowed annotation of 36 secretory peptide genes so far. Peptide precursor genes are, however, poorly predicted by this algorithm, thus prompting an alternative approach described here. With the described searching program we scanned the Drosophila genome for predicted proteins with the structural hallmarks of neuropeptide precursors. As a result, 76 additional putative secretory peptide genes were predicted in addition to the 43 annotated ones. These putative (neuro)peptide genes contain conserved motifs reminiscent of known neuropeptides from other animal species. Peptides that display sequence similarities to the mammalian vasopressin, atrial natriuretic peptide, and prolactin precursors and the invertebrate peptides orcokinin, prothoracicotropic hormones, trypsin modulating oostatic factor, and Drosophila immune induced peptides (DIMs) among others were discovered. Our data hence provide further evidence that many neuropeptide genes were already present in the ancestor of Protostomia and Deuterostomia prior to their divergence. This bioinformatic study opens perspectives for the genome-wide analysis of peptide genes in other eukaryotic model organisms.

Algorithms↗