Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Combinatorial microRNA target predictions.

MicroRNAs are small noncoding RNAs that recognize and bind to partially complementary sites in the 3' untranslated regions of target genes in animals and, by unknown mechanisms, regulate protein production of the target transcript. Different combinations of microRNAs are expressed in different cell types and may coordinately regulate cell-specific target genes. Here, we present PicTar, a computational method for identifying common targets of microRNAs. Statistical tests using genome-wide alignments of eight vertebrate genomes, PicTar's ability to specifically recover published microRNA targets, and experimental validation of seven predicted targets suggest that PicTar has an excellent success rate in predicting targets for single microRNAs and for combinations of microRNAs. We find that vertebrate microRNAs target, on average, roughly 200 transcripts each. Furthermore, our results suggest widespread coordinate control executed by microRNAs. In particular, we experimentally validate common regulation of Mtpn by miR-375, miR-124 and let-7b and thus provide evidence for coordinate microRNA control in mammals.

Algorithms↗

[Study of functional L1 retrotransposon in human type 2 diabetes susceptibility loci].

OBJECTIVE: To investigate the susceptibility gene of type 2 diabetes mellitus (T2DM) through a novel strategy. METHODS: Firstly, the common feature of the putative susceptibility genes in the reported susceptibility loci was searched by using NCBI BLAST, and a functional L1 retrotransposon in the loci was found. Secondly, the mRNA expression level of the functional L1 retrotransposon in 25 Han T2DM patients and 22 normal controls was investigated by reverse transcription-polymerase chain reaction, and statistical analysis was implemented in statistical package SPSS10.0. Thirdly, L1 retrotransponson genome mutation screening was performed via sequencing. RESULTS: Screening the human genome for the retrotransposon genome via alignment with the L1 genome using NCBI BLAST showed the functional L1 retrotransposons distribute on most chromosomes except for chromosomes 19, 21 and Y on which rare type 2 diabetes susceptibility loci were reported to reside, and their distribution sites are consistent with the locations of the reported candidate type 2 diabetes susceptibility loci. The mRNA expression level of the functional L1 retrotransposon in the T2DM patients was significantly lower than that in normal subjects (P<0.001). Nonsense mutations including deletion and/or point mutations were observed in all of the 6 T2DM patients tested, but no mutation was observed in all of the 4 normal controls tested. CONCLUSION: The functional L1 retrotransposon may be a candidate susceptibility gene of type 2 diabetes or a key regulator of the susceptibility genes, and it may be an ideal candidate biomarker for screening type 2 diabetes.

Adult↗

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software↗

VISTA: computational tools for comparative genomics.

Comparison of DNA sequences from different species is a fundamental method for identifying functional elements in genomes. Here, we describe the VISTA family of tools created to assist biologists in carrying out this task. Our first VISTA server at http://www-gsd.lbl.gov/vista/ was launched in the summer of 2000 and was designed to align long genomic sequences and visualize these alignments with associated functional annotations. Currently the VISTA site includes multiple comparative genomics tools and provides users with rich capabilities to browse pre-computed whole-genome alignments of large vertebrate genomes and other groups of organisms with VISTA Browser, to submit their own sequences of interest to several VISTA servers for various types of comparative analysis and to obtain detailed comparative analysis results for a set of cardiovascular genes. We illustrate capabilities of the VISTA site by the analysis of a 180 kb interval on human chromosome 5 that encodes for the kinesin family member 3A (KIF3A) protein.

Binding Sites↗

A fast and sensitive algorithm for aligning ESTs to the human genome.

There is a pressing need to align the growing set of expressed sequence tags (ESTs) with the newly sequenced human genome. However, the problem is complicated by the exon/intron structure of eukaryotic genes misread nucleotides in ESTs, and the millions of repetitive sequences in genomic sequences. To solve this problem, algorithms that use dynamic programming have been proposed. In reality, however, these algorithms require an enormous amount of processing time. In an effort to improve the computational efficiency of these classical DP algorithms, we developed software that fully utilizes lookup-tables to detect the start- and endpoints of an EST within a given DNA sequence efficiently, and subsequently promptly identify exons and introns. In addition, the locations of all splice sites must be calculated correctly with high sensitivity and accuracy, while retaining high computational efficiency. This goal is hard to accomplish in practice, due to misread nucleotides in ESTs and repetitive sequences in the genome. Nevertheless, we present two heuristics that effectively settle this issue. Experimental results confirm that our technique improves the overall computation time by orders of magnitude compared with common tools, such as SIM4 and BLAT, and simultaneously attains high sensitivity and accuracy against a clean dataset of documented genes.

Algorithms↗

Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.

The spliced alignment of expressed sequence data to genomic sequence has proven a key tool in the comprehensive annotation of genes in eukaryotic genomes. A novel algorithm was developed to assemble clusters of overlapping transcript alignments (ESTs and full-length cDNAs) into maximal alignment assemblies, thereby comprehensively incorporating all available transcript data and capturing subtle splicing variations. Complete and partial gene structures identified by this method were used to improve The Institute for Genomic Research Arabidopsis genome annotation (TIGR release v.4.0). The alignment assemblies permitted the automated modeling of several novel genes and >1000 alternative splicing variations as well as updates (including UTR annotations) to nearly half of the approximately 27 000 annotated protein coding genes. The algorithm of the Program to Assemble Spliced Alignments (PASA) tool is described, as well as the results of automated updates to Arabidopsis gene annotations.

Algorithms↗

OWEN: aligning long collinear regions of genomes.

OWEN is an interactive tool for aligning two long DNA sequences that represents similarity between them by a chain of collinear local similarities. OWEN employs several methods for constructing and editing local similarities and for resolving conflicts between them. Alignments of sequences of lengths over 10(6) can often be produced in minutes. OWEN requires memory below 20 L, where L is the sum of lengths of the compared sequences.

Algorithms↗

Multimodal alignment improves generalizability of genomic biomarker prediction in computational pathology.

Computational pathology models that use digitized histopathology whole-slide images have the potential to become a cost-effective and scalable alternative to molecular assays for the prediction of genomic biomarkers, a key task in precision oncology. However, as new genomic biomarkers are discovered or quantified, large, labeled datasets must be prospectively collected to train new models. To address this challenge, we developed multimodal alignment for biomarker learning and generalization (MARBLE), a multimodal contrastive pretraining strategy that integrates structured biomarker knowledge into representation learning of histopathology images. MARBLE aligns histopathology-derived representations with representations of genomic biomarkers generated by a large language model (LLM) and a protein language model (PLM). This biologically informed alignment enables data-efficient generalization to novel, out-of-distribution biomarkers. Using the MSK-IMPACT cohort of over 40,000 patients across multiple biomarker panel versions, we design experiments grounded in real-world data to demonstrate the value of our proposed approach.

CP: computational biology↗

Differential regulation of the IL-10 gene in Th1 and Th2 T cells.

Interleukin-10 (IL-10), an immunoregulatory cytokine, modulates the function of various immune and nonimmune cells, yet little information is available on the molecular mechanism of transcriptional regulation at the chromatin level. During T cell differentiation from naive T cells into Th1 and Th2 cells, the expression of IL-10 in Th1 cells slowly disappears, whereas Th2 cells produce more IL-10. We examined the chromatin structural changes associated with IL-10 gene transcription by naive and differentiated murine Th1 and Th2 cells. Naive T cells lack DNase I hypersensitivity (HS) sites in the vicinity of the IL-10 gene, whereas differentiated T cells display a strong 3' constitutive HS site as well as several inducible sites. In committed Th1 cells, the mechanism of IL-10 gene silencing is associated with a closed chromatin structure, the lack of an HS site at the promoter region, and the development of repressive histone modification near the IL-10 promoter and introns 3 and 4. We confirm that the majority of HS sites coincide with conserved noncoding sequences (CNSs) identified by comparative genomic sequence alignment between human and mouse genomes. Potential transcription factor binding sites were located by comparing CNSs with the TRANSFAC database. Predicted in vivo binding of specific factors on the CNS locus were confirmed by chromatin immunoprecipitation assays. Our results suggest that the combination of HS site and comparative genomic approaches allows identification of regulatory elements involved in differential IL-10 gene expression between Th1 and Th2 cells during T cell differentiation.

Animals↗

Experimental validation of predicted mammalian erythroid cis-regulatory modules.

Multiple alignments of genome sequences are helpful guides to functional analysis, but predicting cis-regulatory modules (CRMs) accurately from such alignments remains an elusive goal. We predict CRMs for mammalian genes expressed in red blood cells by combining two properties gleaned from aligned, noncoding genome sequences: a positive regulatory potential (RP) score, which detects similarity to patterns in alignments distinctive for regulatory regions, and conservation of a binding site motif for the essential erythroid transcription factor GATA-1. Within eight target loci, we tested 75 noncoding segments by reporter gene assays in transiently transfected human K562 cells and/or after site-directed integration into murine erythroleukemia cells. Segments with a high RP score and a conserved exact match to the binding site consensus are validated at a good rate (50%-100%, with rates increasing at higher RP), whereas segments with lower RP scores or nonconsensus binding motifs tend to be inactive. Active DNA segments were shown to be occupied by GATA-1 protein by chromatin immunoprecipitation, whereas sites predicted to be inactive were not occupied. We verify four previously known erythroid CRMs and identify 28 novel ones. Thus, high RP in combination with another feature of a CRM, such as a conserved transcription factor binding site, is a good predictor of functional CRMs. Genome-wide predictions based on RP and a large set of well-defined transcription factor binding sites are available through servers at http://www.bx.psu.edu/.

Amino Acid Motifs↗

Alignment of Sfi I sites with the Not I restriction map of Schizosaccharomyces pombe genome.

A Sfi I restriction map of the fission yeast Schizosaccharomyces pombe genome was aligned with the Not I restriction map. There are 16 Sfi I sites in the S. pombe genome. Three Sfi I sites are on chromosome III which is devoid of Not I sites. The sizes of the entire genome and individual chromosomes, calculated from the Sfi I fragment sizes, are consistent with that calculated from the Not I fragment sizes. The Sfi I map provides greater physical characterization of the S. pombe genome and further validates the use of S. pombe chromosomal DNA as size standard. These maps have allowed detection of polymorphism on all three chromosomes.

Base Sequence↗

The UniMarker (UM) method for synteny mapping of large genomes.

MOTIVATION: Synteny mapping, or detecting regions that are orthologous between two genomes, is a key step in studies of comparative genomics. For completely sequenced genomes, this is increasingly accomplished by whole-genome sequence alignment. However, such methods are computationally expensive, especially for large genomes, and require rather complicated post-processing procedures to filter out non-orthologous sequence matches. RESULTS: We have developed a novel method that does not require sequence alignment for synteny mapping of two large genomes, such as the human and mouse. In this method, the occurrence spectra of genome-wide unique 16mer sequences present in both the human and mouse genome are used to directly detect orthologous genomic segments. Being sequence alignment-free, the method is very fast and able to map the two mammalian genomes in one day of computing time on a single Pentium IV personal computer. The resulting human-mouse synteny map was shown to be in excellent agreement with those produced by the Mouse Genome Sequencing Consortium (MGSC) and by the Ensembl team; furthermore, the syntenic relationship of segments found only by our method was supported by BLASTZ sequence alignment.

Algorithms↗

A hierarchical approach to aligning collinear regions of genomes.

MOTIVATION: As a first approximation, similarity between two long orthologous regions of genomes can be represented by a chain of local similarities. Within such a chain, pairs of successive similarities are collinear (non-conflicting), i.e. segments involved in the nth similarity precede in both sequences segments involved in the (n+1)th similarity. However, when all similarities between two long sequences are considered, usually there are many conflicts between them. Although some conflicts can be avoided by masking transposons or low-complexity sequences, selecting only those similarities that reflect orthology and, thus, belong to the evolutionarily true chain is not trivial. RESULTS: We propose a simple, hierarchical algorithm of finding the true chain of local similarities. Starting from similarities with low P-values, we resolve each pairwise conflict by deleting a similarity with a higher P-value. This greedy approach constructs a chain of similarities faster than when a chain optimal with respect to some global criterion is sought, and makes more sense biologically.

Algorithms↗

The sequence of rice chromosomes 11 and 12, rich in disease resistance genes and recent gene duplications.

BACKGROUND: Rice is an important staple food and, with the smallest cereal genome, serves as a reference species for studies on the evolution of cereals and other grasses. Therefore, decoding its entire genome will be a prerequisite for applied and basic research on this species and all other cereals. RESULTS: We have determined and analyzed the complete sequences of two of its chromosomes, 11 and 12, which total 55.9 Mb (14.3% of the entire genome length), based on a set of overlapping clones. A total of 5,993 non-transposable element related genes are present on these chromosomes. Among them are 289 disease resistance-like and 28 defense-response genes, a higher proportion of these categories than on any other rice chromosome. A three-Mb segment on both chromosomes resulted from a duplication 7.7 million years ago (mya), the most recent large-scale duplication in the rice genome. Paralogous gene copies within this segmental duplication can be aligned with genomic assemblies from sorghum and maize. Although these gene copies are preserved on both chromosomes, their expression patterns have diverged. When the gene order of rice chromosomes 11 and 12 was compared to wheat gene loci, significant synteny between these orthologous regions was detected, illustrating the presence of conserved genes alternating with recently evolved genes. CONCLUSION: Because the resistance and defense response genes, enriched on these chromosomes relative to the whole genome, also occur in clusters, they provide a preferred target for breeding durable disease resistance in rice and the isolation of their allelic variants. The recent duplication of a large chromosomal segment coupled with the high density of disease resistance gene clusters makes this the most recently evolved part of the rice genome. Based on syntenic alignments of these chromosomes, rice chromosome 11 and 12 do not appear to have resulted from a single whole-genome duplication event as previously suggested.

Chromosome Mapping↗

Generating multiple alignments on a pangenomic scale.

MOTIVATION: Since novel long read sequencing technologies allow for de novo assembly of many individuals of a species, high-quality assemblies are becoming widely available. For example, the recently published draft human pangenome reference was based on assemblies composed of contigs. There is an urgent need for a software-tool that is able to generate a multiple alignment of genomes of the same species because current multiple sequence alignment programs cannot deal with such a volume of data. RESULTS: We show that the combination of a well-known anchor-based method with the technique of prefix-free parsing yields an approach that is able to generate multiple alignments on a pangenomic scale, provided that large-scale structural variants are rare. Furthermore, experiments with real world data show that our software tool PANgenomic Anchor-based Multiple Alignment significantly outperforms current state-of-the art programs. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://gitlab.com/qwerzuiop/panama, archived at swh:1:dir:e90c9f664995acca9063245cabdd97549cf39694.

Software↗

Molecular variability analysis of five new complete cacao swollen shoot virus genomic sequences.

Cacao swollen shoot virus (CSSV), a member of the family Caulimovi-ridae, genus Badnavirus occurs in all the main cacao-growing areas of West Africa. We amplified, cloned and sequenced complete genomes of five new isolates, two originating from Togo and three originating from Ghana. The genome of these five newly sequenced isolates all contain the five putative open reading frames I, II, III, X and Y described for the first sequenced CSSV isolate, Agou1 originating from Togo. Their genomes have been aligned with the genome of Agou1. The nucleotide and amino acid sequence identities between isolates have been calculated and a phylogenetic analysis has been made including other pararetroviruses. Maximum nucleotide sequence variability between complete genomes of CSSV isolates was 29.4%. Geographical differentiation between isolates appears more important than differentiation between mild and severe isolates. ORF X differs greatly in size and sequence between the Togolese isolates Nyongbo2 and Agou1, and the four other isolates, its functional role is therefore clearly questionable.

Badnavirus↗

Structure of the human Nkx2.1 gene.

NKX2.1 is a member of the NK2 family of homeodomain-containing transcriptional factors which binds to and activates the promoters of thyroid and pulmonary epithelial genes. We have cloned and sequenced twelve human lung NKx2.1 cDNAs. To elucidate the origin of Nkx2.1 transcripts, we also cloned and sequenced a 12 kb human Nkx2.1 genomic clone. Alignment of cDNA sequences with the genomic clone showed that contrary to previous reports, the human Nkx2.1 gene is organized into three exons and two introns. The newly discovered exon I contains an ATG codon that falls in frame with the previously identified Nkx2.1 initiator ATG codon on one of the cDNAs, designated 5E. Northern blot analysis shows that an mRNA of approximately 2.5 kb in size, homologous to 5E, is expressed in both lung and thyroid. The deduced amino acid sequence of the longest open reading frame on 5E is identical to NKX2.1 with the exception of a 30 amino acid N-terminal extension. Coupled in vitro transcription/translation of the 5E cDNA confirms that the open reading frame is translated into a contiguous polypeptide of 44 kDa. Analysis of Nkx2.1 genomic DNA fragments suggest that at least two independent regions, one within the first intron and the other 5' of the first exon may mediate the basal promoter activity of the Nkx2.1 gene in lung epithelial cells.

Amino Acid Sequence↗

A probabilistic approach to consensus multiple alignment.

We consider the problem of obtaining the maximum a posteriori probability (MAP) estimate of a consensus ancestral sequence for a set of DNA sequences. Our maximization method, called ASA (dnA Sequence Alignment), can be applied to the refinement of noisy regions of a DNA assembly, to the alignment of genomic functional sites, or to the alignment of any set of DNA sequences related by a star-like phylogeny. Along with the optimal consensus, ASA finds suboptimal solutions together with their relative probabilities. The probabilistic approach makes it possible to establish the limits to which an ancestor can in principle be recovered from diverged sequences. In simulations on rather short synthetic sequences (of length up to 80) with different coverage and error rates ranging from 5% to 30%, ASA restored the consensus from noisy observations essentially as best as is theoretically possible for the given error rates. We also illustrate the performance of ASA on the alignment of E.Coli promoters and the Alu-Sb subfamily of human repeat sequences. Since our model is a special case of a profile HMM, we give a comparison between these two approaches, as well as with other DNA alignment methods.

Base Sequence↗