Search PubMed⌕ Search

PubMed · 10380193

A probabilistic approach to consensus multiple alignment.

Abstract

We consider the problem of obtaining the maximum a posteriori probability (MAP) estimate of a consensus ancestral sequence for a set of DNA sequences. Our maximization method, called ASA (dnA Sequence Alignment), can be applied to the refinement of noisy regions of a DNA assembly, to the alignment of genomic functional sites, or to the alignment of any set of DNA sequences related by a star-like phylogeny. Along with the optimal consensus, ASA finds suboptimal solutions together with their relative probabilities. The probabilistic approach makes it possible to establish the limits to which an ancestor can in principle be recovered from diverged sequences. In simulations on rather short synthetic sequences (of length up to 80) with different coverage and error rates ranging from 5% to 30%, ASA restored the consensus from noisy observations essentially as best as is theoretically possible for the given error rates. We also illustrate the performance of ASA on the alignment of E.Coli promoters and the Alu-Sb subfamily of human repeat sequences. Since our model is a special case of a profile HMM, we give a comparison between these two approaches, as well as with other DNA alignment methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B Lazareva-Ulitsky, D Haussler. 1999. A probabilistic approach to consensus multiple alignment.. https://doi.org/10.1142/9789814447300_0015

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence↗

Identification and characterization of Prp45p and Prp46p, essential pre-mRNA splicing factors.

Through exhaustive two-hybrid screens using a budding yeast genomic library, and starting with the splicing factor and DEAH-box RNA helicase Prp22p as bait, we identified yeast Prp45p and Prp46p. We show that as well as interacting in two-hybrid screens, Prp45p and Prp46p interact with each other in vitro. We demonstrate that Prp45p and Prp46p are spliceosome associated throughout the splicing process and both are essential for pre-mRNA splicing. Under nonsplicing conditions they also associate in coprecipitation assays with low levels of the U2, U5, and U6 snRNAs that may indicate their presence in endogenous activated spliceosomes or in a postsplicing snRNP complex.

Base Sequence↗

Large-scale evaluation of imprinting status in the Prader-Willi syndrome region: an imprinted direct repeat cluster resembling small nucleolar RNA genes.

Loss of paternal gene expression at the imprinted domain on proximal human chromosome 15 causes Prader-Willi syndrome (PWS), a complex multiple-anomaly disorder involving variable mental retardation, hyperphasia leading to obesity and infantile hypotonia with failure to thrive. Although numerous paternally expressed transcripts have been identified that reside in the candidate region, the individual contributions to the development of PWS have not been firmly established. Recent studies of mouse models carrying a cytogenetic deletion suggest that paternal deficiency of the SNRPN-IPW interval is critical for perinatal lethality of potential relevance to PWS. Here we determined the allelic expression profiles of a total of 118 cDNA clones using monochromosomal hybrids retaining either a paternal or maternal human chromosome 15. Our results demonstrated a preponderance of unusual transcripts lacking protein-coding potential that were expressed exclusively from the paternal copy of the critical interval. This interval was also found to encompass a large direct repeat (DR) cluster displaying a potentially active chromatin conformation of paternal origin, as suggested by enhanced sensitivity to nuclease digestion. Database searches revealed an unexpected organization of tandemly repeated consensus elements, all of which possessed well-defined box C and D sequences characteristic of small nucleolar RNAs (snoRNAs). Southern blot analysis further demonstrated a considerable degree of phylogenetic conservation of the DR locus in the genomes of all mammalian species tested, but not in chicken, Xenopus and Drosophila. These findings imply a potential direct contribution of the DR locus, representing a cluster of multiple snoRNA genes, to certain phenotypic features of PWS.

Base Sequence↗