Search PubMed⌕ Search

Biomedical subjects

Panayiotis V Benos

Publications and source records attributed to Panayiotis V Benos.

14 recordsLinked to original sources

Web-based primer design software for genome-scale genotyping by pyrosequencing.

Design of locus-specific primers for use during genetic analysis requires combining information from multiple sources and can be a time-consuming process when validating large numbers of assays. Data warehousing of genomic DNA sequences and genetic variations when coupled with software applications for optimizing the generation of locus-specific primers can increase the efficiency of assay development. Selection of oligonucleotide primers for PCR and Pyrosequencing (SOP3) software allows user-directed queries of warehoused data collected from the human and mouse genome sequencing projects. The software automates collection of DNA sequence flanking single-nucleotide polymorphisms (SNPs) as well as the incorporation of locus-associated functional information, such as whether the SNP occurs in an exon, intron, or untranslated region. SOP3 software accepts three types of user-directed input consisting of gene locus symbols, SNP reference sequence numbers, or chromosomal physical location. For human polymorphisms, SOP3 incorporates haplotype, ethnicity, and SNP validation attributes. The output is a list of oligonucleotide primers recommended for Pyrosequencing-based typing of genetic variations. SOP3 is available at the Division of Immunogenetics computational server found at http://imgen.ccbb.pitt.edu.

Animals↗

Self-organizing neural networks to support the discovery of DNA-binding motifs.

Identification of the short DNA sequence motifs that serve as binding targets for transcription factors is an important challenge in bioinformatics. Unsupervised techniques from the statistical learning theory literature have often been applied to motif discovery, but effective solutions for large genomic datasets have yet to be found. We present here three self-organizing neural networks that have applicability to the motif-finding problem. The core system in this study is a previously described SOM-based motif-finder named SOMBRERO. The motif-finder is integrated in this work with a SOM-based method that automatically constructs generalized models for structurally related motifs and initializes SOMBRERO with relevant biological knowledge. A self-organizing tree method that displays the relationships between various motifs is also presented, and it is shown that such a method can act as an effective structural classifier of novel motifs. The performance of the three self-organizing neural networks is evaluated here using various datasets.

Algorithms↗

Reconstructing an ancestral mammalian immune supercomplex from a marsupial major histocompatibility complex.

The first sequenced marsupial genome promises to reveal unparalleled insights into mammalian evolution. We have used the Monodelphis domestica (gray short-tailed opossum) sequence to construct the first map of a marsupial major histocompatibility complex (MHC). The MHC is the most gene-dense region of the mammalian genome and is critical to immunity and reproductive success. The marsupial MHC bridges the phylogenetic gap between the complex MHC of eutherian mammals and the minimal essential MHC of birds. Here we show that the opossum MHC is gene dense and complex, as in humans, but shares more organizational features with non-mammals. The Class I genes have amplified within the Class II region, resulting in a unique Class I/II region. We present a model of the organization of the MHC in ancestral mammals and its elaboration during mammalian evolution. The opossum genome, together with other extant genomes, reveals the existence of an ancestral "immune supercomplex" that contained genes of both types of natural killer receptors together with antigen processing genes and MHC genes.

Animals↗

Deregulation of common genes by c-Myc and its direct target, MT-MC1.

In addition to its role in cancer, the c-Myc oncoprotein controls many normal cellular processes as a consequence of its function as a basic helix-loop-helix leucine zipper transcription factor. Determining which of the myriad genes under c-Myc control are relevant for these various roles is thus a major challenge. mt-mc1 is a direct c-Myc target gene whose overexpression recapitulates multiple c-Myc phenotypes, including transformation. Using transcriptional profiling, we now show that MT-MC1-overexpressing myeloid cells misregulate a total of 47 distinct transcripts, a large proportion of which are involved in signal transduction and/or cancer. Analysis of these genes reveals a consensus promoter structure consisting of multiple, often closely spaced c-Myc binding sites and three additional Wilm's tumor and Egr1-like motifs. More than one-third of MT-MC1 target genes are also clustered on six cancer-associated chromosomal loci. Most surprisingly, all of the transcripts examined also are regulated by c-Myc. Finally, an estrogen receptor-MT-MC1 fusion protein was used to establish that all examined transcripts were regulated directly by the chimeric protein. Our results thus indicate that MT-MC1 target genes largely comprise a subset of those regulated by c-Myc. We propose that the properties imparted by MT-MC1 are the result of its control of a small and select c-Myc target gene population.

Amino Acid Motifs↗

FOOTER: a web tool for finding mammalian DNA regulatory regions using phylogenetic footprinting.

FOOTER is a newly developed algorithm that analyzes homologous mammalian promoter sequences in order to identify transcriptional DNA regulatory 'signals'. FOOTER uses prior knowledge about the binding site preferences of the transcription factors (TFs) in the form of position-specific scoring matrices (PSSMs). The PSSM models are generated from known mammalian binding sites from the TRANSFAC database. In a test set of 72 confirmed binding sites (most of them not present in TRANSFAC) of 19 TFs, it exhibited 83% sensitivity and 72% specificity. FOOTER is accessible over the web at http://biodev.hgen.pitt.edu/Footer/.

Algorithms↗

enoLOGOS: a versatile web tool for energy normalized sequence logos.

enoLOGOS is a web-based tool that generates sequence logos from various input sources. Sequence logos have become a popular way to graphically represent DNA and amino acid sequence patterns from a set of aligned sequences. Each position of the alignment is represented by a column of stacked symbols with its total height reflecting the information content in this position. Currently, the available web servers are able to create logo images from a set of aligned sequences, but none of them generates weighted sequence logos directly from energy measurements or other sources. With the advent of high-throughput technologies for estimating the contact energy of different DNA sequences, tools that can create logos directly from binding affinity data are useful to researchers. enoLOGOS generates sequence logos from a variety of input data, including energy measurements, probability matrices, alignment matrices, count matrices and aligned sequences. Furthermore, enoLOGOS can represent the mutual information of different positions of the consensus sequence, a unique feature of this tool. Another web interface for our software, C2H2-enoLOGOS, generates logos for the DNA-binding preferences of the C2H2 zinc-finger transcription factor family members. enoLOGOS and C2H2-enoLOGOS are accessible over the web at http://biodev.hgen.pitt.edu/enologos/.

Amino Acids↗

Improved detection of DNA motifs using a self-organized clustering of familial binding profiles.

MOTIVATION: One of the limiting factors in deciphering transcriptional regulatory networks is the effectiveness of motif-finding software. An emerging avenue for improving motif-finding accuracy aims to incorporate generalized binding constraints of related transcription factors (TFs), named familial binding profiles (FBPs), as priors in motif identification methods. A motif-finder can thus be 'biased' towards finding motifs from a particular TF family. However, current motif-finders allow only a single FBP to be used as a prior in a given motif-finding run. In addition, current FBP construction methods are based on manual clustering of position specific scoring matrices (PSSMs) according to the known structural properties of the TF proteins. Manual clustering assumes that the binding preferences of structurally similar TFs will also be similar. This assumption is not true, at least not for some TF families. Automatic PSSM clustering methods are thus required for augmenting the usefulness of FBPs. RESULTS: A novel method is developed for automatic clustering of PSSM models. The resulting FBPs are incorporated into the SOMBRERO motif-finder, significantly improving its performance when finding motifs related to those that have been incorporated. SOMBRERO is thus the only existing de novo motif-finder that can incorporate knowledge of all known PSSMs in a given motif-finding run. AVAILABILITY: The methods outlined will be incorporated into the next release of SOMBRERO, which is available from http://bioinf.nuigalway.ie/sombrero

Algorithms↗

Footer: a quantitative comparative genomics method for efficient recognition of cis-regulatory elements.

The search for mammalian DNA regulatory regions poses a challenging problem in computational biology. The short length of the DNA patterns compared with the size of the promoter regions and the degeneracy of the patterns makes their identification difficult. One way to overcome this problem is to use evolutionary information to reduce the number of false-positive predictions. We developed a novel method for pattern identification that compares a pair of putative binding sites in two species (e.g., human and mouse) and assigns two probability scores based on the relative position of the sites in the promoter and their agreement with a known model of binding preferences. We tested the algorithm's ability to predict known binding sites on various promoters. Overall, it exhibited 83% sensitivity and the specificity was 72%, which is a clear improvement over existing methods. Our algorithm also successfully predicted two novel NF-kappaB binding sites in the promoter region of the mouse autotaxin gene (ATX, ENPP2), which we were able to verify by using chromatin immunoprecipitation assay coupled with quantitative real-time PCR.

Algorithms↗

SOP3: a web-based tool for selection of oligonucleotide primers for single nucleotide polymorphism analysis by Pyrosequencing.

SOP3 is a web-based software tool for designing oligonucleotide primers for use in the analysis of single nucleotide polymorphisms (SNPs). Accessible via the Internet, the application is optimized for developing the PCR and sequencing primers that are necessary for Pyrosequencing. The application accepts as input gene name, SNP reference sequence number, or chromosomal nucleotide location. Output can be parsed by gene name, SNP reference number, heterozygosity value, location, chromosome, or function. The location of an individual polymorphism, such as an intron, exon, or 5' or 3' untranslated region is indicated, as are whether nucleotide changes in an exon are associated with a change in an amino acid sequence. SOP3 presents for each entry a set of forward and biotinylated reverse PCR primers as well as a sequencing primer for use during the analysis of SNPs by Pyrosequencing. Theoretical pyrograms for each allele are calculated and presented graphically. The method has been tested in the development of Pyrosequencing assays for determining SNPs and for deletion/insertion polymorphisms in the human genome. Of the SOP3-designed primer sets that were tested, a large majority of the primer sets have successfully produced PCR products and Pyrosequencing data.

Algorithms↗

A sequence alignment-independent method for protein classification.

Annotation of the rapidly accumulating body of sequence data relies heavily on the detection of remote homologues and functional motifs in protein families. The most popular methods rely on sequence alignment. These include programs that use a scoring matrix to compare the probability of a potential alignment with random chance and programs that use curated multiple alignments to train profile hidden Markov models (HMMs). Related approaches depend on bootstrapping multiple alignments from a single sequence. However, alignment-based programs have limitations. They make the assumption that contiguity is conserved between homologous segments, which may not be true in genetic recombination or horizontal transfer. Alignments also become ambiguous when sequence similarity drops below 40%. This has kindled interest in classification methods that do not rely on alignment. An approach to classification without alignment based on the distribution of contiguous sequences of four amino acids (4-grams) was developed. Interest in 4-grams stemmed from the observation that almost all theoretically possible 4-grams (20(4)) occur in natural sequences and the majority of 4-grams are uniformly distributed. This implies that the probability of finding identical 4-grams by random chance in unrelated sequences is low. A Bayesian probabilistic model was developed to test this hypothesis. For each protein family in Pfam-A and PIR-PSD, a feature vector called a probe was constructed from the set of 4-grams that best characterised the family. In rigorous jackknife tests, unknown sequences from Pfam-A and PIR-PSD were compared with the probes for each family. A classification result was deemed a true positive if the probe match with the highest probability was in first place in a rank-ordered list. This was achieved in 70% of cases. Analysis of false positives suggested that the precision might approach 85% if selected families were clustered into subsets. Case studies indicated that the 4-grams in common between an unknown and the best matching probe correlated with functional motifs from PRINTS. The results showed that remote homologues and functional motifs could be identified from an analysis of 4-gram patterns.

Algorithms↗

Probabilistic code for DNA recognition by proteins of the EGR family.

A recognition code for protein-DNA interactions would allow for the prediction of binding sites based on protein sequence, and the identification of binding proteins for specific DNA targets. Crystallographic studies of protein-DNA complexes showed that a simple, deterministic recognition code does not exist. Here, we present a probabilistic recognition code (P-code) that assigns energies to all possible base-pair-amino acid interactions for the early growth response factor (EGR) family of zinc-finger transcription factors. The specific energy values are determined by a maximum likelihood method using examples from in vitro randomisation experiments (namely, SELEX and phage display) reported in the literature. The accuracy of the model is tested in several ways, including the ability to predict in vivo binding sites of EGR proteins and other non-EGR zinc-finger proteins, and the correlation between predicted and measured binding affinities of various EGR proteins to several different DNA sites. We also show that this model improves significantly upon the prediction capabilities of previous qualitative and quantitative models. The probabilistic code we develop uses information about the interacting positions between the protein and DNA, but we show that such information is not necessary, although it reduces the number of parameters to be determined. We also employ the assumption that the total binding energy is the sum of the energies of the individual contacts, but we describe how that assumption can be relaxed at the cost of additional parameters.

Algorithms↗

Additivity in protein-DNA interactions: how good an approximation is it?

Man and Stormo and Bulyk et al. recently presented their results on the study of the DNA binding affinity of proteins. In both of these studies the main conclusion is that the additivity assumption, usually applied in methods to search for binding sites, is not true. In the first study, the analysis of binding affinity data from the Mnt repressor protein bound to all possible DNA (sub)targets at positions 16 and 17 of the binding site, showed that those positions are not independent. In the second study, the authors analysed DNA binding affinity data of the wild-type mouse EGR1 protein and four variants differing on the middle finger. The binding affinity of these proteins was measured to all 64 possible trinucleotide (sub)targets of the middle finger using microarray technology. The analysis of the measurements also showed interdependence among the positions in the DNA target. In the present report, we review the data of both studies and we re- analyse them using various statistical methods, including a comparison with a multiple regression approach. We conclude that despite the fact that the additivity assumption does not fit the data perfectly, in most cases it provides a very good approximation of the true nature of the specific protein-DNA interactions. Therefore, additive models can be very useful for the discovery and prediction of binding sites in genomic DNA.

Animals↗

Is there a code for protein-DNA recognition? Probab(ilistical)ly. . .

Transcriptional regulation of all genes is initiated by the specific binding of regulatory proteins called transcription factors to specific sites on DNA called promoter regions. Transcription factors employ a variety of mechanisms to recognise their DNA target sites. In the last few decades, attempts have been made to describe these mechanisms by general sets of rules and associated models. We give an overview of these models, starting with a historical review of the somewhat controversial issue of a "recognition code" governing protein-DNA interaction. We then present a probabilistic framework in which advantages and disadvantages of various models can be discussed. Finally, we conclude that simplifying assumptions about additivity of interactions are sufficiently justified in many situations (and can be suitably extended in other situations) to allow a unifying concept of a "probabilistic code" for protein-DNA recognition to be defined.

Animals↗

Mapping and identification of essential gene functions on the X chromosome of Drosophila.

The Drosophila melanogaster genome consists of four chromosomes that contain 165 Mb of DNA, 120 Mb of which are euchromatic. The two Drosophila Genome Projects, in collaboration with Celera Genomics Systems, have sequenced the genome, complementing the previously established physical and genetic maps. In addition, the Berkeley Drosophila Genome Project has undertaken large-scale functional analysis based on mutagenesis by transposable P element insertions into autosomes. Here, we present a large-scale P element insertion screen for vital gene functions and a BAC tiling map for the X chromosome. A collection of 501 X-chromosomal P element insertion lines was used to map essential genes cytogenetically and to establish short sequence tags (STSs) linking the insertion sites to the genome. The distribution of the P element integration sites, the identified genes and transcription units as well as the expression patterns of the P-element-tagged enhancers is described and discussed.

Animals↗