Search PubMed⌕ Search

Biomedical subjects

Gary Benson

Publications and source records attributed to Gary Benson.

10 recordsLinked to original sources

TRDB--the Tandem Repeats Database.

Tandem repeats in DNA have been under intensive study for many years, first, as a consequence of their usefulness as genomic markers and DNA fingerprints and more recently as their role in human disease and regulatory processes has become apparent. The Tandem Repeats Database (TRDB) is a public repository of information on tandem repeats in genomic DNA. It contains a variety of tools for repeat analysis, including the Tandem Repeats Finder program, query and filtering capabilities, repeat clustering, polymorphism prediction, PCR primer selection, data visualization and data download in a variety of formats. In addition, TRDB serves as a centralized research workbench. It provides user storage space and permits collaborators to privately share their data and analysis. TRDB is available at https://tandem.bu.edu/cgi-bin/trdb/trdb.exe.

Animals↗

Oligonucleotide fingerprint identification for microarray-based pathogen diagnostic assays.

MOTIVATION: Advances in DNA microarray technology and computational methods have unlocked new opportunities to identify 'DNA fingerprints', i.e. oligonucleotide sequences that uniquely identify a specific genome. We present an integrated approach for the computational identification of DNA fingerprints for design of microarray-based pathogen diagnostic assays. We provide a quantifiable definition of a DNA fingerprint stated both from a computational as well as an experimental point of view, and the analytical proof that all in silico fingerprints satisfying the stated definition are found using our approach. RESULTS: The presented computational approach is implemented in an integrated high-performance computing (HPC) software tool for oligonucleotide fingerprint identification termed TOFI. We employed TOFI to identify in silico DNA fingerprints for several bacteria and plasmid sequences, which were then experimentally evaluated as potential probes for microarray-based diagnostic assays. Results and analysis of approximately 150 in silico DNA fingerprints for Yersinia pestis and 250 fingerprints for Francisella tularensis are presented. AVAILABILITY: The implemented algorithm is available upon request.

Algorithms↗

Indel seeds for homology search.

We are interested in detecting homologous genomic DNA sequences with the goal of locating approximate inverted, interspersed, and tandem repeats. Standard search techniques start by detecting small matching parts, called seeds, between a query sequence and database sequences. Contiguous seed models have existed for many years. Recently, spaced seeds were shown to be more sensitive than contiguous seeds without increasing the random hit rate. To determine the superiority of one seed model over another, a model of homologous sequence alignment must be chosen. Previous studies evaluating spaced and contiguous seeds have assumed that matches and mismatches occur within these alignments, but not insertions and deletions (indels). This is perhaps appropriate when searching for protein coding sequences (<5% of the human genome), but is inappropriate when looking for repeats in the majority of genomic sequence where indels are common. In this paper, we assume a model of homologous sequence alignment which includes indels and we describe a new seed model, called indel seeds, which explicitly allows indels. We present a waiting time formula for computing the sensitivity of an indel seed and show that indel seeds significantly outperform contiguous and spaced seeds when homologies include indels. We discuss the practical aspect of using indel seeds and finally we present results from a search for inverted repeats in the dog genome using both indel and spaced seeds.

Algorithms↗

Evaluating distance functions for clustering tandem repeats.

Tandem repeats are an important class of DNA repeats and much research has focused on their efficient identification, their use in DNA typing and fingerprinting, and their causative role in trinucleotide repeat diseases such as Huntington Disease, myotonic dystrophy, and Fragile-X mental retardation. We are interested in clustering tandem repeats into groups or families based on sequence similarity so that their biological importance may be further explored. To cluster tandem repeats we need a notion of pairwise distance which we obtain by alignment. In this paper we evaluate five distance functions used to produce those alignments--Consensus, Euclidean, Jensen-Shannon Divergence, Entropy-Surface, and Entropy-weighted. It is important to analyze and compare these functions because the choice of distance metric forms the core of any clustering algorithm. We employ a novel method to compare alignments and thereby compare the distance functions themselves. We rank the distance functions based on the cluster validation techniques--Average Cluster Density and Average Silhouette Width. Finally, we propose a multi-phase clustering method which produces good-quality clusters. In this study, we analyze clusters of tandem repeats from five sequences: Human Chromosomes 3, 5, 10 and X and C. elegans Chromosome III.

Algorithms↗

Inverted repeat structure of the human genome: the X-chromosome contains a preponderance of large, highly homologous inverted repeats that contain testes genes.

We have performed the first genome-wide analysis of the Inverted Repeat (IR) structure in the human genome, using a novel and efficient software package called Inverted Repeats Finder (IRF). After masking of known repetitive elements, IRF detected 22,624 human IRs characterized by arm size from 25 bp to >100 kb with at least 75% identity, and spacer length up to 100 kb. This analysis required 6 h on a desktop PC. In all, 166 IRs had arm lengths >8 kb. From this set, IRs were excluded if they were in unfinished/unassembled regions of the genome, or clustered with other closely related IRs, yielding a set of 96 large IRs. Of these, 24 (25%) occurred on the X-chromosome, although it represents only approximately 5% of the genome. Of the X-chromosome IRs, 83.3% were >/=99% identical, compared with 28.8% of autosomal IRs. Eleven IRs from Chromosome X, one from Chromosome 11, and seven already described from Chromosome Y contain genes predominantly expressed in testis. PCR analysis of eight of these IRs correctly amplified the corresponding region in the human genome, and six were also confirmed in gorilla or chimpanzee genomes. Similarity dot-plots revealed that 22 IRs contained further secondary homologous structures partially categorized into three distinct patterns. The prevalence of large highly homologous IRs containing testes genes on the X- and Y-chromosomes suggests a possible role in male germ-line gene expression and/or maintaining sequence integrity by gene conversion.

Animals↗

Minimal entropy probability paths between genome families.

We develop a metric for probability distributions with applications to biological sequence analysis. Our distance metric is obtained by minimizing a functional defined on the class of paths over probability measures on N categories. The underlying mathematical theory is connected to a constrained problem in the calculus of variations. The solution presented is a numerical solution, which approximates the true solution in a set of cases called rich paths where none of the components of the path is zero. The functional to be minimized is motivated by entropy considerations, reflecting the idea that nature might efficiently carry out mutations of genome sequences in such a way that the increase in entropy involved in transformation is as small as possible. We characterize sequences by frequency profiles or probability vectors, in the case of DNA where N is 4 and the components of the probability vector are the frequency of occurrence of each of the bases A, C, G and T. Given two probability vectors a and b, we define a distance function based as the infimum of path integrals of the entropy function H( p) over all admissible paths p(t), 0 < or = t< or =1, with p(t) a probability vector such that p(0)=a and p(1)=b. If the probability paths p(t) are parameterized as y(s) in terms of arc length s and the optimal path is smooth with arc length L, then smooth and "rich" optimal probability paths may be numerically estimated by a hybrid method of iterating Newton's method on solutions of a two point boundary value problem, with unknown distance L between the abscissas, for the Euler-Lagrange equations resulting from a multiplier rule for the constrained optimization problem together with linear regression to improve the arc length estimate L. Matlab code for these numerical methods is provided which works only for "rich" optimal probability vectors. These methods motivate a definition of an elementary distance function which is easier and faster to calculate, works on non-rich vectors, does not involve variational theory and does not involve differential equations, but is a better approximation of the minimal entropy path distance than the distance //b-a//(2). We compute minimal entropy distance matrices for examples of DNA myostatin genes and amino-acid sequences across several species. Output tree dendograms for our minimal entropy metric are compared with dendograms based on BLAST and BLAST identity scores.

Algorithms↗

Predicting human minisatellite polymorphism.

We seek to define sequence-based predictive criteria to identify polymorphic and hypermutable minisatellites in the human genome. Polymorphism of a representative pool of minisatellites, selected from human chromosomes 21 and 22, was experimentally measured by PCR typing in a population of unrelated individuals. Two predictive approaches were tested. One uses simple repeat characteristics (e.g., unit length, copy number, nucleotide bias) and a more complex measure, termed HistoryR, based on the presence of variant motifs in the tandem array. We find that HistoryR and percentage of GC are strongly correlated with polymorphism and, as predictive criteria, reduce by half the number of repeats to type while enriching the proportion with heterozygosity >/=0.5, from a background level of 43% to 59%. The second approach uses length differences between minisatellites in the two releases of the human genome sequence (from the public consortium and Celera). As a predictor, this similarly enriches the number of polymorphic minisatellites, but fails to identify an unexpectedly large number of these. Finally, typing of the highly polymorphic minisatellites in large families identified one new hypermutable minisatellite, located in a predicted coding sequence. This may represent the first coding human hypermutable minisatellite.

Chromosomes, Human, Pair 21↗

Mutation Master: profiles of substitutions in hepatitis C virus RNA of the core, alternate reading frame, and NS2 coding regions.

The RNA genome of the hepatitis C virus (HCV) undergoes rapid evolutionary change. Efforts to control this virus would benefit from the advent of facile methods to identify characteristic features of HCV RNA and proteins, and to condense the vast amount of mutational data into a readily interpretable form. Many HCV sequences are available in GenBank. To facilitate analysis, consensus sequences were constructed to eliminate the overrepresentation of certain genotypes, such as genotype 1, and a novel package of sequence analysis tools was developed. Mutation Master generates profiles of point mutations in a population of sequences and produces a set of visual displays and tables indicating the number, frequency, and character of substitutions. It can be used to analyze hundreds of sequences at a time. When applied to 255 HCV core protein sequences, Mutation Master identified variable domains and a series of mutations meriting further investigation. It flagged position 4, for example, where 90% or more of all sequences in genotypes 1, 2, 4, and 5, have N4, whereas those in genotypes 3, 6, 7, 8, 9, and 10 have L4. This pattern is noteworthy: L (hydrophobic) to N (polar) substitutions are generally rare, and genotypes 1, 2, 4, and 5 do not form a recognized super family of sequences. Thus, the L4N substitution probably arose independently several times. Moreover, not one member of genotypes 1, 2, 4, or 5 has L4 and not one member of genotypes 3, 6, 7, 8, 9, or 10 has N4. This nonoverlapping pattern suggests that coordinated changes at position 4 and a second site are required to yield a viable virus. The package generated a table of genotype-specific substitutions whose future analysis may help to identify interacting amino acids. Three substitutions were present in 100% of genotype 2 members and absent from all others: A68D, R74K, and R114H. Finally, this study revealed thatARFP, a novel protein encoded in an overlapping reading frame, is as conserved as conventional HCV proteins, a result supporting a role for ARFP in the viral life cycle. Whereas most conventional programs for phylogenetic analysis of sequences provide information about overall relatedness of genes or genomes, this program highlights and profiles point mutations. This is important because determinants of pathogenicity and drug susceptibility are likely to result from changes at only one or two key nucleotides or amino acid sites, and would not be detected by the type of pairwise comparisons that have usually been performed on HCV to date. This study is the first application of Mutation Master, which is now available upon request (http://tandem.biomath.mssm.edu/mutationmaster.html).

Amino Acid Motifs↗

Replication and compartmentalization of HIV-1 in kidney epithelium of patients with HIV-associated nephropathy.

HIV-associated nephropathy is a clinicopathologic entity that includes proteinuria, focal segmental glomerulosclerosis often of the collapsing variant, and microcystic tubulointerstitial disease. Increasing evidence supports a role for HIV-1 infection of renal epithelium in the pathogenesis of HIV-associated nephropathy. Using in situ hybridization, we previously demonstrated HIV-1 gag and nef mRNA in renal epithelial cells of patients with HIV-associated nephropathy. Here, to investigate whether renal epithelial cells were productively infected by HIV-1, we examined renal tissue for the presence of HIV-1 DNA and mRNA by in situ hybridization and PCR, and we molecularly characterized the HIV-1 quasispecies in the renal compartment. Infected renal epithelial cells were removed by laser-capture microdissection from biopsies of two patients, DNA was extracted, and HIV-1 V3-loop or gp120-envelope sequences were amplified from individually dissected cells by nested PCR. Phylogenetic analysis of kidney-derived sequences as well as corresponding sequences from peripheral blood mononuclear cells of the same patients revealed evidence of tissue-specific viral evolution. In phylogenetic trees constructed from V3 and gp120 sequences, kidney-derived sequences formed tissue-specific subclusters within the radiation of blood mononuclear cell-derived viral sequences from both patients. These data, along with the detection of HIV-1-specific proviral DNA and mRNA in tubular epithelium cells, argue strongly for localized replication of HIV-1 in the kidney and the existence of a renal viral reservoir.

Base Sequence↗

A new distance measure for comparing sequence profiles based on path lengths along an entropy surface.

We describe a new distance measure for comparing DNA sequence profiles. For this measure, columns in a multiple alignment are treated as character frequency vectors (sum of the frequencies equal to one). The distance between two vectors is based on minimum path length along an entropy surface. Path length is estimated using a random graph generated on the entropy surface and Dijkstra's algorithm for all shortest paths to a source. We use the new distance measure to analyze similarities within familes of tandem repeats in the C. elegans genome and show that this new measure gives more accurate refinement of family relationships than a method based on comparing consensus sequences.

Algorithms↗