Search PubMed⌕ Search

Biomedical subjects

W Hide

Publications and source records attributed to W Hide.

10 recordsLinked to original sources

Conserved domains of subtype C nef from South African HIV type 1-infected individuals include cytotoxic T lymphocyte epitope-rich regions.

We have characterized 43 nef sequences from subtype C HIV-1-infected South Africans and compared deduced amino acid sequences with other subtypes to identify areas of conservation. Our Nef amino acid sequences were aligned with a consensus subtype B, HXB2 reference strain and a consensus subtype C sequence. All were found to be highly homologous to subtype B in the central region of Nef, but more variable at the N and C termini of the molecule. Alignment of a consensus amino acid sequence generated from South African subtype C Nef with subtypes A, B, and D underscores cross-clade conservation in the central domain of the molecule. This domain is also rich in previously described cytotoxic T lymphocyte (CTL) epitopes that are restricted by commonly found HLA molecules in the South African population.

Amino Acid Sequence↗

The ESAT-6 gene cluster of Mycobacterium tuberculosis and other high G+C Gram-positive bacteria.

BACKGROUND: The genome of Mycobacterium tuberculosis H37Rv has five copies of a cluster of genes known as the ESAT-6 loci. These clusters contain members of the CFP-10 (lhp) and ESAT-6 (esat-6) gene families (encoding secreted T-cell antigens that lack detectable secretion signals) as well as genes encoding secreted, cell-wall-associated subtilisin-like serine proteases, putative ABC transporters, ATP-binding proteins and other membrane-associated proteins. These membrane-associated and energy-providing proteins may function to secrete members of the ESAT-6 and CFP-10 protein families, and the proteases may be involved in processing the secreted peptide. RESULTS: Finished and unfinished genome sequencing data of 98 publicly available microbial genomes has been analyzed for the presence of orthologs of the ESAT-6 loci. The multiple duplicates of the ESAT-6 gene cluster found in the genome of M. tuberculosis H37Rv are also conserved in the genomes of other mycobacteria, for example M. tuberculosis CDC1551, M. tuberculosis 210, M. bovis, M. leprae, M. avium, and the avirulent strain M. smegmatis. Phylogenetic analyses of the resulting sequences have established the duplication order of the gene clusters and demonstrated that the gene cluster known as region 4 (Rv3444c-3450c) is ancestral. Region 4 is also the only region for which an ortholog could be found in the genomes of Corynebacterium diphtheriae and Streptomyces coelicolor. CONCLUSIONS: Comparative genomic analysis revealed that the presence of the ESAT-6 gene cluster is a feature of some high-G+C Gram-positive bacteria. Multiple duplications of this cluster have occurred and are maintained only within the genomes of members of the genus Mycobacterium.

Antigens, Bacterial↗

STACK: Sequence Tag Alignment and Consensus Knowledgebase.

STACK is a tool for detection and visualisation of expressed transcript variation in the context of developmental and pathological states. The datasystem organizes and reconstructs human transcripts from available public data in the context of expression state. The expression state of a transcript can include developmental state, pathological association, site of expression and isoform of expressed transcript. STACK consensus transcripts are reconstructed from clusters that capture and reflect the growing evidence of transcript diversity. The comprehensive capture of transcript variants is achieved by the use of a novel clustering approach that is tolerant of sub-sequence diversity and does not rely on pairwise alignment. This is in contrast with other gene indexing projects. STACK is generated at least four times a year and represents the exhaustive processing of all publicly available human EST data extracted from GenBank. This processed information can be explored through 15 tissue-specific categories, a disease-related category and a whole-body index and is accessible via WWW at http://www.sanbi.ac.za/Dbases.html. STACK represents a broadly applicable resource, as it is the only reconstructed transcript database for which the tools for its generation are also broadly available (http://www.sanbi.ac.za/CODES).

Base Sequence↗

Molecular evolution of Mycobacterium tuberculosis: phylogenetic reconstruction of clonal expansion.

SETTING: M. tuberculosis isolates were collected from patients attending health clinics in a high incidence urban community and in a low incidence rural setting in South Africa. OBJECTIVE: To reconstruct the evolutionary history of a group of closely related M. tuberculosis isolates using IS6110, DRr and MTB484(1) restriction fragment length polymorphism (RFLP) data. DESIGN: Mycobacterium tuberculosis isolates containing an average of ten IS6110 elements, with a similarity index of > or = 65% were genotypically classified by DNA fingerprinting using the IS6110 derived probes IS-3' and IS-5', as well as the DRr and MTB484(1) probes, in combination with PvuII or Hinfl endonuclease digestion. These RFLP data were subjected to phylogenetic analysis using both genetic distance and parsimony algorithms. RESULTS: Phylogenetic analysis predicted the existence of two independently evolving lineages, possibly evolving from a common ancestral strain. The topology of the phylogenetic tree was supported by comprehensive bootstrapping and the specific partitioning of DNA methylation phenotypes. The observed difference in the branch lengths of the two lineages may suggest differential evolutionary rates. Isolates collected from different geographical regions demonstrate independent evolution, suggesting that it is highly unlikely that strains have been recently transmitted between the two regions. The number of evolutionary events identified in this strain family differs significantly from that of previously characterized strain families, implying that evolutionary rate may be strain family dependent. CONCLUSION: Based on this analysis we propose that the algorithm used to calculate recent epidemiological events should be revised to incorporate the evolutionary characteristics of individual strain families, thereby enhancing the accuracy of molecular epidemiological calculations.

Algorithms↗

d2_cluster: a validated method for clustering EST and full-length cDNAsequences.

Several efforts are under way to condense single-read expressed sequence tags (ESTs) and full-length transcript data on a large scale by means of clustering or assembly. One goal of these projects is the construction of gene indices where transcripts are partitioned into index classes (or clusters) such that they are put into the same index class if and only if they represent the same gene. Accurate gene indexing facilitates gene expression studies and inexpensive and early partial gene sequence discovery through the assembly of ESTs that are derived from genes that have yet to be positionally cloned or obtained directly through genomic sequencing. We describe d2_cluster, an agglomerative algorithm for rapidly and accurately partitioning transcript databases into index classes by clustering sequences according to minimal linkage or "transitive closure" rules. We then evaluate the relative efficiency of d2_cluster with respect to other clustering tools. UniGene is chosen for comparison because of its high quality and wide acceptance. It is shown that although d2_cluster and UniGene produce results that are between 83% and 90% identical, the joining rate of d2_cluster is between 8% and 20% greater than UniGene. Finally, we present the first published rigorous evaluation of under and over clustering (in other words, of type I and type II errors) of a sequence clustering algorithm, although the existence of highly identical gene paralogs means that care must be taken in the interpretation of the type II error. Upper bounds for these d2_cluster error rates are estimated at 0.4% and 0.8%, respectively. In other words, the sensitivity and selectivity of d2_cluster are estimated to be >99.6% and 99.2%.

Algorithms↗

Alternative gene form discovery and candidate gene selection from gene indexing projects.

Several efforts are under way to partition single-read expressed sequence tag (EST), as well as full-length transcript data, into large-scale gene indices, where transcripts are in common index classes if and only if they share a common progenitor gene. Accurate gene indexing facilitates gene expression studies, as well as inexpensive and early gene sequence discovery through assembly of ESTs that are derived from genes that have not been sequenced by classical methods. We extend, correct, and enhance the information obtained from index groups by splitting index classes into subclasses based on sequence dissimilarity (diversity). Two applications of this are highlighted in this report. First it is shown that our method can ameliorate the damage that artifacts, such as chimerism, inflict on index integrity. Additionally, we demonstrate how the organization imposed by an effective subpartition can greatly increase the sensitivity of gene expression studies by accounting for the existence and tissue- or pathology-specific regulation of novel gene isoforms and polymorphisms. We apply our subpartitioning treatment to the UniGene gene indexing project to measure a marked increase in information quality and abundance (in terms of assembly length and insertion/deletion error) after treatment and demonstrate cases where new levels of information concerning differential expression of alternate gene forms, such as regulated alternative splicing, are discovered. [Tables 2 and 3 can be viewed in their entirety as Online Supplements at http://www.genome.org.]

Alternative Splicing↗

Biological evaluation of d2, an algorithm for high-performance sequence comparison.

A number of algorithms exist for searching sequence databases for biologically significant similarities based on the primary sequence similarity of aligned sequences. We have determined the biological sensitivity and selectivity of d2, a high-performance comparison algorithm that rapidly determines the relative dissimilarity of large datasets of genetic sequences. d2 uses sequence-word multiplicity as a simple measure of dissimilarity. It is not constrained by the comparison of direct sequence alignments and so can use word contexts to yield new information on relationships. It is extremely efficient, comparing a query of length 884 bases (INS1ECLAC) with 19,540,603 bases of the bacterial division of GenBank (release 76.0) in 51.77 CPU seconds on a Cray Y/MP-48 supercomputer. It is unique in that subsequences (words) of biological interest can be weighted to improve the sensitivity and selectivity of a search over existing methods. We have determined the ability of d2 to detect biologically significant matches between a query and large datasets of DNA sequences while varying parameters such as word-length and window size. We have also determined the distribution of dissimilarity scores within eukaryotic and prokaryotic divisions of GenBank. We have optimized parameters of the d2 program using Cray hardware and present an analysis of the sensitivity and selectivity of the algorithm. A theoretical analysis of the expectation for scores is presented. This work demonstrates that d2 is a unique, sensitive, and selective method of rapid sequence comparison that can detect novel sequence relationships which remain undetected by alternate methodologies.

Algorithms↗