Search PubMed⌕ Search

Biomedical subjects

France Denoeud

Publications and source records attributed to France Denoeud.

7 recordsLinked to original sources

EGASP: the human ENCODE Genome Annotation Assessment Project.

BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a 'reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40% to 50% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.

Alternative Splicing↗

GENCODE: producing a reference annotation for ENCODE.

BACKGROUND: The GENCODE consortium was formed to identify and map all protein-coding genes within the ENCODE regions. This was achieved by a combination of initial manual annotation by the HAVANA team, experimental validation by the GENCODE consortium and a refinement of the annotation based on these experimental results. RESULTS: The GENCODE gene features are divided into eight different categories of which only the first two (known and novel coding sequence) are confidently predicted to be protein-coding genes. 5' rapid amplification of cDNA ends (RACE) and RT-PCR were used to experimentally verify the initial annotation. Of the 420 coding loci tested, 229 RACE products have been sequenced. They supported 5' extensions of 30 loci and new splice variants in 50 loci. In addition, 46 loci without evidence for a coding sequence were validated, consisting of 31 novel and 15 putative transcripts. We assessed the comprehensiveness of the GENCODE annotation by attempting to validate all the predicted exon boundaries outside the GENCODE annotation. Out of 1,215 tested in a subset of the ENCODE regions, 14 novel exon pairs were validated, only two of them in intergenic regions. CONCLUSION: In total, 487 loci, of which 434 are coding, have been annotated as part of the GENCODE reference set available from the UCSC browser. Comparison of GENCODE annotation with RefSeq and ENSEMBL show only 40% of GENCODE exons are contained within the two sets, which is a reflection of the high number of alternative splice forms with unique exons annotated. Over 50% of coding loci have been experimentally verified by 5' RACE for EGASP and the GENCODE collaboration is continuing to refine its annotation of 1% human genome with the aid of experimental validation.

Chromosome Mapping↗

Evaluation and selection of tandem repeat loci for a Brucella MLVA typing assay.

BACKGROUND: The classification of Brucella into species and biovars relies on phenotypic characteristics and sometimes raises difficulties in the interpretation of the results due to an absence of standardization of the typing reagents. In addition, the resolution of this biotyping is moderate and requires the manipulation of the living agent. More efficient DNA-based methods are needed, and this work explores the suitability of multiple locus variable number tandem repeats analysis (MLVA) for both typing and species identification. RESULTS: Eighty tandem repeat loci predicted to be polymorphic by genome sequence analysis of three available Brucella genome sequences were tested for polymorphism by genotyping 21 Brucella strains (18 reference strains representing the six 'classical' species and all biovars as well as 3 marine mammal strains currently recognized as members of two new species). The MLVA data efficiently cluster the strains as expected according to their species and biovar. For practical use, a subset of 15 loci preserving this clustering was selected and applied to the typing of 236 isolates. Using this MLVA-15 assay, the clusters generated correspond to the classical biotyping scheme of Brucella spp. The 15 markers have been divided into two groups, one comprising 8 user-friendly minisatellite markers with a good species identification capability (panel 1) and another complementary group of 7 microsatellite markers with higher discriminatory power (panel 2). CONCLUSION: The MLVA-15 assay can be applied to large collections of Brucella strains with automated or manual procedures, and can be proposed as a complement, or even a substitute, of classical biotyping methods. This is facilitated by the fact that MLVA is based on non-infectious material (DNA) whereas the biotyping procedure itself requires the manipulation of the living agent. The data produced can be queried on a dedicated MLVA web service site.

Animals↗

Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource.

BACKGROUND: Polymorphic tandem repeat typing is a new generic technology which has been proved to be very efficient for bacterial pathogens such as B. anthracis, M. tuberculosis, P. aeruginosa, L. pneumophila, Y. pestis. The previously developed tandem repeats database takes advantage of the release of genome sequence data for a growing number of bacteria to facilitate the identification of tandem repeats. The development of an assay then requires the evaluation of tandem repeat polymorphism on well-selected sets of isolates. In the case of major human pathogens, such as S. aureus, more than one strain is being sequenced, so that tandem repeats most likely to be polymorphic can now be selected in silico based on genome sequence comparison. RESULTS: In addition to the previously described general Tandem Repeats Database, we have developed a tool to automatically identify tandem repeats of a different length in the genome sequence of two (or more) closely related bacterial strains. Genome comparisons are pre-computed. The results of the comparisons are parsed in a database, which can be conveniently queried over the internet according to criteria of practical value, including repeat unit length, predicted size difference, etc. Comparisons are available for 16 bacterial species, and the orthopox viruses, including the variola virus and three of its close neighbors. CONCLUSIONS: We are presenting an internet-based resource to help develop and perform tandem repeats based bacterial strain typing. The tools accessible at http://minisatellites.u-psud.fr now comprise four parts. The Tandem Repeats Database enables the identification of tandem repeats across entire genomes. The Strain Comparison Page identifies tandem repeats differing between different genome sequences from the same species. The "Blast in the Tandem Repeats Database" facilitates the search for a known tandem repeat and the prediction of amplification product sizes. The "Bacterial Genotyping Page" is a service for strain identification at the subspecies level.

Bacteria↗

Variable number of tandem repeats in Salmonella enterica subsp. enterica for typing purposes.

The genomic sequences of Salmonella enterica subsp. enterica strains CT18, Ty2 (serovar Typhi), and LT2 (serovar Typhimurium) were analyzed for potential variable number tandem repeats (VNTRs). A multiple-locus VNTR analysis (MLVA) of 99 strains of S. enterica supsp. enterica based on 10 VNTRs distinguished 52 genotypes and placed them into four groups. All strains tested were independent human isolates from France and did not reflect isolates from outbreak episodes. Of these 10 VNTRs, 7 showed variability within serovar Typhi, whereas 1 showed variability within serovar Typhimurium. Four VNTRs showed high Nei's diversity indices (DIs) of 0.81 to 0.87 within serovar Typhi (n = 27). Additionally, three of these more variable VNTRs showed DIs of 0.18 to 0.58 within serovar Paratyphi A (n = 10). The VNTR polymorphic site within multidrug-resistant (MDR) serovar Typhimurium isolates (n = 39; resistance to ampicillin, chloramphenicol, spectinomycin, sulfonamides, and tetracycline) showed a DI of 0.81. Cluster analysis not only identified three genetically distinct groups consistent with the present serovar classification of salmonellae (serovars Typhi, Paratyphi A, and Typhimurium) but also discriminated 25 subtypes (93%) within serovar Typhi isolates. The analysis discriminated only eight subtypes within serovar Typhimurium isolates resistant to ampicillin, chloramphenicol, spectinomycin, sulfonamides, and tetracycline, possibly reflecting the emergence in the mid-1990s of the DT104 phage type, which often displays such an MDR spectrum. Coupled with the ongoing improvements in automated procedures offered by capillary electrophoresis, use of these markers is proposed in further investigations of the potential of MLVA in outbreaks of salmonellosis, especially outbreaks of typhoid fever.

Alleles↗

Predicting human minisatellite polymorphism.

We seek to define sequence-based predictive criteria to identify polymorphic and hypermutable minisatellites in the human genome. Polymorphism of a representative pool of minisatellites, selected from human chromosomes 21 and 22, was experimentally measured by PCR typing in a population of unrelated individuals. Two predictive approaches were tested. One uses simple repeat characteristics (e.g., unit length, copy number, nucleotide bias) and a more complex measure, termed HistoryR, based on the presence of variant motifs in the tandem array. We find that HistoryR and percentage of GC are strongly correlated with polymorphism and, as predictive criteria, reduce by half the number of repeats to type while enriching the proportion with heterozygosity >/=0.5, from a background level of 43% to 59%. The second approach uses length differences between minisatellites in the two releases of the human genome sequence (from the public consortium and Celera). As a predictor, this similarly enriches the number of polymorphic minisatellites, but fails to identify an unexpectedly large number of these. Finally, typing of the highly polymorphic minisatellites in large families identified one new hypermutable minisatellite, located in a predicted coding sequence. This may represent the first coding human hypermutable minisatellite.

Chromosomes, Human, Pair 21↗

High resolution, on-line identification of strains from the Mycobacterium tuberculosis complex based on tandem repeat typing.

BACKGROUND: Currently available reference methods for the molecular epidemiology of the Mycobacterium tuberculosis complex either lack sensitivity or are still too tedious and slow for routine application. Recently, tandem repeat typing has emerged as a potential alternative. This report contributes to the development of tandem repeat typing for M. tuberculosis by summarising the existing data, developing additional markers, and setting up a freely accessible, fast, and easy to use, internet-based service for strain identification. RESULTS: A collection of 21 VNTRs incorporating 13 previously described loci and 8 newly evaluated markers was used to genotype 90 strains from the M. tuberculosis complex (M. tuberculosis (64 strains), M. bovis (9 strains including 4 BCG representatives), M. africanum (17 strains)). Eighty-four different genotypes are defined. Clustering analysis shows that the M. africanum strains fall into three main groups, one of which is closer to the M. tuberculosis strains, and an other one is closer to the M. bovis strains. The resulting data has been made freely accessible over the internet http://bacterial-genotyping.igmors.u-psud.fr/bnserver to allow direct strain identification queries. CONCLUSIONS: Tandem-repeat typing is a PCR-based assay which may prove to be a powerful complement to the existing epidemiological tools for the M. tuberculosis complex. The number of markers to type depends on the identification precision which is required, so that identification can be achieved quickly at low cost in terms of consumables, technical expertise and equipment.

DNA, Bacterial↗