Search PubMed⌕ Search

PubMed · 16925838

GENCODE: producing a reference annotation for ENCODE.

Abstract

BACKGROUND: The GENCODE consortium was formed to identify and map all protein-coding genes within the ENCODE regions. This was achieved by a combination of initial manual annotation by the HAVANA team, experimental validation by the GENCODE consortium and a refinement of the annotation based on these experimental results. RESULTS: The GENCODE gene features are divided into eight different categories of which only the first two (known and novel coding sequence) are confidently predicted to be protein-coding genes. 5' rapid amplification of cDNA ends (RACE) and RT-PCR were used to experimentally verify the initial annotation. Of the 420 coding loci tested, 229 RACE products have been sequenced. They supported 5' extensions of 30 loci and new splice variants in 50 loci. In addition, 46 loci without evidence for a coding sequence were validated, consisting of 31 novel and 15 putative transcripts. We assessed the comprehensiveness of the GENCODE annotation by attempting to validate all the predicted exon boundaries outside the GENCODE annotation. Out of 1,215 tested in a subset of the ENCODE regions, 14 novel exon pairs were validated, only two of them in intergenic regions. CONCLUSION: In total, 487 loci, of which 434 are coding, have been annotated as part of the GENCODE reference set available from the UCSC browser. Comparison of GENCODE annotation with RefSeq and ENSEMBL show only 40% of GENCODE exons are contained within the two sets, which is a reflection of the high number of alternative splice forms with unique exons annotated. Over 50% of coding loci have been experimentally verified by 5' RACE for EGASP and the GENCODE collaboration is continuing to refine its annotation of 1% human genome with the aid of experimental validation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jennifer Harrow, France Denoeud, Adam Frankish, Alexandre Reymond, Chao-Kung Chen, Jacqueline Chrast, Julien Lagarde, James G R Gilbert, Roy Storey, David Swarbreck, Colette Rossier, Catherine Ucla, Tim Hubbard, Stylianos E Antonarakis, Roderic Guigo. 2006-08-07. GENCODE: producing a reference annotation for ENCODE.. https://doi.org/10.1186/gb-2006-7-s1-s4

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Mitotic karyotyping and FISH mapping of the gender-specific locus indicate an advanced XY system in Hippophae rhamnoides.

Hippophae rhamnoides ssp. turkestanica, a subdioecious plant inhabiting the cold desert of the Indian Himalaya, has gained immense recognition for its nutritional and medicinal values. In recent years, the plant species has proven to be a suitable system to understand the evolution of dioecy. Despite its biological significance, the cytogenetics of this dioecious plant is unclear due to various conflicting accounts of its X-Y chromosome system, particularly the length of Y-chromosome. In this study, we resolved these ambiguities through comprehensive cytogenetic analyses across diverse western Himalayan populations. Using morphometric analysis and fluorescence in situ hybridization (FISH) with a gender-specific marker (HRMSSR), we confirmed homomorphic XX chromosomes in females and heteromorphic sex-chromosomes in males with a notably smaller Y-chromosome. The investigation also revealed a predominant somatic chromosome number of 2n = 24, although minor deviations (2n = 18, 20, 22) appeared at the seed level. These findings highlight an evolutionarily advanced sex-chromosome system. This first detailed cytogenetic investigation of Himalayan Seabuckthorn provides critical insights into the chromosomal architecture, laying a crucial foundation for future evolutionary, genomic, and conservation studies in the species.

Chromosome Mapping↗

Methods for linkage disequilibrium mapping in crops.

Linkage disequilibrium (LD) mapping in plants detects and locates quantitative trait loci (QTL) by the strength of the correlation between a trait and a marker. It offers greater precision in QTL location than family-based linkage analysis and should therefore lead to more efficient marker-assisted selection, facilitate gene discovery and help to meet the challenge of connecting sequence diversity with heritable phenotypic differences. Unlike family-based linkage analysis, LD mapping does not require family or pedigree information and can be applied to a range of experimental and non-experimental populations. However, care must be taken during analysis to control for the increased rate of false positive results arising from population structure and variety interrelationships. In this review, we discuss how suitable the recently developed alternative methods of LD mapping are for crops.

Chromosome Mapping↗

An efficient method for producing an indexed, insertional-mutant library in rice.

Generation of an indexed, saturated, insertional-mutant library is an aid to understanding the functions of genes in an organism. However, 10 years of work by many investigators have not yet yielded such a library in rice. The major reason is that determining the chromosomal locations of a very large number of random insertion mutants by flanking sequence analysis is highly labor intensive, and therefore, libraries that do exist have not been indexed. We report here an efficient procedure to construct an indexed, region-specific, insertional-mutant library of rice. The procedure makes use of efficient long-PCR-based high-throughput indexing, coupled with a random but anchored population of Ds transposants. Long-PCR indexing allows rapid and simultaneous determination of the chromosomal locations of a large number of mutants that surround a particular anchor line, thus converting a random library into an indexed one. Such a library can be used directly, without the need to screen a large random library for a desired mutant plant.

Chromosome Mapping↗