Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Apparent homology of expressed genes from wood-forming tissues of loblolly pine (Pinus taeda L.) with Arabidopsis thaliana.

Pinus taeda L. (loblolly pine) and Arabidopsis thaliana differ greatly in form, ecological niche, evolutionary history, and genome size. Arabidopsis is a small, herbaceous, annual dicotyledon, whereas pines are large, long-lived, coniferous forest trees. Such diverse plants might be expected to differ in a large number of functional genes. We have obtained and analyzed 59,797 expressed sequence tags (ESTs) from wood-forming tissues of loblolly pine and compared them to the gene sequences inferred from the complete sequence of the Arabidopsis genome. Approximately 50% of pine ESTs have no apparent homologs in Arabidopsis or any other angiosperm in public databases. When evaluated by using contigs containing long, high-quality sequences, we find a higher level of apparent homology between the inferred genes of these two species. For those contigs 1,100 bp or longer, approximately 90% have an apparent Arabidopsis homolog (E value < 10-10). Pines and Arabidopsis last shared a common ancestor approximately 300 million years ago. Few genes would be expected to retain high sequence similarity for this time if they did not have essential functions. These observations suggest substantial conservation of gene sequence in seed plants.

3' Untranslated Regions↗

Comparative sequencing of human and chimpanzee MHC class I regions unveils insertions/deletions as the major path to genomic divergence.

Despite their high degree of genomic similarity, reminiscent of their relatively recent separation from each other ( approximately 6 million years ago), the molecular basis of traits unique to humans vs. their closest relative, the chimpanzee, is largely unknown. This report describes a large-scale single-contig comparison between human and chimpanzee genomes via the sequence analysis of almost one-half of the immunologically critical MHC. This 1,750,601-bp stretch of DNA, which encompasses the entire class I along with the telomeric part of the MHC class III regions, corresponds to an orthologous 1,870,955 bp of the human HLA region. Sequence analysis confirms the existence of a high degree of sequence similarity between the two species. However, and importantly, this 98.6% sequence identity drops to only 86.7% taking into account the multiple insertions/deletions (indels) dispersed throughout the region. This is functionally exemplified by a large deletion of 95 kb between the virtual locations of human MICA and MICB genes, which results in a single hybrid chimpanzee MIC gene, in a segment of the MHC genetically linked to species-specific handling of several viral infections (HIV/SIV, hepatitis B and C) as well as susceptibility to various autoimmune diseases. Finally, if generalized, these data suggest that evolution may have used the mechanistically more drastic indels instead of the more subtle single-nucleotide substitutions for shaping the recently emerged primate species.

Animals↗

An Eulerian path approach to DNA fragment assembly.

For the last 20 years, fragment assembly in DNA sequencing followed the "overlap-layout-consensus" paradigm that is used in all currently available assembly tools. Although this approach proved useful in assembling clones, it faces difficulties in genomic shotgun assembly. We abandon the classical "overlap-layout-consensus" approach in favor of a new euler algorithm that, for the first time, resolves the 20-year-old "repeat problem" in fragment assembly. Our main result is the reduction of the fragment assembly to a variation of the classical Eulerian path problem that allows one to generate accurate solutions of large-scale sequencing problems. euler, in contrast to the celera assembler, does not mask such repeats but uses them instead as a powerful fragment assembly tool.

Algorithms↗

A novel amplicon at 8p22-23 results in overexpression of cathepsin B in esophageal adenocarcinoma.

Cathepsin B (CTSB) is overexpressed in tumors of the lung, prostate, colon, breast, and stomach. However, evidence of primary genomic alterations in the CTSB gene during tumor initiation or progression has been lacking. We have found a novel amplicon at 8p22-23 that results in CTSB overexpression in esophageal adenocarcinoma. Amplified genomic NotI-HinfI fragments were identified by two-dimensional DNA electrophoresis. Two amplified fragments (D4 and D5) were cloned and yielded unique sequences. Using bacterial artificial chromosome clones containing either D4 or D5, fluorescent in situ hybridization defined a single region of amplification involving chromosome bands 8p22-23. We investigated the candidate cancer-related gene CTSB, and potential coamplified genes from this region including farnesyl-diphosphate farnesyltransferase (FDFT1), arylamine N-acetyltransferase (NAT-1), lipoprotein lipase (LPL), and an uncharacterized expressed sequence tag (D8S503). Southern blot analysis of 66 esophageal adenocarcinomas demonstrated only CTSB and FDFT1 were consistently amplified in eight (12.1%) of the tumors. Neither NAT-1 nor LPL were amplified. Northern blot analysis showed overexpression of CTSB and FDFT1 mRNA in all six of the amplified esophageal adenocarcinomas analyzed. CTSB mRNA overexpression also was present in two of six nonamplified tumors analyzed. However, FDFT1 mRNA overexpression without amplification was not observed. Western blot analysis confirmed CTSB protein overexpression in tumor specimens with CTSB mRNA overexpression compared with either normal controls or tumors without mRNA overexpression. Abundant extracellular expression of CTSB protein was found in 29 of 40 (72. 5%) of esophageal adenocarcinoma specimens by using immunohistochemical analysis. The finding of an amplicon at 8p22-23 resulting in CTSB gene amplification and overexpression supports an important role for CTSB in esophageal adenocarcinoma and possibly in other tumors.

Adenocarcinoma↗

The Lateral suppressor (Ls) gene of tomato encodes a new member of the VHIID protein family.

The ability of the shoot apical meristem to multiply and distribute its meristematic potential through the formation of axillary meristems is essential for the diversity of forms and growth habits of higher plants. In the lateral suppressor mutant of tomato the initiation of axillary meristems is prevented, thus offering the unique opportunity to study the molecular mechanisms underlying this important function of the shoot apical meristem. We report here the isolation of the Lateral suppressor gene by positional cloning and show that the mutant phenotype is caused by a complete loss of function of a new member of the VHIID family of plant regulatory proteins.

Amino Acid Sequence↗

Characterization of a 190-kilobase pair domain of human type I hair keratin genes.

Polymerase chain reaction-based screening of an arrayed human P1 artificial chromosome (PAC) library using primer pairs specific for the human type I hair keratins hHa3-II or hHa6, led to the isolation of two PAC clones, which covered 190 kilobase pairs (kbp) of genomic DNA and contained nine human type I hair keratin genes, one transcribed hair keratin pseudogene, as well as one orphan exon. The hair keratin genes are 4-7 kbp in size, exhibit intergenic distances of 5-8 kbp, and display the same direction of transcription. With one exception, all hair keratin genes are organized into 7 exons and 6 positionally conserved introns. On the basis of sequence homologies, the genes can be grouped into three subclusters of tandemly arranged genes. One subcluster harbors the highly related genes hHa1, hHa3-I, hHa3-II, and hHa4. A second subcluster of highly related genes comprises the novel genes hHa7 and hHa8, as well as pseudogene PsihHaA, while the structurally less related genes hHa6, hHa5, and hHa2 are constituents of the third subcluster. As shown by reverse transcription-polymerase chain reaction, all hair keratin genes, including the pseudogene, are expressed in the human hair follicle. The transcribed pseudogene PsihHaA contains a premature stop codon in exon 4 and exhibits aberrant pre-mRNA splicing. Evolutionary tree construction reveals an early divergence of hair keratin genes from cytokeratin genes, followed by the segregation of the genes into the three subclusters. We suspect that the 190-kbp domain contains the entire complement of human type I hair keratin genes.

Amino Acid Sequence↗

Structural organization and regulation of the small proline-rich family of cornified envelope precursors suggest a role in adaptive barrier function.

The protective barrier provided by stratified squamous epithelia relies on the cornified cell envelope (CE), a structure synthesized at late stages of keratinocyte differentiation. It is composed of structural proteins, including involucrin, loricrin, and the small proline-rich (SPRR) proteins, all encoded by genes localized at human chromosome 1q21. The genetic characterization of the SPRR locus reveals that the various members of this multigene family can be classified into two distinct groups with separate evolutionary histories. Whereas group 1 genes have diverged in protein structure and are composed of three different classes (SPRR1 (2x), SPRR3, and SPRR4), an active process of gene conversion has counteracted diversification of the protein sequences of group 2 genes (SPRR2 class, seven genes). Contrasting with this homogenization process, all individual members of the SPRR gene family show specific in vivo and in vitro expression patterns and react selectively to UV irradiation. Apparently, creation of regulatory rather than structural diversity has been the driving force behind the evolution of the SPRR gene family. Differential regulation of highly homologous genes underlines the importance of SPRR protein dosage in providing optimal barrier function to different epithelia, while allowing adaptation to diverse external insults.

Amino Acid Sequence↗

Characterization of a cluster of human high/ultrahigh sulfur keratin-associated protein genes embedded in the type I keratin gene domain on chromosome 17q12-21.

Low stringency screening of a human P1 artificial chromosome library using a human hair keratin-associated protein (hKAP1.1A) gene probe resulted in the isolation of six P1 artificial chromosome clones. End sequencing and EMBO/GenBank(TM) data base analysis showed these clones to be contained in four previously sequenced human bacterial artificial chromosome clones present on chromosome 17q12-21 and arrayed into two large contigs of 290 and 225 kilobase pairs (kb) in size. A fifth, partially sequenced human bacterial artificial chromosome clone data base sequence overlapped and closed both of these contigs. One end of this 600-kb cluster harbored six gene loci for previously described human type I hair keratin genes. The other end of this cluster contained the human type I cytokeratin K20 and K12 gene loci. The center of the cluster, starting 35 kb downstream of the hHa3-I hair keratin gene, contained 37 genes for high/ultrahigh sulfur hair keratin-associated proteins (KAPs), which could be divided into a total of 7 KAP multigene families based on amino acid homology comparisons with previously identified sheep, mouse, and rabbit KAPs. To date, 26 human KAP cDNA clones have been isolated through screening of an arrayed human scalp cDNA library by means of specific 3'-noncoding region polymerase chain reaction probes derived from the identified KAP gene sequences. This screening also yielded four additional cDNA sequences whose genes were not present on this gene cluster but belonged to specific KAP gene families present on this contig. Hair follicle in situ hybridization data for single members of five different KAP multigene families all showed localization of the respective mRNAs to the upper cortex of the hair shaft.

Amino Acid Sequence↗

Serglycin is essential for maturation of mast cell secretory granule.

To address the biological function of the scarcely studied intracellular proteoglycans, we targeted the gene for serglycin (SG), the only known committed intracellular proteoglycan. SG-/- mice developed normally and were fertile, but their mast cells (MCs) were severely affected. In peritoneum there was a complete absence of normal granulated MCs. Furthermore, peritoneal cells and ear tissue from SG-/- animals were devoid of the various MC-specific proteases. However, mRNA for the proteases was present in SG+/+, SG+/-, and SG-/- tissues, indicating that SG is essential for the storage, but not expression, of the MC proteases. Experiments, in which the differentiation of bone marrow stem cells into mature MCs was followed, showed that secretory granule maturation was compromised in SG-/- cells. Moreover, SG+/+ and SG+/- cells, but not SG-/- cells, synthesized proteoglycans of high anionic charge density. Taken together, we demonstrate a key role for SG proteoglycan in MC function.

Animals↗

Chromosome breakage in the Prader-Willi and Angelman syndromes involves recombination between large, transcribed repeats at proximal and distal breakpoints.

Prader-Willi syndrome (PWS) and Angelman syndrome (AS) are distinct neurobehavioral disorders that most often arise from a 4-Mb deletion of chromosome 15q11-q13 during paternal or maternal gametogenesis, respectively. At a de novo frequency of approximately.67-1/10,000 births, these deletions represent a common structural chromosome change in the human genome. To elucidate the mechanism underlying these events, we characterized the regions that contain two proximal breakpoint clusters and a distal cluster. Novel DNA sequences potentially associated with the breakpoints were positionally cloned from YACs within or near these regions. Analyses of rodent-human somatic-cell hybrids, YAC contigs, and FISH of normal or rearranged chromosomes 15 identified duplicated sequences (the END repeats) at or near the breakpoints. The END-repeat units are derived from large genomic duplications of a novel gene (HERC2), many copies of which are transcriptionally active in germline tissues. One of five PWS/AS patients analyzed to date has an identifiable, rearranged HERC2 transcript derived from the deletion event. We postulate that the END repeats flanking 15q11-q13 mediate homologous recombination resulting in deletion. Furthermore, we propose that active transcription of these repeats in male and female germ cells may facilitate the homologous recombination process.

Angelman Syndrome↗

Disruption of a novel imprinted zinc-finger gene, ZNF215, in Beckwith-Wiedemann syndrome.

The genetics of Beckwith-Wiedemann syndrome (BWS) is complex and is thought to involve multiple genes. It is known that three regions on chromosome 11p15 (BWSCR1, BWSCR2, and BWSCR3) may play a role in the development of BWS. BWSCR2 is defined by two BWS breakpoints. Here we describe the cloning and sequence analysis of 73 kb containing BWSCR2. Within this region, we detected a novel zinc-finger gene, ZNF215. We show that two of its five alternatively spliced transcripts are disrupted by both BWSCR2 breakpoints. Parts of the 3' end of these splice forms are transcribed from the antisense strand of a second zinc-finger gene, ZNF214. We show that ZNF215 is imprinted in a tissue-specific manner.

Alleles↗

Linkage of tuberculosis to chromosome 2q35 loci, including NRAMP1, in a large aboriginal Canadian family.

An epidemic of tuberculosis occurred in a community of Aboriginal Canadians during the period 1987-89. Genetic and epidemiologic data were collected on an extended family from this community, and the evidence for linkage to NRAMP1, a candidate gene for susceptibility to mycobacterial diseases, was assessed. Individuals were grouped into risk (liability) classes based on vaccination, age, previous disease, and tuberculin skin-test results. Under the assumption of a dominant mode of inheritance and a relative risk of 10, which is associated with the high-risk genotypes, a maximum LOD score of 3.81 was observed for linkage between a tuberculosis-susceptibility locus and D2S424, which is located just distal to NRAMP1, in chromosome region 2q35. Significant linkage was also observed between a tuberculosis-susceptibility locus and a haplotype of 10 NRAMP1 intragenic variants. No linkage to the major histocompatibility-complex region on chromosome 6p was observed, despite distortion of transmission from one member of the oldest couple to their affected offspring. The ability to assign individuals to risk classes was crucial to the success of this study.

Alleles↗

Olfactory receptor-gene clusters, genomic-inversion polymorphisms, and common chromosome rearrangements.

The olfactory receptor (OR)-gene superfamily is the largest in the mammalian genome. Several of the human OR genes appear in clusters with > or = 10 members located on almost all human chromosomes, and some chromosomes contain more than one cluster. We demonstrate, by experimental and in silico data, that unequal crossovers between two OR gene clusters in 8p are responsible for the formation of three recurrent chromosome macrorearrangements and a submicroscopic inversion polymorphism. The first two macrorearrangements are the inverted duplication of 8p, inv dup(8p), which is associated with a distinct phenotype, and a supernumerary marker chromosome, +der(8)(8p23.1pter), which is also a recurrent rearrangement and is associated with minor anomalies. We demonstrate that it is the reciprocal of the inv dup(8p). The third macrorearrangment is a recurrent 8p23 interstitial deletion associated with heart defect. Since inv dup(8p)s originate consistently in maternal meiosis, we investigated the maternal chromosomes 8 in eight mothers of subjects with inv dup(8p) and in the mother of one subject with +der(8), by means of probes included between the two 8p-OR gene clusters. All the mothers were heterozygous for an 8p submicroscopic inversion that was delimited by the 8p-OR gene clusters and was present, in heterozygous state, in 26% of a population of European descent. Thus, inversion heterozygosity may cause susceptibility to unequal recombination, leading to the formation of the inv dup(8p) or to its reciprocal product, the +der(8p). After the Yp inversion polymorphism, which is the preferential background for the PRKX/PRKY translocation in XX males and XY females, the OR-8p inversion is the second genomic polymorphism that confers susceptibility to the formation of common chromosome rearrangements. Accordingly, it may be possible to develop a profile of the individual risk of having progeny with chromosome rearrangements.

Chromosome Breakage↗

Mutations in a novel gene with transmembrane domains underlie Usher syndrome type 3.

Usher syndrome type 3 (USH3) is an autosomal recessive disorder characterized by progressive hearing loss, severe retinal degeneration, and variably present vestibular dysfunction, assigned to 3q21-q25. Here, we report on the positional cloning of the USH3 gene. By haplotype and linkage-disequilibrium analyses in Finnish carriers of a putative founder mutation, the critical region was narrowed to 250 kb, of which we sequenced, assembled, and annotated 207 kb. Two novel genes-NOPAR and UCRP-and one previously identified gene-H963-were excluded as USH3, on the basis of mutational analysis. USH3, the candidate gene that we identified, encodes a 120-amino-acid protein. Fifty-two Finnish patients were homozygous for a termination mutation, Y100X; patients in two Finnish families were compound heterozygous for Y100X and for a missense mutation, M44K, whereas patients in an Italian family were homozygous for a 3-bp deletion leading to an amino acid deletion and substitution. USH3 has two predicted transmembrane domains, and it shows no homology to known genes. As revealed by northern blotting and reverse-transcriptase PCR, it is expressed in many tissues, including the retina.

Abnormalities, Multiple↗

A tool for analyzing mate pairs in assemblies (TAMPA).

The current generation of genome assembly programs uses distance and orientation relationships of paired end reads of clones (mate pairs) to order and orient contigs. Mate pair data can also be used to evaluate and compare assemblies after the fact. Earlier work employed a simple heuristic to detect assembly problems by scanning across an assembly to locate peak concentrations of unsatisfied mate pairs. TAMPA is a novel, computational geometry-based approach to detecting assembly breakpoints by exploiting constraints that mate pairs impose on each other. The method can be used to improve assemblies and determine which of two assemblies is correct in the case of sequence disagreement. Results from several human genome assemblies are presented.

Algorithms↗

Searching the expressed sequence tag (EST) databases: panning for genes.

The genomes of living organisms contain many elements, including genes coding for proteins. The portions of the genes expressed as mature mRNA, collectively known as the transcriptome, represent only a small part of the genome. The expressed sequence tag (EST) databases contain an increasingly large part of the transcriptome of many species. For this reason, these databases are probably the most abundant source of new coding sequences available today. However, the raw data deposited in the EST databases are to a large extent unorganised, unannotated, redundant and of relatively low quality. This paper reviews some of the characteristics of the EST data, and the methods that can be used to find novel protein sequences within them. It also documents a collection of databases, software and web sites that can be useful to biologists interested in mining the EST databases over the Internet, or in establishing a local environment for such analyses.

Algorithms↗

Evaluation of gene prediction software using a genomic data set: application to Arabidopsis thaliana sequences.

MOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: Pierre.Rouze@gengenp.rug.ac.be.

Alternative Splicing↗