Sequence interpretation. Making sense of the sequence.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Yeast artificial chromosomes (YACs) are a common tool for cloning eukaryotic DNA. The manner by which large pieces of foreign DNA are assimilated by yeast cells into a functional chromosome is poorly understood, as is the reason why some of them are stably maintained and some are not. We examined the replication of a stable YAC containing a 240-kb insert of DNA from the human T-cell receptor beta locus. The human insert contains multiple sites that serve as origins of replication. The activity of these origins appears to require the yeast ARS consensus sequence and, as with yeast origins, additional flanking sequences. In addition, the origins in the human insert exhibit a spacing, a range of activation efficiencies, and a variation in times of activation during S phase similar to those found for normal yeast chromosomes. We propose that an appropriate combination of replication origin density, activation times, and initiation efficiencies is necessary for the successful maintenance of YAC inserts.
BACKGROUND: The development of clinically significant disease in South Africa is associated with the vacuolating cytotoxin gene (vacA) s1 genotype but not with the presence of the cytotoxin associated gene cagA. cagA occurs in >95% of South African isolates and is a variable marker for the entire cag pathogenicity island (PAI). AIM: To characterise the cagPAI in South African isolates and to investigate if structural variants of this multigene locus were associated with variations in vacA status and clinical outcome. PATIENTS AND METHODS: We studied 109 Helicobacter pylori strains (36 from patients with peptic ulceration, 26 with gastric adenocarcinoma, and 47 with no pathology other than gastritis) for differences in selected genes of the cagPAI and alleles of vacA by polymerase chain reaction. RESULTS: All strains were cagA(+). Sixty five (60%) strains had an intact contiguous cagPAI; 78% of peptic ulcer isolates, 73% of gastric adenocarcinoma isolates, but only 40% of gastritis alone isolates (p< 0.01). The entire cagII region was undetectable in 23% of gastritis alone isolates but in only 8% of peptic ulceration isolates (p<0.05). The vacA signal sequence and mid region demonstrated a strong relationship between the virulence associated vacA s1 (p<0.005) and vacA m1 (p=0.05) alleles and an intact cagPAI. CONCLUSION: Although a complete cagPAI was a feature of most infected individuals, deletions in the 5' region of this genetic locus were associated with gastritis alone and with the non-cytotoxic s2/m2 vacA genotype.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Whole-genome sequencing is fundamental to understanding the genetic composition of an organism. Given the size and complexity of the soybean genome, an alternative approach is targeted random-gene sequencing, which provides an immediate and productive method of gene discovery. In this study, more than 120000 soybean expressed sequence tags (ESTs) generated from more than 50 cDNA libraries were evaluated. These ESTs coalesced into 16928 contigs and 17336 singletons. On average, each contig was composed of 6 ESTs and spanned 788 bases. The average sequence length submitted to dbEST was 414 bases. Using only those libraries generating more than 800 ESTs each and only those contigs with 10 or more ESTs each, correlated patterns of gene expression among libraries and genes were discerned. Two-dimensional qualitative representations of contig and library similarities were generated based on expression profiles. Genes with similar expression patterns and, potentially, similar functions were identified. These studies provide a rich source of publicly available gene sequences as well as valuable insight into the structure, function, and evolution of a model crop legume genome.
Surveying the soybean genome with 683 bacterial artificial chromosome (BAC) contiguous groups (contigs) anchored by restriction fragment length polymorphisms (RFLPs) enabled us to explore microsyntenic relationships among duplicated regions and also to examine the physical organization of hypomethylated (and presumably gene-rich) genomic regions. Numerous cases where nonhomologous RFLPs hybridized to common BAC clones indicated that RFLPs were physically clustered in soybean, apparently in less than 25% of the genome. By extension, we speculate that most of the genes are clustered in less than 275 M of the soybean genome. Approximately 40%-45% of this gene-rich portion is associated with the RFLP-anchored contigs described in this study. Similarities in genome organization among BAC contigs from duplicate genomic regions were also examined. Homoeologous BAC contigs often exhibited extensive microsynteny. Furthermore, paralogs recovered from duplicate contigs shared 86%-100% sequence identity.
The US Wheat Genome Project, funded by the National Science Foundation, developed the first large public Triticeae expressed sequence tag (EST) resource. Altogether, 116,272 ESTs were produced, comprising 100,674 5' ESTs and 15 598 3' ESTs. These ESTs were derived from 42 cDNA libraries, which were created from hexaploid bread wheat (Triticum aestivum L.) and its close relatives, including diploid wheat (T. monococcum L. and Aegilops speltoides L.), tetraploid wheat (T. turgidum L.), and rye (Secale cereale L.), using tissues collected from various stages of plant growth and development and under diverse regimes of abiotic and biotic stress treatments. ESTs were assembled into 18,876 contigs and 23,034 singletons, or 41,910 wheat unigenes. Over 90% of the contigs contained fewer than 10 EST members, implying that the ESTs represented a diverse selection of genes and that genes expressed at low and moderate to high levels were well sampled. Statistical methods were used to study the correlation of gene expression patterns, based on the ESTs clustered in the 1536 contigs that contained at least 10 5' EST members and thus representing the most abundant genes expressed in wheat. Analysis further identified genes in wheat that were significantly upregulated (p < 0.05) in tissues under various abiotic stresses when compared with control tissues. Though the function annotation cannot be assigned for many of these genes, it is likely that they play a role associated with the stress response. This study predicted the possible functionality for 4% of total wheat unigenes, which leaves the remaining 96% with their functional roles and expression patterns largely unknown. Nonetheless, the EST data generated in this project provide a diverse and rich source for gene discovery in wheat.
Its accessibility, unique evolutionary position, and recently assembled genome sequence have advanced the chicken to the forefront of comparative genomics and developmental biology research as a model organism. Several chicken expressed sequence tag (EST) projects have placed the chicken in 10th place for accrued ESTs among all organisms in GenBank. We have completed the single-pass 5'-end sequencing of 37,557 chicken cDNA clones from several single and multiple tissue cDNA libraries and have entered 35,407 EST sequences into GenBank. Our chicken EST sequences and those found in public databases (on July 1, 2004) provided a total of 517,727 public chicken ESTs and mRNAs. These sequences were used in the CAP3 assembly of a chicken gene index composed of 40,850 contigs and 79,192 unassembled singlets. The CAP3 contigs show a 96.7% match to the chicken genome sequence. The University of Delaware (UD) EST collection (43,928 clones) was assembled into 19,237 nonredundant sequences (13,495 contigs and 5,742 unassembled singlets). The UD collection contains 6,223 unique sequences that are not found in other public EST collections but show a 76% match to the chicken genome sequence. Our chicken contig and singlet sequences were annotated according to the highest BlastX and/or BlastN hits. The UD CAP3 contig assemblies and singlets are searchable by nucleotide sequence or key word (http://cogburn.dbi.udel.edu), and the cDNA clones are readily available for distribution from the chick EST website and clone repository (http://www.chickest.udel.edu). The present paper describes the construction and normalization of single and multiple tissue chicken cDNA libraries, high-throughput EST sequencing from these libraries, the CAP3 assembly of a chicken gene index from all public ESTs, and the identification of several nonredundant chicken gene sets for production of custom DNA microarrays.
A family of negative regulators of JAK signaling pathway referred to as suppressor of cytokines signaling (SOCS) or cytokine-inducible SH2 protein (CIS) has been recently identified. In order to find additional members of this family, we have used a consensus amino acid sequence contained in the well-conserved central SH2 domain to search DNA databases. We isolated cDNA coding for the human homologue of SOCS-5, referred to as CIS6. Northern blot analysis revealed CIS6 mRNA expression in various tissues such as heart, muscle, spleen, and thymus and in all myeloma cell lines examined. The gene was assigned to human chromosome bands 2p21 and 3p22 by in situ hybridization. CIS6 is structurally related to other members of the CIS family and therefore could act as a negative regulator of signal transduction.
We report on a small de novo interstitial deletion of the short arm of chromosome 20, 46,XY,del(20)(p12.3p13), in a young boy with hypotonia, moderate development delay, mild facial dysmorphism and severe growth failure. This patient did not show major features of Alagille-Watson Syndrome (AWS) which are common in more proximal 20p deletions. Standard and high resolution chromosome banding analysis revealed an apparent terminal deletion. Nevertheless, using chromosomal fluorescent in situ hybridization (FISH) and molecular analysis with polymorphic markers, we demonstrated that the abnormal chromosome resulted from a de novo interstitial deletion of paternal origin spanning from D20S842 to D20S900 and covering approximately 6 Mb. These findings indicate that a karyotype can lead to insufficient characterization of an apparently terminal deletion, and that one or a few genes in 20p13-->p12.3 bands are important for normal growth.
Human chromosome 11p15.3 is associated with chromosome aberrations in the Beckwith Wiedemann Syndrome and implicated in the pathogenesis of different tumor types including lung cancer and leukemias. To date, only single tumor-relevant genes with linkage to this region (e.g. LMO1) have been found suggesting that this region may harbor additional potential disease associated genes. Although this genomic area has been studied for years, the exact order of genes/chromosome markers between D11S572 and the WEE1 gene locus remained unclear. Using the FISH technique and PAC clones of the flanking markers we determined the order of the genomic markers. Based on these clones we established a PAC contig of the respective region. To analyse the chromosome area in detail the synteny of the orthologous region on distal mouse chromosome 7 was determined and a corresponding mouse clone contig established, proving the conserved order of the genes and markers in both species: "cen-WEE1-D11S2043-ZNF143-RANBP7-CEGF1- ST5-D11S932-LMO1-D11S572-TUB-tel", with inverted order of the murine genes with respect to the telomere/centromere orientation. The region covered by these contigs comprises roughly 1.6 MB in human as well as in mouse. The genomic sequence of the two subregions (around WEE1 and LMO1) in both species was determined using a shotgun sequencing strategy. Comparative sequence analysis techniques demonstrate that the content of repetitive elements seems to decline from centromere to telomere (52.6% to 34.5%) in human and in the corresponding murine region from telomere to centromere (41.87% to 27.82%). Genomic organisation of the regions around WEE1 and LMO1 was conserved, although the length of gene regions varied between the species in an unpredictable ratio. CpG islands were found conserved in putative promoter regions of the known genes but also in regions which so far have not been described as harboring expressed sequences.
Comparative genomics is a superior way to identify phylogenetically conserved features like genes or regions involved in gene regulation. The comparison of extended orthologous chromosomal regions should also reveal other characteristic traits essential for chromosome or gene function. In the present study we have sequenced and compared a region of conserved synteny from human chromosome 11p15.3 and mouse chromosome 7. In human, this region is known to contain several genes involved in the development of various disorders like Beckwith-Wiedemann overgrowth syndrome and other tumor diseases. Furthermore, in the neighboring chromosome region 11p15.5 extensive imprinting of genes has been reported which might extend to region 11p15.3. The analysis of approximately 730 kb in human and 620 kb in mouse led to the identification of eleven genes. All putative genes found in the mouse DNA were also present in the same order and orientation in the human chromosome. However, in the human DNA one putative gene of unknown function could be identified which is not present in the orthologous position of the mouse chromosome. The sequence similarity between human and mouse is higher in transcribed and exon regions than in non-transcribed segments. Dot plot analysis, however, reveals a surprisingly well-conserved sequence similarity over the entire analyzed region. In particular, the positions of CpG islands, short regions of very high GC content in the 5' region of putative genes, are similar in human and mouse. With respect to base composition, two distinct segments of significantly different GC content exist as well in human as in the mouse. With a GC content of 45% the one segment would correspond to "isochore H1" and the other segment (39% GC in human, 40% GC in mouse) to "isochore L1/L2". The gene density (one gene per 66 kb) is slightly higher than the average calculated for the complete human genome (one gene per 90 kb). The comparison of the number and distribution of repetitive elements shows that the proportion of human DNA made up by interspersed repeats (43.8%) is significantly higher than in the corresponding mouse DNA (30.1%). This partly explains why the human DNA is longer between the landmark genes used to define the orthologous positions in human and mouse.
Compared to other regions on the human Y chromosome, the genomic segment encompassing the functionally defined AZFa locus has undergone higher X-Y sequence divergence, which is detectable by fluorescence in-situ hybridisation. This allows an evolutionary definition of an interval enclosing AZFa with a size of about 1.1 Mb. The region includes the genes USP9Y, DBY and UTY and is limited by evolutionary breakpoints within the PAC clones 41L06 and 46M11. These breakpoints restrict an area of possible male specific evolution that may have resulted in the acquisition of male specific functions, including a role in spermatogenesis.
Mutation of genetic material is a necessary component of evolutionary change. There is evidence for both intragenome and intergenome heterogeneity in terms of mutation frequencies. Reported comparisons of DNA sequence differences between human and chimpanzee (Pan troglodytes) suggest that human chromosome 21 may exhibit mutational hypervariability relative to the other autosomes. In the present study, further evidence is provided for such hypervariability based on large-scale analyses of amino acid composition of (translated) human genes and pseudogenes. A comparison of the variation in the above cases (i.e., DNA sequence differences and amino acid composition differences) yields similar ratios (1.2-1.4) for chromosome 21 relative to the other autosomes, e.g., human chromosome 22 - an autosome that is more typical in this respect and is of similar size to 21. Human chromosome 21 is also presented in this study as being atypical in terms of reported associations between mutation rates and GC content or CpG dinucleotides. In terms of GC distribution patterns, a comparison of NT_011512 and NT_011520 contigs revealed a lower heterogeneity for human chromosome 21 relative to 22. Possible hypermutability of chromosome 21 is further discussed in the context of GC patterns, reported long interspersed nuclear element content (LINE1s), and the implications of these parameters for chromatin structure.
Chromosome 11q deletions are frequently observed in chronic lymphocytic leukemia (CLL) in association with progressive disease and a poor prognosis. A minimal region of deletion has been assigned to 11q22-q23. Trinucleotide repeats have been associated with anticipation in disease, and evidence of anticipation has been observed in various malignancies including CLL. Loss of heterozygosity at 11q22-23 is common in a wide range of cancers, suggesting this is an unstable area prone to chromosome breakage. The location of 8 CCG-trinucleotide repeats on 11q was determined by Southern blot analysis of a 40-Mb YAC and PAC contig spanning 11q22-qter. Deletion breakpoints in CLL are found to co-localize at specific sites on 11q where CCG repeats are located. In addition, a CCG repeat has been identified within the minimal region of deletion. Specific alleles of this repeat are associated with worse prognosis. Folate-sensitive fragile sites are regions of late replication and are characterized by CCG repeats. The mechanism for chromosome deletion at 11q could be explained by a delay in replication. Described here is an association between CCG repeats and chromosome loss suggesting that in vivo "fragile sites" exist on 11q and that the instability of CCG repeats may play an important role in the pathogenesis of CLL.
BACKGROUND: Identifying reliable oligonucleotide sequences for use in microarray experiments is a complex process. Two key issues are the accuracy of the input sequences and the specificity of the oligonucleotide sequences. RESULTS: We provide a suite of Perl scripts that facilitates the search for gene-specific oligonucleotides for microarray experiments. Genes of interest are first identified in the form of UniGene clusters. The sequences of these clusters were extracted and assembled into contigs to increase their accuracy. The 3' untranslated region (3'UTR) of the contig was parsed. Then, multiple 50mer oligonucleotide sequences with similar melting temperature were obtained from each 3'UTR. These sequences were analyzed for gene specificity. Five Cy3-labeled cDNAs were used to empirically verify the specificity of a set of 1814 50mers. CONCLUSION: Oliz can be used to select oligonucleotide sequences for microarrays. Oliz is freely available for academic users at http://www.utmem.edu/pharmacology/otherlinks/oliz.html
BACKGROUND: Assignment of function to new molecular sequence data is an essential step in genomics projects. The usual process involves similarity searches of a given sequence against one or more databases, an arduous process for large datasets. RESULTS: We present AutoFACT, a fully automated and customizable annotation tool that assigns biologically informative functions to a sequence. Key features of this tool are that it (1) analyzes nucleotide and protein sequence data; (2) determines the most informative functional description by combining multiple BLAST reports from several user-selected databases; (3) assigns putative metabolic pathways, functional classes, enzyme classes, GeneOntology terms and locus names; and (4) generates output in HTML, text and GFF formats for the user's convenience. We have compared AutoFACT to four well-established annotation pipelines. The error rate of functional annotation is estimated to be only between 1-2%. Comparison of AutoFACT to the traditional top-BLAST-hit annotation method shows that our procedure increases the number of functionally informative annotations by approximately 50%. CONCLUSION: AutoFACT will serve as a useful annotation tool for smaller sequencing groups lacking dedicated bioinformatics staff. It is implemented in PERL and runs on LINUX/UNIX platforms. AutoFACT is available at http://megasun.bch.umontreal.ca/Software/AutoFACT.htm.