Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

RabbitSketch: a high-performance sketching library for genome analysis.

SUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest.

Software↗

Heteroduplex resolution using T7 endonuclease I in microbial community analyses.

Microbial community analyses using molecular techniques, such as PCR followed by genomic library construction, have been helpful in better understanding microbial communities. This is especially critical in ecological systems where most of the microbes present cannot be cultured using traditional techniques. Unfortunately, there are problems associated with the use of such molecular techniques for the analysis of microbial community structure, primarily from the frequent formation of PCR artifacts. Multitemplate PCR is often subject to errors such as heteroduplex formation that is generated during the amplification of a particular gene from a mixed community of DNA. Based on work in this laboratory, heteroduplexes may be resolved before carrying out genomic library construction by including a digestion step with T7 endonuclease I. Here, the 18S rDNA gene of fungi was amplified from soil community DNA and digested with T7 endonuclease I to resolve any heteroduplexes present in the PCR product before cloning. These samples were compared with replicates that did not receive the T7 endonuclease I treatment. Digestion of the amplified community 18S rDNA with 10 U T7 endonuclease I/microgram DNA prior to cloning eliminated heteroduplexes, leaving only the desired clones. Without the T7 endonuclease I treatment, heteroduplexes were produced in approximately 10% of the recombinants screened. The addition of this step may eliminate heteroduplexes from PCR products and ensure that subsequent genomic library construction is not compromised.

Artifacts↗

The complete structure of human class IV alcohol dehydrogenase (retinol dehydrogenase) determined from the ADH7 gene.

A novel human alcohol dehydrogenase (ADH) gene called ADH7 has been characterized and determined to encode class IV ADH, an ADH isozyme which is very active as a retinol dehydrogenase. A nearly full-length cDNA for ADH7 was isolated from a human stomach cDNA library, and a 5' genomic clone containing exons 1 and 2 was isolated from a human genomic library. DNA sequence analysis of the cDNA and genomic clones revealed the complete coding region of the gene and the deduced full-length amino acid sequence of human class IV ADH composed of 373 amino acids following the initiator methionine. The class IV identity of the sequence was confirmed by agreement with previously determined sequences for several human stomach class IV ADH peptides. Alignment of the full-length predicted amino acid sequence of human class IV ADH with the full-length sequences of the other four known human ADH classes revealed sequence identities of 69% (class I), 59% (class II), 61% (class III), and 60% (class V). The higher sequence identity shared with human class I ADH suggests that the genes for ADH classes I and IV may have diverged from a common ancestor after the separation of the other classes, and may still share common physiological functions. Discussed is the possibility that one of these functions is retinol oxidation for the synthesis of retinoic acid, a hormone important for cellular differentiation.

Alcohol Oxidoreductases↗

The TIGR Maize Database.

Maize is a staple crop of the grass family and also an excellent model for plant genetics. Owing to the large size and repetitiveness of its genome, we previously investigated two approaches to accelerate gene discovery and genome analysis in maize: methylation filtration and high C(0)t selection. These techniques allow the construction of gene-enriched genomic libraries by minimizing repeat sequences due to either their methylation status or their copy number, yielding a 7-fold enrichment in genic sequences relative to a random genomic library. Approximately 900,000 gene-enriched reads from maize were generated and clustered into Assembled Zea mays (AZM) sequences. Here we report the current AZM release, which consists of approximately 298 Mb representing 243,807 sequence assemblies and singletons. In order to provide a repository of publicly available maize genomic sequences, we have created the TIGR Maize Database (http://maize.tigr.org). In this resource, we have assembled and annotated the AZMs and used available sequenced markers to anchor AZMs to maize chromosomes. We have constructed a maize repeat database and generated draft sequence assemblies of 287 maize bacterial artificial chromosome (BAC) clone sequences, which we annotated along with 172 additional publicly available BAC clones. All sequences, assemblies and annotations are available at the project website via web interfaces and FTP downloads.

Chromosome Mapping↗

A DNA probe to distinguish the species Anopheles quadriannulatus from other species of the Anopheles gambiae complex.

DNA probes used previously to distinguish the species Anopheles gambiae sensu stricto, An.arabiensis, An.melas and An.merus were tested against An.quadriannulatus. Using these DNA probes, An.gambiae s.s. and An.quadriannulatus were indistinguishable. A genomic library was constructed for An.quadriannulatus. Differential screening of this genomic library with An.gambiae s.s. and An.quadriannulatus genomic DNAs identified a species-specific, repeated DNA sequence. When used as a hybridization probe, this DNA sequence clearly distinguished An.gambiae s.s. from An.quadriannulatus. A simplified protocol for the use of DNA probes is described which may be used to identify material squashed directly on to nitrocellulose filter paper.

Animals↗

The Neurospora crassa genome: cosmid libraries sorted by chromosome.

A Neurospora crassa cosmid library of 12,000 clones (at least nine genome equivalents) has been created using an improved cosmid vector pLorist6Xh, which contains a bacteriophage lambda origin of replication for low-copy-number replication in bacteria and the hygromycin phosphotransferase marker for direct selection in fungi. The electrophoretic karyotype of the seven chromosomes comprising the 42.9-Mb N. crassa genome was resolved using two translocation strains. Using gel-purified chromosomal DNAs as probes against the new cosmid library and the commonly used medium-copy-number pMOcosX N. crassa cosmid library in two independent screenings, the cosmids were assigned to chromosomes. Assignments of cosmids to linkage groups on the basis of the genetic map vs. the electrophoretic karyotype are 93 +/- 3% concordant. The size of each chromosome-specific subcollection of cosmids was found to be linearly proportional to the size of the particular chromosome. Sequencing of an entire cosmid containing the qa gene cluster indicated a gene density of 1 gene per 4 kbp; by extrapolation, 11,000 genes would be expected to be present in the N. crassa genome. By hybridizing 79 nonoverlapping cosmids with an average insert size of 34 kbp against cDNA arrays, the density of previously characterized expressed sequence tags (ESTs) was found to be slightly <1 per cosmid (i.e., 1 per 40 kbp), and most cosmids, on average, contained an identified N. crassa gene sequence as a starting point for gene identification.

Bacteriophage lambda↗

Construction of a human Ig combinatorial library from genomic V segments and synthetic CDR3 fragments.

A naive combinatorial Ig library was constructed from semi-synthetic V genes consisting of human genomic V segments and synthetic CDR3 fragments. VH and V kappa segments were amplified from human genomic DNA by polymerase chain reaction using V subgroup-specific primers. The amplified VH and V kappa segments were combined with synthetic oligonucleotides containing a J region and CDR3 with amino acid sequence variations, resulting in complete V genes. These V genes were cloned into a phagemid expression vector in a single-chain form fused to the carboxyl-terminus of the M13 minor coat protein III. Phagemid particles displaying the single chain hybrid proteins on their surface were screened with Con A as Ag. Several clones showing specific binding to Con A were obtained after four rounds of selection and were further analyzed for their binding properties and DNA sequences. This method provides a novel way to create a naive combinatorial library without using mRNA from B lymphocytes as template. The method should be useful to isolate human antibodies that react with self-Ag.

Amino Acid Sequence↗

Characterization and molecular genetic mapping of microsatellite loci in pepper.

Microsatellites or simple sequence repeats are highly variable DNA sequences that can be used as informative markers for the genetic analysis of plants and animals. For the development of microsatellite markers in Capsicum, microsatellites were isolated from two small-insert genomic libraries and the GenBank database. Using five types of oligonucleotides, (AT)(15), (GA)(15), (GT)(15), (ATT)(10) and (TTG)(10), as probes, positive clones were isolated from the genomic libraries, and sequenced. Out of 130 positive clones, 77 clones showed microsatellite motifs, out of which 40 reliable microsatellite markers were developed. (GA)(n) and (GT)(n) sequences were found to occur most frequently in the pepper genome, followed by (TTG)(n) and (AT)(n). Additional 36 microsatellite primers were also developed from GenBank and other published data. To measure the information content of these markers, the polymorphism information contents (PICs) were calculated. Capsicum microsatellite markers from the genomic libraries have shown a high level of PIC value, 0.76, twice the value for markers from GenBank data. Forty six microsatellite loci were placed on the SNU-RFLP linkage map, which had been derived from the interspecific cross between Capsicum annuum "TF68" and Capsicum chinense "Habanero". The current "SNU2" pepper map with 333 markers in 15 linkage groups contains 46 SSR and 287 RFLP markers covering 1,761.5 cM with an average distance of 5.3 cM between markers.

Capsicum↗

Detecting cellulase and esterase enzyme activities encoded by novel genes present in environmental DNA libraries.

A genomic DNA library was made from the alkaliphilic cellulase-producing Bacillus agaradhaerans in order to prove our technologies for gene isolation prior to using them with samples of DNA isolated directly from environmental samples. Clones expressing a cellulase activity were identified and sequenced. A new cellulase gene was identified. Genomic DNA libraries were then made from DNA isolated directly from the Kenyan soda lakes, Lake Elmenteita and Crater Lake. Crater Lake clones expressing a cellulase activity and Lake Elmenteita clones expressing a lipase/esterase activity were identified and sequenced. These were encoded by novel genes as judged by DNA sequence comparisons. Genomic DNA libraries were also made from laboratory enrichment cultures of Lake Nakuru and Lake Elmenteita samples. Selective enrichment cultures were grown in the presence of carboxymethylcellulose (CMC) and olive oil. A number of new cellulase and lipase/esterase genes were discovered in these libraries. Cellulase-positive clones from Lake Nakuru were isolated at a frequency of 1 in 15,000 from a library made from a CMC enrichment as compared to 1 in 60,000 from a minimal medium enrichment. Esterase/lipase-positive clones from Lake Elmenteita were isolated with a frequency of 1 in 30,000 from a library made from an olive-oil enrichment as compared to 1 in 100,000 from an environmental library.

Amino Acid Sequence↗

[The establishment of genomic DNA libraries for the human malaria parasite Plasmodium falciparum].

The DNA of Plasmodium falciparum has been purified and fragmented with restriction endonuclease BamHI. The fragments have been incorporated in vitro into derivatives of bacteriophage lambda EMBL4 digested with BamHI and Sal I. The recombinant mixture has been ligated and packaged in vitro. The recombinant phages have been identified in E. coli L95 host cell and the libraries have been established in which most of the parasite DNA is represented. The ligation proportion of vector to insert is 3:1. The recombinant phages of 4 x 10(5) have been obtained. By plaque hybridization, we have been able to recover from these libraries specific clones containing repetitive DNA sequences.

Animals↗

Gene discovery within the planctomycete division of the domain Bacteria using sequence tags from genomic DNA libraries.

BACKGROUND: The planctomycetes comprise a distinct group of the domain Bacteria, forming a separate division by phylogenetic analysis. The organization of their cells into membrane-defined compartments including membrane-bounded nucleoids, their budding reproduction and complete absence of peptidoglycan distinguish them from most other Bacteria. A random sequencing approach was applied to the genomes of two planctomycete species, Gemmata obscuriglobus and Pirellula marina, to discover genes relevant to their cell biology and physiology. RESULTS: Genes with a wide variety of functions were identified in G. obscuriglobus and Pi. marina, including those of metabolism and biosynthesis, transport, regulation, translation and DNA replication, consistent with established phenotypic characters for these species. The genes sequenced were predominantly homologous to those in members of other divisions of the Bacteria, but there were also matches with nuclear genomic genes of the domain Eukarya, genes that may have appeared in the planctomycetes via horizontal gene transfer events. Significant among these matches are those with two genes atypical for Bacteria and with significant cell-biology implications - integrin alpha-V and inter-alpha-trypsin inhibitor protein - with homologs in G. obscuriglobus and Pi. marina respectively. CONCLUSIONS: The random-sequence-tag approach applied here to G. obscuriglobus and Pi. marina is the first report of gene recovery and analysis from members of the planctomycetes using genome-based methods. Gene homologs identified were predominantly similar to genes of Bacteria, but some significant best matches to genes from Eukarya suggest that lateral gene transfer events between domains may have involved this division at some time during its evolution.

Amino Acids↗

Molecular characterization of equine prostaglandin G/H synthase-2 and regulation of its messenger ribonucleic acid in preovulatory follicles.

To increase our understanding of the molecular control of PG synthesis in equine preovulatory follicles, the specific objectives of this study were to clone and determine the primary structure of equine prostaglandin G/H synthase-2 (PGHS-2) and to characterize the regulation of PGHS-2 messenger RNA (mRNA) in follicles before ovulation. A complementary DNA (cDNA) library prepared from follicular mRNA and a genomic library were screened with a mouse PGHS-2 cDNA probe to isolate the equine PGHS-2 cDNA and gene, respectively. The expression library yielded three nearly full-length clones that differed only in their 5'-ends; clones 3, 5, and 6 were 2946, 3138, and 3398 bp in length, respectively. The longest clone was shown to start 9 bp downstream of the transcription initiation site, as determined by primer extension analysis, and to contain 120 bp of 5'-untranslated region (UTR), 1812 bp of open reading frame, and 1466 bp of 3'-UTR. The open reading frame encodes a 604-amino acid protein that is more than 80% identical to PGHS-2 homologs in other species. Numerous repeats (n = 11) of the Shaw-Kamen's sequence (ATTTA) are present in the 3'-UTR, a motif typically indicative of mRNAs with a short half-life. The complete equine PGHS-2 gene was isolated and sequenced from a approximately 17-kilobase clone obtained from the genomic library. The equine PGHS-2 gene structure (10 exons and 9 introns; total length of 6991 bp) is similar to its human homolog except for lacking sequence elements in introns 4, 8, and 9 and in the 3'-UTR region of exon 10. To characterize the regulation of PGHS-2 mRNA in equine follicles before ovulation, preovulatory follicles were isolated during estrus, 0, 12, 24, 30, 33, 36, and 39 h (n = 4-5 follicles/time point) after an ovulatory dose of hCG. Results from Northern blots showed significant changes in steady state levels of PGHS-2 mRNA in preovulatory follicles after hCG treatment (P < 0.05). The transcript remained undetectable between 0-24 h post-hCG, first appeared (approximately 4 kilobases) only at 30 h, and reached maximal levels 33 h post-hCG. PGHS-2 mRNA was selectively induced in granulosa cells and not in theca interna. Thus, this study provides for the first time the primary structure of the equine PGHS-2 gene, transcript, and protein. It also demonstrates that the induction of PGHS-2 gene expression in equine granulosa cells is a long molecular process (30 h post-hCG), thereby providing a model to study the molecular basis for the late transcriptional activation of PGHS-2 in species with a long ovulatory process.

Amino Acid Sequence↗

Isolation and Characterization of Frankia sp. Strain FaC1 Genes Involved in Nitrogen Fixation.

Genomic DNA was isolated from Frankia sp. strain FaC1, an Alnus root nodule endophyte, and used to construct a genomic library in the cosmid vector pHC79. The genomic library was screened by in situ colony hybridization to identify clones of Frankia nitrogenase (nif) genes based on DNA sequence homology to structural nitrogenase genes from Klebsiella pneumoniae. Several Frankia nif clones were isolated, and hybridization with individual structural nitrogenase gene fragments (nifH, nifD, and nifK) from K. pneumoniae revealed that they all contain the nifD and nifK genes, but lack the nifH gene. Restriction endonuclease mapping of the nifD and nifK hybridizing region from one clone revealed that the nifD and nifK genes in Frankia sp. are contiguous, while the nifH gene is absent from a large region of DNA on either side of the nifDK gene cluster. Additional hybridizations with gene fragments derived from K. pneumoniae as probes and containing other genes involved in nitrogen fixation demonstrated that the Frankia nifE and nifN genes, which play a role in the biosynthesis of the iron-molybdenum cofactor, are located adjacent to the nifDK gene cluster.

Journal Article↗

Cloning and expression of a trypomastigote-specific 85-kilodalton surface antigen gene from Trypanosoma cruzi.

An 85-kDa trypomastigote-specific surface antigen gene from Trypanosoma cruzi has been identified by screening a genomic library in lambda gt10 with trypomastigote and epimastigote cDNA. The 1.3-kb genomic clone (pTt34) hybridizes to a single trypomastigote mRNA of 3.7 kb and to multiple bands in genomic Southern blots. Dot-blot experiments show that there are 5-10 copies of this sequence per haploid genome, and these are arranged in a non-tandem manner. pTt34 has been expressed as an anthranilate synthetase fusion protein in Escherichia coli, and inclusion bodies have been used to raise antiserum in rabbits. This antiserum immunoprecipitates a cell surface trypomastigote-specific protein of 85 kDa. The DNA and predicted amino acid sequences of pTt34 are given. Four further clones obtained from a PvuII/HpaI partial genomic library in pUC13 have extended the sequence of the 3' end of pTt34; each of these clones has regions of sequence divergence and each could represent a different member of the gene family.

Amino Acid Sequence↗

Identification and Cloning of Genes Involved in Specific Desulfurization of Dibenzothiophene by Rhodococcus sp. Strain IGTS8.

The gram-positive bacterium Rhodococcus sp. strain IGTS8 is able to remove sulfur from certain aromatic compounds without breaking carbon-carbon bonds. In particular, sulfur is removed from dibenzothiophene (DBT) to give the final product, 2-hydroxybiphenyl. A genomic library of IGTS8 was constructed in the cosmid vector pLAFR5, but no desulfurization phenotype was imparted to Escherichia coli. Therefore, IGTS8 was mutagenized, and a new strain (UV1) was selected that had lost the ability to desulfurize DBT. The genomic library was transferred into UV1, and several colonies that had regained the desulfurization phenotype were isolated, though free plasmid could not be isolated. Instead, vector DNA had integrated into either the chromosome or a large resident plasmid. DNA on either side of the inserted vector sequences was cloned and used to probe the original genomic library in E. coli. This procedure identified individual cosmid clones that, when electroporated into strain UV1, restored desulfurization. When the origin of replication from a Rhodococcus plasmid was inserted, the efficiency with which these clones transformed UV1 increased 20- to 50-fold and they could be retrieved as free plasmids. Restriction mapping and subcloning indicated that the desulfurization genes reside on a 4.0-kb DNA fragment. Finally, the phenotype was transferred to Rhodococcus fascians D188-5, a species normally incapable of desulfurizing DBT. The mutant strain, UV1, and R. fascians produced 2-hydroxybiphenyl from DBT when they contained appropriate clones, indicating that the genes for the entire pathway have been isolated.

Journal Article↗