Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Libraries”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

T-cell epitope prediction with combinatorial peptide libraries.

T cell receptors (TCR) recognize antigenic peptides in complex with the major histocompatibility complex (MHC) molecules and this trimolecular interaction initiates antigen-specific signaling pathways in the responding T lymphocytes. For the study of autoimmune diseases and vaccine development, it is important to identify peptides (epitopes) that can stimulate a given TCR. The use of combinatorial peptide libraries has recently been introduced as a powerful tool for this purpose. A combinatorial library of n-mer peptides is a set of complex mixtures each characterized by one position fixed to be a specified amino acid and all other positions randomized. A given TCR can be fingerprinted by screening a variety of combinatorial libraries using a proliferation assay. Here, we present statistical models for elucidating the recognition profile of a TCR using combinatorial library proliferation assay data and known MHC binding data.

Combinatorial Chemistry Techniques↗

Phage library-derived human anti-TETA and anti-DOTA ScFv for pretargeting RIT.

Pretargeting techniques have promise for radioimmunotherapy (RIT) in cancer because of the potential for markedly increasing the therapeutic ratio. Human antibody fragments can be retrieved from phage libraries and used to realize this potential. The library can be used to select the genetic material required to generate molecules with binding sites for both radiochelates and tumor antigens. In this study, human anti-chelate scFvs (single-chain fragments) for two metal chelates (Cu-TETA and Y-DOTA) were selected from a large, naive human scFv library. These anti-chelate scFvs were intended to serve as one arm of bispecific pretargeting molecules and to bind radiochelates given subsequently as Cu-67-TETA or Y-90-DOTA. Phage that displayed the anti-chelate scFv were selected by absorption to antibody (Lym-1) bound Cu-TETA or Y-DOTA. Enzyme-linked immunosorbent assays (ELISA) were performed to assess the intensity and specificity of phage binding to the specific chelate. Ninety-six clones demonstrating metal chelate binding seven times greater than to Lym-1 alone were chosen for diversity analysis. BstN I restriction digests were performed on DNA from these clones. Twenty-three and 43 different DNA fingerprint patterns were identified for anti-TETA and anti-DOTA clones, respectively. DNA sequencing of 39 anti-TETA clones for 23 different BstN I fingerprint patterns revealed 22 distinct sequences. Eleven of the anti-TETA clones were selected for further study. Five hundred to 1000 microg (100 to 320 microg per liter of culture) of purified scFv was produced from each of the 11 anti-TETA clones. Preliminary studies by BIAcore demonstrated evidence of 25- to 200-nM affinities. Comparable examination of the anti-DOTA clones is in progress. This study provides evidence that human scFv against unique synthetic targets can be readily selected from a large, naive human immunoglobulin phage library. Selections against metal chelated antibodies provided a wealth of scFvs with diverse binding affinities useful for engineering molecules for pretargeting RIT.

Antibodies, Monoclonal↗

HTSinfer: inferring metadata from bulk Illumina RNA-Seq libraries.

SUMMARY: The Sequencing Read Archive is one of the largest and fastest-growing repositories of sequencing data, containing tens of petabytes of sequenced reads. Its data is used by a wide scientific community, often beyond the primary study that generated them. Such analyses rely on accurate metadata concerning the type of experiment and library, as well as the organism from which the sequenced reads were derived. These metadata are typically entered manually by contributors in an error-prone process, and are frequently incomplete. In addition, easy-to-use computational tools that verify the consistency and completeness of metadata describing the libraries to facilitate data reuse, are largely unavailable. Here, we introduce HTSinfer, a Python-based tool to infer metadata directly and solely from bulk RNA-sequencing data generated on Illumina platforms. HTSinfer leverages genome sequence information and diagnostic genes to rapidly and accurately infer the library source and library type, as well as the relative read orientation, 3' adapter sequence and read length statistics. HTSinfer is written in a modular manner, published under a permissible free and open-source license and encourages contributions by the community, enabling easy addition of new functionalities, e.g. for the inference of additional metrics, or the support of different experiment types or sequencing platforms. AVAILABILITY AND IMPLEMENTATION: HTSinfer is released under the Apache License 2.0. Latest code is available via GitHub at https://github.com/zavolanlab/htsinfer, while releases are published on Bioconda. A snapshot of the HTSinfer version described in this article was deposited at Zenodo at 10.5281/zenodo.13985958.

Metadata↗

Differential expression in SAGE: accounting for normal between-library variation.

MOTIVATION: In contrasting levels of gene expression between groups of SAGE libraries, the libraries within each group are often combined and the counts for the tag of interest summed, and inference is made on the basis of these larger 'pseudolibraries'. While this captures the sampling variability inherent in the procedure, it fails to allow for normal variation in levels of the gene between individuals within the same group, and can consequently overstate the significance of the results. The effect is not slight: between-library variation can be hundreds of times the within-library variation. RESULTS: We introduce a beta-binomial sampling model that correctly incorporates both sources of variation. We show how to fit the parameters of this model, and introduce a test statistic for differential expression similar to a two-sample t-test.

Algorithms↗

Construction of libraries for methylation sites by in-gel competitive reassociation (IGCR).

The in-gel competitive reassociation (IGCR) procedure was successfully applied to construct a comprehensive library enriched in DNA fragments containing C5mCGG sequences from mouse liver and brain genomic DNA. For IGCR, methylation-insensitive restriction enzyme (Msp I) digests were used as target DNA and methylation-sensitive restriction enzyme (Hpa II) digests as competitor DNA. Southern blot analysis indicated that 60 to 70% of the clones in the library were derived from the methylated sites and overall enrichment was 200- to 1000-fold. IGCR was further applied to construct a library for the sites differentially methylated between brain and liver DNA. In the library, approximately 20% of the Hpa II sites exhibited different degrees of methylation between these tissues.

Animals↗

Construction and characterization of human brain cDNA libraries suitable for analysis of cDNA clones encoding relatively large proteins.

Analysis of proteins registered in the PIR protein database implied that most of relatively large proteins are related to important functions in higher multicellular organisms, but not many large proteins have been registered to date. To establish a protocol for efficient analysis of cDNA clones coding for large proteins, we constructed a series of strictly size-fractionated cDNA libraries of human brain, where the average insert sizes of cDNA clones ranged from 3.3 kb to 10 kb. As judged by hybridization analysis with probes derived from mRNAs of known sizes, the libraries with insert sizes up to 7 kb, at least, contained the clones corresponding to full-length transcripts in addition to truncated products of longer transcripts, but few chimeric clones. Using one of the fractionated libraries with an average insert size of 7 kb, the single-pass sequences from both the ends of randomly sampled clones were determined and sarched against DNA databases. Approximately 90% of the clones were found to be new with respect to their 5'-sequences while their 3'-sequences were frequently similar to the registered expression sequence tags. Examination of the protein-coding capacity in an in vitro transcription/translation system showed that about 20% of the clones direct the synthesis of proteins with apparent molecular masses larger than 50 kDa. The set of libraries constructed here should be very useful for the accumulation of sequence data on large proteins in the human brain.

Brain↗

Tuatara (Sphenodon) genomics: BAC library construction, sequence survey, and application to the DMRT gene family.

The tuatara (Sphenodon punctatus) is of "extraordinary biological interest" as the most distinctive surviving reptilian lineage (Rhyncocephalia) in the world. To provide a genomic resource for an understanding of genome evolution in reptiles, and as part of a larger project to produce genomic resources for various reptiles (evogen.jgi.doe.gov/second_levels/BACs/our_libraries.html), a large-insert bacterial artificial chromosome (BAC) library from a male tuatara was constructed. The library consists of 215 424 individual clones whose average insert size was empirically determined to be 145 kb, yielding a genomic coverage of approximately 6.3x. A BAC-end sequencing analysis of 121 420 bp of sequence revealed a genomic GC content of 46.8%, among the highest observed thus far for vertebrates, and identified several short interspersed repetitive elements (mammalian interspersed repeat-type repeats) and long interspersed repetitive elements, including chicken repeat 1 element. Finally, as a quality control measure the arrayed library was screened with probes corresponding to 2 conserved noncoding regions of the candidate sex-determining gene DMRT1 and the DM domain of the related DMRT2 gene. A deep coverage contig spanning nearly 300 kb was generated, supporting the deep coverage and utility of the library for exploring tuatara genomics.

Animals↗

Interplay of selective pressure and stochastic events directs evolution of the MEL172 satellite DNA library in root-knot nematodes.

According to the library model, related species can have in common satellite DNA (satDNA) families amplified in differing abundances, but reasons for persistence of particular sequences in the library during long periods of time are poorly understood. In this paper, we characterize 3 related satDNAs coexisting in the form of a library in mitotic parthenogenetic root-knot nematodes of the genus Meloidogyne. Due to sequence similarity and conserved monomer length of 172 bp, this group of satDNAs is named MEL172. Analysis of sequence variability patterns among monomers of the 3 MEL172 satellites revealed 2 low-variable (LV) domains highly reluctant to sequence changes, 2 moderately variable (MV) domains characterized by limited number of mutations, and 1 highly variable (HV) domain. The latter domain is prone to rapid spread and homogenization of changes. Comparison of the 3 MEL172 consensus sequences shows that the LV domains have 6% changed nucleotide positions, the MV domains have 48%, whereas 78% divergence is concentrated in the HV domain. Conserved distribution of intersatellite variability might indicate a complex pattern of interactions in heterochromatin, which limits the range and phasing of allowed changes, implying a possible selection imposed on monomer sequences. The lack of fixed species-diagnostic mutations in each of the examined MEL172 satellites suggests that they existed in unaltered form in a common ancestor of extant species. Consequently, the evolution of these satellites seems to be driven by interplay of selective constraints and stochastic events. We propose that new satellites were derived from an ancestral progenitor sequence by nonrandom accumulation of mutations due to selective pressure on particular sequence segments. In the library of particular taxa, established satellites might be subject to differential amplification at chance due to stochastic mechanisms of concerted evolution.

Animals↗

Efficient isolation and mapping of rad genes of the fungus Coprinus cinereus using chromosome-specific libraries.

We have constructed cosmid libraries from electrophoretically separated chromosomes of the basidiomycete Coprinus cinereus. These libraries greatly facilitate the isolation of genes by complementation of mutant phenotypes and are particularly useful for map-based cloning strategies. From a library constructed from two co-migrating C.cinereus chromosomes, we isolated a clone that complements the C.cinereus rad9-1 mutation. Examination of this clone showed that it complements both the repair and meiotic defects of this mutant. Restriction fragment length polymorphism mapping using a portion of this clone showed that it maps to the rad9 locus. In addition, a single copy of transforming DNA is sufficient to complement the rad9-1 defects. Thus, we believe we have cloned the rad9 gene itself. We also used a chromosome-specific library and backcrossed isolates to rapidly identify a cosmid clone which is tightly linked to the rad11 locus and is therefore a suitable starting point for a chromosome walk. These rapid methods of gene mapping and isolation should be applicable to any organism with separable chromosomes.

Cell Cycle Proteins↗

Preparation of a differentially expressed, full-length cDNA expression library by RecA-mediated triple-strand formation with subtractively enriched cDNA fragments.

We have developed a fast and general method to obtain an enriched, full-length cDNA expression library with subtractively enriched cDNA fragments. The procedure relies on RecA-mediated triple-helix formation of single-stranded cDNA fragments with a double-stranded cDNA plasmid library. The complexes were then captured from the solutions using the digoxigenin-antidigoxigenin paramagnetic beads followed by recovery of the enriched double-stranded cDNA expression library. We have observed a linear relation between the capture of full-length cDNAs in the library and the fold enrichment in the subtracted cDNA population.

Affinity Labels↗

Libraries for genomic SELEX.

An increasing number of proteins are being identified that regulate gene expression by binding specific nucleic acidsin vivo. A method termed genomic SELEX facilitates the rapid identification of networks of protein-nucleic acid interactions by identifying within the genomic sequences of an organism the highest affinity sites for any protein of the organism. As with its progenitor, SELEX of random-sequence nucleic acids, genomic SELEX involves iterative binding, partitioning, and amplification of nucleic acids. The two methods differ in that the variable region of the nucleic acid library for genomic SELEX is derived from the genome of an organism. We have used a quick and simple method to construct Escherichia coli, Saccharomyces cerevisiae, and human genomic DNA PCR libraries that can be transcribed with T7 RNA polymerase. We present evidence that the libraries contain overlapping inserts starting at most of the positions within the genome, making these libraries suitable for genomic SELEX.

Binding Sites↗

Solid-phase cDNA library construction, a versatile approach.

A rapid and versatile method for cDNA library construction was developed. It is based on conventional cDNA library synthesis including all enzymatic steps usually required, but is performed on a solid support. The cDNA is immobilised via a biotin residue to streptavidin coupled magnetic beads, which allows rapid and easy to perform changes of buffers and enzymes. Therefore, it combines speed (library construction within a single day) with high quality libraries, making it ideally suited for most purposes.

Base Sequence↗

Construction of a directed hammerhead ribozyme library: towards the identification of optimal target sites for antisense-mediated gene inhibition.

Antisense-mediated gene inhibition uses short complementary DNA or RNA oligonucleotides to block expression of any mRNA of interest. A key parameter in the success or failure of an antisense therapy is the identification of a suitable target site on the chosen mRNA. Ultimately, the accessibility of the target to the antisense agent determines target suitability. Since accessibility is a function of many complex factors, it is currently beyond our ability to predict. Consequently, identification of the most effective target(s) requires examination of every site. Towards this goal, we describe a method to construct directed ribozyme libraries against any chosen mRNA. The library contains nearly equal amounts of ribozymes targeting every site on the chosen transcript and the library only contains ribozymes capable of binding to that transcript. Expression of the ribozyme library in cultured cells should allow identification of optimal target sites under natural conditions, subject to the complexities of a fully functional cell. Optimal target sites identified in this manner should be the most effective sites for therapeutic intervention.

Base Sequence↗

Identification and prevention of a GC content bias in SAGE libraries.

Serial Analysis of Gene Expression (SAGE) is becoming a widely used gene expression profiling method for the study of development, cancer and other human diseases. Investigators using SAGE rely heavily on the quantitative aspect of this method for cataloging gene expression and comparing multiple SAGE libraries. We have developed additional computational and statistical tools to assess the quality and reproducibility of a SAGE library. Using these methods, a critical variable in the SAGE protocol was identified that has the potential to bias the Tag distribution relative to the GC content of the 10 bp SAGE Tag DNA sequence. We also detected this bias in a number of publicly available SAGE libraries. It is important to note that the GC content bias went undetected by quality control procedures in the current SAGE protocol and was only identified with the use of these statistical analyses on as few as 750 SAGE Tags. In addition to keeping any solution of free DiTags on ice, an analysis of the GC content should be performed before sequencing large numbers of SAGE Tags to be confident that SAGE libraries are free from experimental bias.

Animals↗

Rapid generation of incremental truncation libraries for protein engineering using alpha-phosphothioate nucleotides.

Incremental truncation for the creation of hybrid enzymes (ITCHY) is a novel tool for the generation of combinatorial libraries of hybrid proteins independent of DNA sequence homology. We herein report a fundamentally different methodology for creating incremental truncation libraries using nucleotide triphosphate analogs. Central to the method is the polymerase catalyzed, low frequency, random incorporation of alpha-phosphothioate dNTPs into the region of DNA targeted for truncation. The resulting phosphothioate internucleotide linkages are resistant to 3'-->5' exonuclease hydrolysis, rendering the target DNA resistant to degradation in a subsequent exonuclease III treatment. From an experimental perspective the protocol reported here to create incremental truncation libraries is simpler and less time consuming than previous approaches by combining the two gene fragments in a single vector and eliminating additional purification steps. As proof of principle, an incremental truncation library of fusions between the N-terminal fragment of Escherichia coli glycinamide ribonucleotide formyltransferase (PurN) and the C-terminal fragment of human glycinamide ribonucleotide formyltransferase (hGART) was prepared and successfully tested for functional hybrids in an auxotrophic E.coli host strain. Multiple active hybrid enzymes were identified, including ones fused in regions of low sequence homology.

Amino Acid Sequence↗

CpG Island microarray probe sequences derived from a physical library are representative of CpG Islands annotated on the human genome.

An effective tool for the global analysis of both DNA methylation status and protein-chromatin interactions is a microarray constructed with sequences containing regulatory elements. One type of array suited for this purpose takes advantage of the strong association between CpG Islands (CGIs) and gene regulatory regions. We have obtained 20,736 clones from a CGI Library and used these to construct CGI arrays. The utility of this library requires proper annotation and assessment of the clones, including CpG content, genomic origin and proximity to neighboring genes. Alignment of clone sequences to the human genome (UCSC hg17) identified 9595 distinct genomic loci; 64% were defined by a single clone while the remaining 36% were represented by multiple, redundant clones. Approximately 68% of the loci were located near a transcription start site. The distribution of these loci covered all 23 chromosomes, with 63% overlapping a bioinformatically identified CGI. The high representation of genomic CGI in this rich collection of clones supports the utilization of microarrays produced with this library for the study of global epigenetic mechanisms and protein-chromatin interactions. A browsable database is available on-line to facilitate exploration of the CGIs in this library and their association with annotated genes or promoter elements.

Base Sequence↗

Biotin-tagged cDNA expression libraries displayed on lambda phage: a new tool for the selection of natural protein ligands.

cDNA expression libraries displayed on lambda phage have been successfully employed to identify partners involved in antibody-antigen, protein- protein and DNA-protein interactions and represent a novel approach to functional genomics. However, as in all other cDNA expression libraries based on fusion to a carrier polypeptide, a major issue of this system is the absence of control over the translation frame of the cDNA. As a consequence, a large number of clones will contain lambda D/cDNA fusions, resulting in the foreign sequence being translated on alternative reading frames. Thus, many phage will not display natural proteins, but could be selected, as they mimic the binding properties of the real ligand, and will hence interfere with the selection outcome. Here we describe a novel lambda vector for display of exogenous peptides at the C-terminus of the capsid D protein. In this vector, translation of fusion peptides in the correct reading frame allows efficient in vivo biotinylation of the chimeric phage during amplification. Using this vector system we constructed three libraries from human hepatoma cells, mouse hepatocytic MMH cells and from human brain. Clones containing open reading frames (ORFs) were rapidly selected by streptavidin affinity chromatography, leading to biological repertoires highly enriched in natural polypeptides. We compared the selection outcome of two independent experiments performed using an anti-GAP-43 monoclonal antibody on the human brain cDNA library before and after ORF enrichment. A significant increase in the efficiency of identification of natural target peptides with very little background of false-positive clones was observed in the latter case.

Amino Acid Sequence↗

Designing gene libraries from protein profiles for combinatorial protein experiments.

Protein combinatorial libraries provide new ways to probe the determinants of folding and to discover novel proteins. Such libraries are often constructed by expressing an ensemble of partially random gene sequences. Given the intractably large number of possible sequences, some limitation on diversity must be imposed. A non-uniform distribution of nucleotides can be used to reduce the number of possible sequences and encode peptide sequences having a predetermined set of amino acid probabilities at each residue position, i.e., the amino acid sequence profile. Such profiles can be determined by inspection, multiple sequence alignment or physically-based computational methods. Here we present a computational method that takes as input a desired sequence profile and calculates the individual nucleotide probabilities among partially random genes. The calculated gene library can be readily used in the context of standard DNA synthesis to generate a protein library with essentially the desired profile. The fidelity between the desired profile and the calculated one coded by these partially random genes is quantitatively evaluated using the linear correlation coefficient and a relative entropy, each of which provides a measure of profile agreement at each position of the sequence. On average, this method of identifying such codon frequencies performs as well or better than other methods with regard to fidelity to the original profile. Importantly, the method presented here provides much better yields of complete sequences that do not contain stop codons, a feature that is particularly important when all or large fractions of a gene are subject to combinatorial mutation.

Amino Acids↗