Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The diversity and evolutionary relationships of the pregnancy-associated glycoproteins, an aspartic proteinase subfamily consisting of many trophoblast-expressed genes.

The pregnancy-associated glycoproteins (PAGs) are structurally related to the pepsins, thought to be restricted to the hooved (ungulate) mammals and characterized by being expressed specifically in the outer epithelial cell layer (chorion/trophectoderm) of the placenta. At least some PAGs are catalytically inactive as proteinases, although each appears to possess a cleft capable of binding peptides. By cloning expressed genes from ovine and bovine placental cDNA libraries, by Southern genomic blotting, by screening genomic libraries, and by using PCR to amplify portions of PAG genes from genomic DNA, we estimate that cattle, sheep, and most probably all ruminant Artiodactyla possess many, possibly 100 or more, PAG genes, many of which are placentally expressed. The PAGs are highly diverse in sequence, with regions of hypervariability confined largely to surface-exposed loops. Nonsynonymous (replacement) mutations in the regions of the genes coding for these hypervariable loop segments have accumulated at a higher rate than synonymous (silent) mutations. Construction of distance phylograms, based on comparisons of PAG and related aspartic proteinase amino acid sequences, suggests that much diversification of the PAG genes occurred after the divergence of the Artiodactyla and Perissodactyla, but that at least one gene is represented outside the hooved species. The results also suggest that positive selection of duplicated genes has acted to provide considerable functional diversity among the PAGs, whose presence at the interface between the placenta and endometrium and in the maternal circulation indicates involvement in fetal-maternal interactions.

Amino Acid Sequence↗

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning↗

Preparation and properties of recombinant corynebacterial sarcosine oxidase: evidence for posttranslational modification during turnover with sarcosine.

The genes encoding the four subunits of sarcosine oxidase from Corynebacterium sp. P-1 were isolated and overexpressed in a single step by using indicator plates to screen a genomic library for colonies that generated hydrogen peroxide in a sarcosine-dependent reaction. The genomic library was constructed by inserting size-fractionated genomic DNA, previously subjected to partial digestion by Sau3AI, into pBluescript II SK (+). At least 1.0 kb, but less than 4.0 kb, can be deleted from the 3' end of the original cornyebacterial insert (7.3 kb) without affecting sarcosine oxidase expression, consistent with the estimated 5.0-kb operon size. Recombinant sarcosine oxidase is isolated as a heterotetramer containing equimolar amounts of covalent and noncovalent flavin, identical to that observed for enzyme isolated from Corynebacterium sp. P-1. Despite its similar flavin content, recombinant enzyme exhibits significantly different spectral properties than enzyme from Corynebacterium sp. P-1 (values shown in parentheses) [epsilon 450 = 9.7 (12.7) mM-1 cm-1; A368/A450 = 1.0 (0.83); A280/A450 = 16.9 (12.2)]. This difference is due to the fact that about half of the covalent flavin in recombinant enzyme forms a reversible covalent 4a-adduct with a cysteine residue (lambda max = 383 nm; epsilon 383 = 7.3 mM-1 cm-1). The equilibrium is shifted in favor of adduct dissociation by oxidizing the cysteine residue with hydrogen peroxide or by alkylation with methyl methanethiosulfonate in a reaction that is fully reversible upon addition of excess dithiothreitol. The cysteine residue is also oxidized during aerobic turnover with sarcosine. Reaction of the cysteine residue with hydrogen peroxide (or a precursor) formed during turnover partially competes with the release of hydrogen peroxide into solution, as judged by the effect of catalase on this reaction. Although the same specific activity is observed for recombinant enzyme and enzyme from Corynebacterium sp. P-1, the recombinant enzyme exhibits a pronounced lag in an NADH peroxidase-coupled assay. The lag is eliminated by prior disruption of the 4a-thiolate adduct via reaction with hydrogen peroxide or methyl methanethiosulfonate. The results show that the 4a-thiolate adduct is an inactive form of sarcosine oxidase that can be activated by reaction with sarcosine in what appears to be the first example of a posttranslational modification associated with turnover. Complete activation occurs in vivo when sarcosine oxidase is produced in Corynebacterium sp. P-1, where enzyme synthesis is induced by growth of the organism with sarcosine as the source of carbon and energy.(ABSTRACT TRUNCATED AT 400 WORDS)

Chromatography, Gel↗

The human 3 beta-hydroxysteroid dehydrogenase (3 beta-HSD) gene cluster on chromosome 1p13 contains a presumptive pseudogene; 3 beta-HSD and CYP17 do not segregate with dominantly inherited hirsutism.

Four hirsute females from a family exhibiting idiopathic dominant hirsutism were examined. Basal blood levels of delta 5 and delta 4 steroids were within the normal range, but ACTH stimulation led to increases in 17-hydroxypregnenolone and dehydroepiandrosterone that were significantly above control levels. Using polymorphic genetic markers, the genes for cytochrome P450c1717 encoded by CYP17, and the type I and II forms of 3 beta-hydroxysteroid dehydrogenase (3 beta-HSD) were found not to segregate with hirsutism in this family, though a base substitution was detected in the 3' end of exon 1 of the gene for 3 beta-HSD type I in three of the four patients investigated. Analysis of PCR patients amplification products by denaturing gradient gel electrophoresis (DGGE) and sequencing revealed a novel homologue of exon 3 of 3 beta-HSD. DNA of one of the affected patients was used to create a genomic library in lambda gem 11 and clones containing the novel homologue were obtained and partially sequenced. The equivalent clone was obtained from a genomic library of an unrelated normal individual. The sequences of the clones from patient and control were identical and homologous to exons 2-4 of human 3 beta-HSD types I and II. No difference was found in the PCR primer sites that flanked the exons 3 homologue which led to its detection on DGGE gels. In both clones, stop codons and deletions were identified in the exon 4 homologue, leading to the deduction that the sequence comes from a pseudogene, which we call 3 beta-HSD psi 1. The pseudogene mapped to chromosome 1p13. It was concluded that dominantly inherited idiopathic hirsutism in this rare kindred was not due to deficiencies in 3 beta-HSD types I, II, or psi or of CYP17).

3-Hydroxysteroid Dehydrogenases↗

The 5' end sequences and exon organization in rat regucalcin gene.

The 5'-flanking region of the gene for a Ca(2+)-binding protein regucalcin was cloned from a rat genomic library which was constructed in lambda EMBL3 SP6/T7 vector. The genomic library was screened by using the radiolabeled probe with the 5' region (0.5 kb) of rat regucalcin complementary deoxyribonucleic acid (cDNA). Positive clone had the 5.5 kb fragment which was hybridized with the 5'-probe. This fragment contained three exons (I-III) of the gene coding for a rat regucalcin. The nucleotide sequence of exons completely agreed with that of a rat regucalcin cDNA clone. A supposed translational initiation site existed in the exon II. Homology analysis showed that a putative transcription start site in the rat regucalcin gene was located at position 26 downstream from a TATA-box. Another upstream element, a CCAAT box-like sequence, was located at -170. Moreover, there were many regulatory elements (Hox, AP-1, AP-2 and AP-4) in the 5'-flanking region of the rat regucalcin gene. The organization of rat regucalcin gene seemed to be about 18 kb in size and consisted of seven exons and six introns.

Amino Acid Sequence↗

The GDI1 genes from Kluyveromyces lactis and Pichia pastoris: cloning and functional expression in Saccharomyces cerevisiae.

The nucleotide sequences of 2.8 kb and 2.9 kb fragments containing the Kluyveromyces lactis and Pichia pastoris GDI1 genes, respectively, were determined. K. lactis GDI1 was found during sequencing of a genomic library clone, whereas the P. pastoris GDI1 was obtained from a genomic library by complementing a Saccharomyces cerevisiae sec19-1 mutant strain. The sequenced DNA fragments contain open reading frames of 1338 bp (K.lactis) and 1344 bp (P. pastoris), coding for polypeptides of 445 and 447 residues, respectively. Both sequences fully complement the S. cerevisiae sec19-1 mutation. They have high degrees of homology with known GDP dissociation inhibitors from yeast species and other eukaryotes.

Amino Acid Sequence↗

Phylogenetics of modern birds in the era of genomics.

In the 14 years since the first higher-level bird phylogenies based on DNA sequence data, avian phylogenetics has witnessed the advent and maturation of the genomics era, the completion of the chicken genome and a suite of technologies that promise to add considerably to the agenda of avian phylogenetics. In this review, we summarize current approaches and data characteristics of recent higher-level bird studies and suggest a number of as yet untested molecular and analytical approaches for the unfolding tree of life for birds. A variety of comparative genomics strategies, including adoption of objective quality scores for sequence data, analysis of contiguous DNA sequences provided by large-insert genomic libraries, and the systematic use of retroposon insertions and other rare genomic changes all promise an integrated phylogenetics that is solidly grounded in genome evolution. The avian genome is an excellent testing ground for such approaches because of the more balanced representation of single-copy and repetitive DNA regions than in mammals. Although comparative genomics has a number of obvious uses in avian phylogenetics, its application to large numbers of taxa poses a number of methodological and infrastructural challenges, and can be greatly facilitated by a 'community genomics' approach in which the modest sequencing throughputs of single PI laboratories are pooled to produce larger, complementary datasets. Although the polymerase chain reaction era of avian phylogenetics is far from complete, the comparative genomics era-with its ability to vastly increase the number and type of molecular characters and to provide a genomic context for these characters-will usher in a host of new perspectives and opportunities for integrating genome evolution and avian phylogenetics.

Animals↗

Mapping of rabbit chromosome 1 markers generated from a microsatellite-enriched chromosome-specific library.

A genomic DNA library was produced from flow-sorted rabbit chromosome 1 and enriched for fragments containing CA-repeats. Clones containing CA-repeats were identified and primers for amplification of the microsatellite were developed after sequencing the clone. The degree of polymorphism was tested in rabbits from different breeds. This approach identified 12 microsatellite markers which could be used for studying linkage relationships in the progeny of an F(2)-intercross: (AX/JUxIIIVO/JU) F(2), and two backcrosses: (OS/JxX/J)X/J and (WH/JxX/J)X/J. Seven of these markers were mapped on chromosome 1.

Animals↗

Isolation and structural analysis of a ribosomal protein gene in D.melanogaster.

By using the cDNA clone containing the sequence for the L1 ribosomal protein gene of Xenopus laevis as probe (1), we have isolated positive phages from a Drosophila melanogaster genomic library. The Drosophila genomic fragment, which gives the hybridization signal with the Xenopus cDNA, was sequenced: a region of 369 bp is 70% homologous to the sequence of X. laevis L1 cDNA. The gene was localized in situ at position 98AB of the right arm of the third polytene chromosome. By S1 mapping and heteroduplex analysis we have found that the gene is interrupted by three introns. A Drosophila cDNA embryonic library was screened and three cDNA clones were isolated (900, 1400 and 1500 nt long). By Northern analysis the cDNAs identify a 1400nt transcript present at every stage of development. By the features described, the clones we have isolated identify the Drosophila rp gene homologous to the L1 rp gene of Xenopus and could code for the L1 ribosomal protein described in D. melanogaster.

Amino Acid Sequence↗

The integrated genome map of Mycobacterium leprae.

The integrated map of the Mycobacterium leprae genome unveiled for the first time the genomic organization of this obligate intracellular parasite. Selected cosmid clones, isolated from a genomic library created in the cosmid vector Lorist6, were identified as representing nearly the complete genome and were subsequently used in the M. leprae genome sequencing project. Now a new version of the integrated map of M. leprae can be presented, combining the mapping results from the Lorist6 cosmids with data obtained from a second genomic library constructed in an Escherichia coli-mycobacterium shuttle cosmid, pYUB18. More than 98% of the M. leprae genome is now covered by overlapping large insert genomic clones representing a renewable source of well defined DNA segments and a powerful tool for functional genomics.

Animals↗

Genomic cloning of feline Fas ligand gene and characterization of the transcription regulatory region.

The feline Fas ligand gene was molecularly cloned from a feline genomic library and its genomic organization was determined. The feline Fas ligand gene contained four exons and spanned approximately 10 kb in the genome, and thus had the same structure as the human Fas ligand gene. The promoter region of the feline Fas ligand gene was further characterized by deletion mapping. The region between nucleotides -459 and -172 relative to the ATG codon was essential for the promoter activity when transfected into human and feline lymphoid cells. The characterization of the feline Fas ligand gene in the present study will be useful for further investigations of the regulatory mechanism of feline Fas ligand expression.

Amino Acid Sequence↗

YAC cloning and identification of rice (Oriza sativa L.) genomic DNA.

Construction of the genomic library by using yeast artificial chromosomes (YAC) has been an approach to a map-based gene cloning strategy. We have obtained more than 2000 YAC clones for the construction of a rice (Oriza sativa L.) genomic library through a procedure that could be summarized as follows: a fractionate by the pulsed-field gel electrophoresis (PFGE), the rice nucleic high molecular weight (HMW) DNA which is partially digested with EcoRI, and recovered fragments larger than 200 kb; ligate the fragments to the EcoRI digested YAC vector pairs pJS97 and pJS98; transform competent spheroplasts prepared from yeast strain YPH252; and select transformants directly on Ura-Trp-double selective media. Southern hybridization results indicated that the sizes of inserts were in the range of 200-820 kb.

Chromosomes, Artificial, Yeast↗

Molecular characterization of Hmg2 gene encoding a 3-hydroxy-methylglutaryl-CoA reductase in rice.

Three genes encoding 3-hydroxy-3-methylglutaryl-CoA reductase (HMGR, EC1.1.1.34), which converts HMG-CoA into mevalonate in the early key step of the plant isoprenoid pathway, were isolated by RT-PCR and rice cDNA and genomic library screening. A genomic Southern blot analysis confirmed that HMGR genes are present in three copies in rice. Of the three, the HMGR 2 gene (Hmg2) obtained as a cDNA clone and its genomic clone had 4 exons and 3 introns, and encoded a 576 amino acid peptide containing an open reading frame of 1,728 bp with a calculated Mw. of 61,150. The structure of rice Hmg2 had common features, based on its nucleotide and deduced amino acid sequence homologies, with other plant HMGR genes published to date. Rice Hmg2 transcripts were constitutively detected in all parts of the rice plant, except in lamina and their levels were high particularly in the leaf part of the dark-grown seedlings and mature flowers. Our result showed that mRNA levels of rice Hmg2 were strongly induced in seedlings and influorescence in the early development stage. Rice Hmg2 possibly has a housekeeping role involved in the sterol biosynthesis, among the possible roles of plant HMGR genes that have been suggested in other plants [Weissenborn et al. (1995)].

Amino Acid Sequence↗

The very large amplifiable element AUD2 from Streptomyces lividans 66 has insertion sequence-like repeats at its ends.

In a spontaneous, chloramphenicol-sensitive (Cms), arginine-auxotrophic (Arg-) mutant of Streptomyces lividans 1326, two amplified DNA sequences were found. One of them was the well-characterized 5.7-kb ADS1 sequence, amplified to about 300 copies per chromosome. The second one was a 92-kb sequence called ADS2. ADS2 encoding the previously isolated mercury resistance genes of S. lividans was amplified to around 20 copies per chromosome. The complete ADS2 sequence was isolated from a genomic library of the mutant S. lividans 1326.32, constructed in the phage vector lambda EMBL4. In addition, the DNA sequences flanking the corresponding amplifiable element called AUD2 in the wild-type strain were isolated by using another genomic library prepared from S. lividans 1326 DNA. Analysis of the ends of AUD2 revealed the presence of an 846-bp sequence on both sides repeated in the same orientation. Each of the direct repeats ended with 18-bp inverted repeated sequences. This insertion sequence-like structure was confirmed by the DNA sequence determined from the amplified copy of the direct repeats which demonstrated a high degree of similarity of 65% identity in nucleic acid sequence to IS112 from Streptomyces albus. The recombination event leading to the amplification of AUD2 occurred within these direct repeats, as shown by DNA sequence analysis. The amplification of AUD2 was correlated with a deletion on one side of the flanking chromosomal region beginning very near or in the amplified DNA. Strains of S. lividans like TK20 and TK21 which are mercury sensitive have completely lost AUD2 together with flanking chromosomal DNA on one or both sides.

Amino Acid Sequence↗

Molecular cloning and chromosomal mapping of the mouse gene encoding cyclin-dependent kinase 5 regulatory subunit p35.

A neural-specific activating subunit, p35, of cyclin-dependent kinase 5 (Cdk5) was recently reported to differ from other mammalian cyclins, suggesting a new type of regulatory subunit for Cdk activity. The mouse gene encoding p35, Cdk5r, was isolated from a mouse 129/SvJ genomic library, and the genomic structure of Cdk5r was characterized. The most notable features of Cdk5r are the absence of introns in the amino acid coding region and the high homology of amino acid sequence among species. The 5'-flanking region of Cdk5r contained no canonical TATA or CAAT box but had several putative promoter elements, including Sp1, AP2, MRE, and NGFIA. The mouse Cdk5r transcript was detected only in the brain by Northern blot analysis. Mouse Cdk5r was mapped to a position on mouse chromosome 11.

Amino Acid Sequence↗

CgT1: a non-LTR retrotransposon with restricted distribution in the fungal phytopathogen Colletotrichum gloeosporioides.

Two genetically distinct biotypes (A and B) of Colletotrichum gloeosporioides that cause different anthracnose diseases on the legumes Stylosanthes spp. have been identified in Australia. A DNA sequence that was present in biotype B and absent in biotype A was isolated by differential hybridisation of a genomic library using total genomic DNA of each biotype as hybridisation probes. This sequence also failed to hybridise to DNA of three biotypes of C. gloeosporioides from other host species and to DNA of three other species of Colletotrichum. This clone was used to isolate two cosmid clones of biotype B. Sequence analysis of these clones revealed a repetitive element of approximately 5.7 kb in length. This element, termed CgT1, was dispersed in the genome and present in about 30 copies. The element contained open reading frames encoding deduced sequence motifs homologous to gag-like proteins, reverse transcriptase and RNase H domains of non-LTR retrotransposons. The termini of CgT1 lacked long terminal repeats (LTRs) but contained a 3' A-rich domain. The insertion site of one copy of the element was flanked by short 13-bp direct repeats. These characteristics of the termini, taken together with the overall structure and sequence homologies, indicate that CgT1 belongs to the non-LTR, LINE-like retrotransposon class of elements that are present in many eukaryotes. PCR primers designed to amplify regions of CgT1 can be used to distinguish biotypes A and B in Australia. DNA fingerprinting analysis of genomic DNA using hybridisation probes derived from the terminal regions of CgT1 revealed that Australian isolates of biotype B are monomorphic. CgT1 was not detected in some isolates causing Type B disease from other countries and when CgT1 was present there was considerable polymorphism in CgT1 organisation in the genome. CgT1 is the first transposon-like element to be identified in the genus Colletotrichum and has considerable potential as a tool for the study of population structure, genome dynamics and evolution in C. gloeosporioides.

Amino Acid Sequence↗

Distribution of interstitial telomere-like repeats and their adjacent sequences in a dioecious plant, Silene latifolia.

The dioecious plant Silene latifolia has large, heteromorphic X and Y sex chromosomes that are thought to be derived from rearrangements of autosomes. To reveal the origin of the sex chromosomes in S. latifolia, we isolated and characterized telomere-homologous sequences from intra-chromosomal regions (interstitial telomere-like repeats; ITRs) and ITR-adjacent sequences (IASs). Nine genomic DNA fragments with degenerate 84- to 175-bp ITRs were isolated from a genomic library and total genome of male plants. Comparing the nucleotide sequences, the IASs of the nine ITRs were classified into seven elements (IAS-a, IAS-b, IAS-c, IAS-d, IAS-e, IAS-f, and IAS-g) by sequence similarity. The ITRs were grouped into two classes (class-I and -II ITRs) according to the classification of IASs. The class-I ITRs were sub-grouped into three subclasses (subclasses-IA, -IB, and -IC ITRs) based on the arrangement of IAS elements. By contrast, the class-II ITR was located between two different IASs (IAS-f and IAS-g). Genomic Southern analyses showed that both the male and female genomes contained six (IAS-f) to 153 (IAS-d) copies of each IAS per haploid genome. Fluorescence in situ hybridization analyses showed that one IAS element, IAS-d, was distributed in the interstitial and proximal regions of the sex chromosomes of S. latifolia. The distribution of IAS-d is important evidence for past telomere-mediated chromosome rearrangements during the evolution of the sex chromosomes of S. latifolia.

Base Sequence↗

Design of Onchocerca DNA probes based upon analysis of a repeated sequence family.

Repeated DNA sequences have been instrumental in the development of DNA probes for many different parasites. Isolation of such DNA probes has generally been accomplished by differential screening of genomic libraries with total genomic DNA preparations. In the current work, a rational design strategy is presented for the development of oligonucleotide probes based upon repeated sequence families. A repeated sequence family present in the genome of Onchocerca parasites, designated O-150, has been amplified from various samples of genomic DNA using PCR. DNA sequence analysis of the resulting PCR products demonstrated that the sequences may be arranged into clusters within which the individual sequences are identical or nearly identical. Differences among the cluster consensus sequences have been exploited to explain the specificities of previously isolated O-150 based probes and to develop two new oligonucleotide probes. One of these probes hybridizes specifically to Onchocerca volvulus O-150 PCR products, while the second hybridizes specifically to O-150 PCR products from the closely related bovine parasite O. ochengi. These oligonucleotide probes have been used to characterize Onchocerca infective larvae isolated from wild caught infected flies in West Africa. Because repeated sequence families are a common feature of most genomes, including those of parasites, this method should be applicable to the rational design of oligonucleotide probes for other parasitic infections.

Animals↗