Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Completion of the genome sequence of Brucella abortus and comparison to the highly similar genomes of Brucella melitensis and Brucella suis.

Brucellosis is a worldwide disease of humans and livestock that is caused by a number of very closely related classical Brucella species in the alpha-2 subdivision of the Proteobacteria. We report the complete genome sequence of Brucella abortus field isolate 9-941 and compare it to those of Brucella suis 1330 and Brucella melitensis 16 M. The genomes of these Brucella species are strikingly similar, with nearly identical genetic content and gene organization. However, a number of insertion-deletion events and several polymorphic regions encoding putative outer membrane proteins were identified among the genomes. Several fragments previously identified as unique to either B. suis or B. melitensis were present in the B. abortus genome. Even though several fragments were shared between only B. abortus and B. suis, B. abortus shared more fragments and had fewer nucleotide polymorphisms with B. melitensis than B. suis. The complete genomic sequence of B. abortus provides an important resource for further investigations into determinants of the pathogenicity and virulence phenotypes of these bacteria.

Bacterial Proteins↗

Analysis of chicken TLR4, CD28, MIF, MD-2, and LITAF genes in a Salmonella enteritidis resource population.

Salmonella enteritidis is a foodborne pathogen that negatively affects both animal and human health. Genetic variations in response to pathogenic SE colonization or to SE vaccination were measured in a chicken resource population. Outbred broiler sires and 3 diverse, highly inbred dam lines produced 508 F1 progeny that were evaluated for either bacterial colonization after pathogenic SE inoculation or circulating antibody level after SE vaccination. Five candidate genes were selected for study, based on their biological function as possibly affecting response to SE: toll-like receptor 4 (TLR4), T-cell specific surface protein (CD28), macrophage migration inhibitory factor (MIF), MD-2, and lipopolysaccharide-induced tumor necrosis factor (TNF)-alpha factor (LITAF). Gene fragments were sequenced from the founder lines of the resource population. The LITAF and MIF genes were homozygous for all sires. Single nucleotide polymorphisms (SNP) were identified in 3 genes (TLR4, CD28, and MD-2) and were used to test for associations of sire SNP with SE response. Linear mixed models were used for statistical analyses. The CD28 broiler sire SNP was associated with both bacterial load in the cecum (P < 0.003) and vaccine antibody response (P < 0.05). The MD-2 SNP was associated (P < 0.04) with the bacterial load in the spleen. The use of these SNP in these genes in marker-assisted selection may result in enhancement of disease resistance.

Animals↗

Conservation genetics of the European brown bear--a study using excremental PCR of nuclear and mitochondrial sequences.

In the Brenta area of northern Italy, a brown bear Ursus arctos population is rapidly going extinct. Restocking of the population is planned. In order to study the genetics of this highly vulnerable population with a minimum of stress to the animals we have developed a PCR-based method that allows the study of mitochondrial and nuclear gene sequences from droppings collected in the field. This method is generally applicable to animals in the wild. Using excremental as well as hair samples, we show that the Brenta population is monomorphic for one mitochondrial lineage and that female as well as male bears exist in the area. In addition, 70 samples from other parts of Europe were studied. As others have previously reported, the mitochondrial gene pool of European bears is divided into two major clades, one with a western and the other with an eastern distribution. Whereas populations generally belong to either one or the other mitochondrial clade, the Romanian population contains both clades. The bears in the Brenta belong to the western clade. The implications for the management of brown bears in the Brenta and elsewhere in Europe are discussed.

Animals↗

Anchoring of rice BAC clones to the rice genetic map in silico.

A wealth of molecular resources have been developed for rice genomics, including dense genetic maps, expressed sequence tags (ESTs), yeast artificial chromosome maps, bacterial artificial chromosome (BAC) libraries and BAC end sequence databases. Integration of genetic and physical maps involves labor-intensive empirical experiments. To accelerate the integration of the bacterial clone resources with the genetic map for the International Rice Genome Sequencing Project, we cleaned and filtered the available EST and BAC end sequences for repetitive sequences and then searched all available rice genetic markers with our filtered databases. We identified 418 genetic markers that aligned with at least one BAC end sequence with >95% sequence identity, providing a set of large insert clones with an average separation of 1 Mb that can serve as nucleation points for the sequencing phase of the International Rice Genome Sequencing Project.

Chromosome Mapping↗

Whole genome sequence comparisons and "full-length" cDNA sequences: a combined approach to evaluate and improve Arabidopsis genome annotation.

To evaluate the existing annotation of the Arabidopsis genome further, we generated a collection of evolutionary conserved regions (ecores) between Arabidopsis and rice. The ecore analysis provides evidence that the gene catalog of Arabidopsis is not yet complete, and that a number of these annotations require re-examination. To improve the Arabidopsis genome annotation further, we used a novel "full-length" enriched cDNA collection prepared from several tissues. An additional 1931 genes were covered by new "full-length" cDNA sequences, raising the number of annotated genes with a corresponding "full-length" cDNA sequence to about 14,000. Detailed comparisons between these "full-length" cDNA sequences and annotated genes show that this resource is very helpful in determining the correct structure of genes, in particular, those not yet supported by "full-length" cDNAs. In addition, a total of 326 genomic regions not included previously in the Arabidopsis genome annotation were detected by this cDNA resource, providing clues for new gene discovery. Because, as expected, the two data sets only partially overlap, their combination produces very useful information for improving the Arabidopsis genome annotation.

Arabidopsis↗

An optimized protocol for analysis of EST sequences.

The vast body of Expressed Sequence Tag (EST) data in the public databases provide an important resource for comparative and functional genomics studies and an invaluable tool for the annotation of genomic sequences. We have developed a rigorous protocol for reconstructing the sequences of transcribed genes from EST and gene sequence fragments. A key element in developing this protocol has been the evaluation of a number of sequence assembly programs to determine which most faithfully reproduce transcript sequences from EST data. The TIGR Gene Indices constructed using this protocol for human, mouse, rat and a variety of other plant and animal models have demonstrated their utility in a variety of applications and are freely available to the scientific research community.

Algorithms↗

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans↗

AgBase: a functional genomics resource for agriculture.

BACKGROUND: Many agricultural species and their pathogens have sequenced genomes and more are in progress. Agricultural species provide food, fiber, xenotransplant tissues, biopharmaceuticals and biomedical models. Moreover, many agricultural microorganisms are human zoonoses. However, systems biology from functional genomics data is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation and agricultural research communities are smaller with limited funding compared to many model organism communities. DESCRIPTION: To facilitate systems biology in these traditionally agricultural species we have established "AgBase", a curated, web-accessible, public resource http://www.agbase.msstate.edu for structural and functional annotation of agricultural genomes. The AgBase database includes a suite of computational tools to use GO annotations. We use standardized nomenclature following the Human Genome Organization Gene Nomenclature guidelines and are currently functionally annotating chicken, cow and sheep gene products using the Gene Ontology (GO). The computational tools we have developed accept and batch process data derived from different public databases (with different accession codes), return all existing GO annotations, provide a list of products without GO annotation, identify potential orthologs, model functional genomics data using GO and assist proteomics analysis of ESTs and EST assemblies. Our journal database helps prevent redundant manual GO curation. We encourage and publicly acknowledge GO annotations from researchers and provide a service for researchers interested in GO and analysis of functional genomics data. CONCLUSION: The AgBase database is the first database dedicated to functional genomics and systems biology analysis for agriculturally important species and their pathogens. We use experimental data to improve structural annotation of genomes and to functionally characterize gene products. AgBase is also directly relevant for researchers in fields as diverse as agricultural production, cancer biology, biopharmaceuticals, human health and evolutionary biology. Moreover, the experimental methods and bioinformatics tools we provide are widely applicable to many other species including model organisms.

Agriculture↗

Construction of a cosmid library of DNA replicated early in the S phase of normal human fibroblasts.

We constructed a subgenomic cosmid library of DNA replicated early in the S phase of normal human diploid fibroblasts. Cells were synchronized by release from confluence arrest and incubation in the presence of aphidicolin. Bromodeoxyuridine (BrdUrd) was added to aphidicolin-containing medium to label DNA replicated as cells entered S phase. Nuclear DNA was partially digested with Sau 3AI, and hybrid density DNA was separated in CsCl gradients. The purified early-replicating DNA was cloned into sCos1 cosmid vector. Clones were transferred individually into the wells of 96 microtiter plates (9,216 potential clones). Vigorous bacterial growth was detected in 8,742 of those wells. High-density colony hybridization filters (1, 536 clones/filter) were prepared from a set of replicas of the original plates. Bacteria remaining in the wells of replica plates were combined, mixed with freezing medium, and stored at -80 degrees C. These pooled stocks were analyzed by polymerase chain reaction to determine the presence of specific sequences in the library. Hybridization of high-density filters was used to identify the clones of interest, which were retrieved from the frozen cultures in the 96-well plates. In testing the library for the presence of 14 known early-replicating genes, we found sequences at or near 5 of them: APRT, beta-actin, beta-tubulin, c-myc, and HPRT. This library is a valuable resource for the isolation and analysis of certain DNA sequences replicated at the beginning of S phase, including potential origins of bidirectional replication.

Aphidicolin↗

Renal transcriptomes: segmental analysis of differential expression.

BACKGROUND/AIMS: Progress accomplished by complete genomes and cDNA-sequencing projects calls for methods that fully use these resources to study gene expression patterns in characterized cell populations. However, since the number of functional genes cannot be readily inferred from the genomic sequence, it is highly desirable to make use of methods enabling to study both known and unknown genes. METHODS: The method of serial analysis of gene expression provides short diagnostic cDNA tags without bias towards known genes. In addition, the frequency of each tag in the library conveys quantitative information on gene expression. A microassay was set-up to perform serial analysis of gene expression in minute samples such as those obtained by microdissecting nephron segments. RESULTS: Studies carried out in the thick ascending limb of Henle's loop and the collecting duct of the mouse kidney provided expression data for several thousand genes. Known markers were found appropriately enriched, and several of the thick ascending limb or collecting duct specific transcripts had no database match. CONCLUSIONS: The microassay for serial analysis of gene expression makes possible large-scale quantitative measurements of mRNA levels in nephron segments. The comprehensive picture generated by analyzing both known and unknown transcripts in defined cell populations should help to discover genes with dedicated functions.

Animals↗

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7&#xa0;Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (&#x223c;33.25 million SNPs and&#xa0;&#x223c;&#xa0;1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

The Aggregated Gut Viral Catalogue (AVrC): A unified resource for exploring the viral diversity of the human gut.

The growing interest in the role of the gut virome in human health and disease, has led to several recent large-scale viral catalogue projects mining human gut metagenomes each using varied computational tools and quality control criteria. Importantly, there has been to date no consistent comparison of these catalogues' quality, diversity, and overlap. In this project, we therefore systematically surveyed nine previously published human gut viral catalogues. While these catalogues collectively screened >40,000 human fecal metagenomes, 82% of the recovered 345,613 viral sequences were unique to one catalogue, highlighting limited redundancy between the ressources and suggesting the need for an aggregated resource bringing these viral sequences together. We further expanded these viral catalogues by mining 7,867 infant gut metagenomes from 12 large-scale infant studies collected in 9 different countries. From these datasets, we constructed the Aggregated Gut Viral Catalogue (AVrC), a unified modular resource containing 1,018,941 dereplicated viral sequences (449,859 species-level vOTUs). Using computational inference tools, annotations were obtained for each vOTU representative sequence quality, viral taxonomy, predicted viral lifestyle, and putative host. This project aims to facilitate the reuse of previously published viral catalogues by the research community and follows a modular framework to enable future expansions as novel data becomes available.

Humans↗

Generation, annotation, analysis and database integration of 16,500 white spruce EST clusters.

BACKGROUND: The sequencing and analysis of ESTs is for now the only practical approach for large-scale gene discovery and annotation in conifers because their very large genomes are unlikely to be sequenced in the near future. Our objective was to produce extensive collections of ESTs and cDNA clones to support manufacture of cDNA microarrays and gene discovery in white spruce (Picea glauca [Moench] Voss). RESULTS: We produced 16 cDNA libraries from different tissues and a variety of treatments, and partially sequenced 50,000 cDNA clones. High quality 3' and 5' reads were assembled into 16,578 consensus sequences, 45% of which represented full length inserts. Consensus sequences derived from 5' and 3' reads of the same cDNA clone were linked to define 14,471 transcripts. A large proportion (84%) of the spruce sequences matched a pine sequence, but only 68% of the spruce transcripts had homologs in Arabidopsis or rice. Nearly all the sequences that matched the Populus trichocarpa genome (the only sequenced tree genome) also matched rice or Arabidopsis genomes. We used several sequence similarity search approaches for assignment of putative functions, including blast searches against general and specialized databases (transcription factors, cell wall related proteins), Gene Ontology term assignation and Hidden Markov Model searches against PFAM protein families and domains. In total, 70% of the spruce transcripts displayed matches to proteins of known or unknown function in the Uniref100 database (blastx e-value < 1e-10). We identified multigenic families that appeared larger in spruce than in the Arabidopsis or rice genomes. Detailed analysis of translationally controlled tumour proteins and S-adenosylmethionine synthetase families confirmed a twofold size difference. Sequences and annotations were organized in a dedicated database, SpruceDB. Several search tools were developed to mine the data either based on their occurrence in the cDNA libraries or on functional annotations. CONCLUSION: This report illustrates specific approaches for large-scale gene discovery and annotation in an organism that is very distantly related to any of the fully sequenced genomes. The ArboreaSet sequences and cDNA clones represent a valuable resource for investigations ranging from plant comparative genomics to applied conifer genetics.

Arabidopsis↗

FLEXGene repository: from sequenced genomes to gene repositories for high-throughput functional biology and proteomics.

The vast amount of information generated by the human genome sequencing project and related projects has given rise to a new paradigm in experimental biology. This new paradigm invokes the experimentation and data analysis at genome-wide scales, as well as the generation of new technologies and resources that take full advantage of the available sequence information. The Institute of Proteomics at Harvard Medical School is building a comprehensive, characterized, arrayed and flexible gene repository that will allow full exploitation of the genomic information by enabling functional genomics as well as protein expression, purification and analysis at genome wide scale. The FLEXGene repository (Full Length EXpression-ready) will contain clones representing the complete set of open reading frames (ORFs) of different organisms including H. sapiens and several pathogens and model organisms. The clones are constructed using recombination-based cloning technology so that hundreds or thousands of coding regions can be transferred into any expression vector in a parallel and timely mode, allowing the broadest variety of experiments to be carried out.

Animals↗

Sequence analysis of the mitochondrial DNA control region of ciscoes (genus Coregonus): taxonomic implications for the Great Lakes species flock.

Sequence variation in the control region (D-loop) of the mitochondrial DNA (mtDNA) was examined to assess the genetic distinctiveness of the shortjaw cisco (Coregonus zenithicus). Individuals from within the Great Lakes Basin as well as inland lakes outside the basin were sampled. DNA fragments containing the entire D-loop were amplified by PCR from specimens of C. zenithicus and the related species C. artedi, C. hoyi, C. kiyi, and C. clupeaformis. DNA sequence analysis revealed high similarity within and among species and shared polymorphism for length variants. Based on this analysis, the shortjaw cisco is not genetically distinct from other cisco species.

Animals↗

Characterization of a 4-Mb region at chromosome 6q21 harboring a replicative senescence gene.

A 4-Mb region containing a senescence gene was defined at 6q21 by fluorescence in situ hybridization and deletion mapping after transfer of a normal human chromosome 6 to a BK virus-transformed mouse cell line. By screening three different yeast artificial chromosome (YAC) libraries, a YAC contig was constructed that covers the deleted region at 6q21. The contig is composed of 18 overlapping YACs with a size of 250-1800 kb and contains 3 CpG islands and 10 expressed sequence tags. By sequencing YACs and P1 artificial chromosomes, nine new sequence tagged sites and three new expressed sequence tags were detected that enrich the genetic resources of the region. The contig may also contain a fragile site, FRA6F, located close to a CpG island, which could be a landmark to localize the senescence gene. This YAC contig will be used to detect expressed sequences to clone and characterize the senescence gene at 6q21.

Animals↗

Recent segmental and gene duplications in the mouse genome.

BACKGROUND: The high quality of the mouse genome draft sequence and its associated annotations are an invaluable biological resource. Identifying recent duplications in the mouse genome, especially in regions containing genes, may highlight important events in recent murine evolution. In addition, detecting recent sequence duplications can reveal potentially problematic regions of the genome assembly. We use BLAST-based computational heuristics to identify large (>/= 5 kb) and recent (>/= 90% sequence identity) segmental duplications in the mouse genome sequence. Here we present a database of recently duplicated regions of the mouse genome found in the mouse genome sequencing consortium (MGSC) February 2002 and February 2003 assemblies. RESULTS: We determined that 33.6 Mb of 2,695 Mb (1.2%) of sequence from the February 2003 mouse genome sequence assembly is involved in recent segmental duplications, which is less than that observed in the human genome (around 3.5-5%). From this dataset, 8.9 Mb (26%) of the duplication content consisted of 'unmapped' chromosome sequence. Moreover, we suspect that an additional 18.5 Mb of sequence is involved in duplication artifacts arising from sequence misassignment errors in this genome assembly. By searching for genes that are located within these regions, we identified 675 genes that mapped to duplicated regions of the mouse genome. Sixteen of these genes appear to have been duplicated independently in the human genome. From our dataset we further characterized a 42 kb recent segmental duplication of Mater, a maternal-effect gene essential for embryogenesis in mice. CONCLUSION: Our results provide an initial analysis of the recently duplicated sequence and gene content of the mouse genome. Many of these duplicated loci, as well as regions identified to be involved in potential sequence misassignment errors, will require further mapping and sequencing to achieve accuracy. A Genome Browser database was set up to display the identified duplication content presented in this work. This data will also be relevant to the growing number of investigators who use the draft genome sequence for experimental design and analysis.

Animals↗

Database resources of the National Center for Biotechnology Information: update.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's website. NCBI resources include Entrez, PubMed, PubMed Central, LocusLink, the NCBI Taxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR, OrfFinder, Spidey, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes and related tools, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SARS Coronavirus Resource, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD) and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗