Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

IRIS: a database surveying known human immune system genes.

We have compiled an online database of known human defense genes: the Immunogenetic Related Information Source (IRIS). As of October 1, 2004, there are 1562 immune genes recorded in IRIS, representing 7% of the human genome. This resource contains searchable information including chromosomal location, sequence data, and a curated functional annotation for each entry. We used IRIS as a basis for analyzing the composition and characteristics of the immune genome, such as gene clustering, polymorphism, and relationship to disease. High protein sequence similarity correlated inversely with distance between immune genes, consistent with clustering of duplicated loci. We also found that, even though some immune genes exhibit high levels of polymorphism, such as MHC class I, the range of levels of polymorphism in immune genes is similar to that of nonimmune genes. Approximately 20% of immune genes have a known disease association. IRIS is available online at .

Databases, Genetic↗

Human epidermal differentiation complex in a single 2.5 Mbp long continuum of overlapping DNA cloned in bacteria integrating physical and transcript maps.

Terminal differentiation of keratinocytes involves the sequential expression of several major proteins which can be identified in distinct cellular layers within the mammalian epidermis and are characteristic for the maturation state of the keratinocyte. Many of the corresponding genes are clustered in one specific human chromosomal region 1q21. It is rare in the genome to find in such close proximity the genes belonging to at least three structurally different families, yet sharing spatial and temporal expression specificity, as well as interdependent functional features. This DNA segment, termed the epidermal differentiation complex, contains 27 genes, 14 of which are specifically expressed during calcium-dependent terminal differentiation of keratinocytes (the majority being structural protein precursors of the cornified envelope) and the other 13 belong to the S100 family of calcium binding proteins with possible signal transduction roles in the differentiation of epidermis and other tissues. In order to provide a bacterial clone resource that will enable further studies of genomic structure, transcriptional regulation, function and evolution of the epidermal differentiation complex, as well as the identification of novel genes, we have constructed a single 2.45 Mbp long continuum of genomic DNA cloned as 45 p1 artificial chromosomes, three bacterial artificial chromosomes, and 34 cosmid clones. The map encompasses all of the 27 genes so far assigned to the epidermal differentiation complex, and integrates the physical localization of these genes at a high resolution on a complete NotI and SalI, and a partial EcoRI restriction map. This map will be the starting resource for the large-scale genomic sequencing of this region by The Sanger Center, Hinxton, U.K.

Bacteria↗

The clinical promise of mass spectrometry-based single-cell proteomics: from bedside to bench.

INTRODUCTION: Single-cell proteomics (SCP) is entering into a transformative phase, moving beyond technically demanding benchmarking studies toward robust and reproducible workflows capable of quantifying thousands of proteins per cell. These advances highlight SCP's potential to address clinically relevant questions by resolving cellular and pathological heterogeneity that remains obscured in bulk proteomics. AREAS COVERED: This review discusses current advances, challenges, and clinical applications of SCP based on literature identified through searches in major scientific databases. Many clinically relevant samples remain underexplored in SCP studies, in part because their application requires careful evaluation of pre-analytical variables that can strongly influence proteomic readouts. Current SCP methodologies vary according to sample type, experimental conditions, and available resources. Compared with single-cell RNA sequencing, SCP remains limited in cellular throughput, making it challenging to define optimal sample sizes and to reliably detect both abundant and rare cell populations. These limitations also make dataset integration difficult, as reduced cellular coverage and sampling depth increase data sparsity. Moreover, implementing quality control strategies across sequential SCP experiments is essential to ensure data robustness, comparability, and accurate biological interpretation. EXPERT OPINION: Applying SCP to clinical samples advances our understanding of biological complexity and holds potential to drive progress in translational and precision medicine.

Humans↗

Evidence standards in experimental and inferential INSDC Third Party Annotation data.

The Third Party Annotation (TPA) project collects and presents high-quality annotation of nucleotide sequence. Annotation is submitted by researchers who have not themselves generated novel nucleotide sequence. In its first few years, the resource has proven to be popular with submitters from a range of biological research areas. Central to the project is the requirement for high-quality data, resulting from experimental and inferred analysis discussed in peer-reviewed publications. The data are divided into two tiers: those with experimental evidence and those with inferential evidence. Standards for TPA are detailed and illustrated with the aid of case studies.

Animals↗

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning↗

Drosophila DNase I footprint database: a systematic genome annotation of transcription factor binding sites in the fruitfly, Drosophila melanogaster.

UNLABELLED: Despite increasing numbers of computational tools developed to predict cis-regulatory sequences, the availability of high-quality datasets of transcription factor binding sites limits advances in the bioinformatics of gene regulation. Here we present such a dataset based on a systematic literature curation and genome annotation of DNase I footprints for the fruitfly, Drosophila melanogaster. Using the experimental results of 201 primary references, we annotated 1367 binding sites from 87 transcription factors and 101 target genes in the D.melanogaster genome sequence. These data will provide a rich resource for future bioinformatics analyses of transcriptional regulation in Drosophila such as constructing motif models, training cis-regulatory module detectors, benchmarking alignment tools and continued text mining of the extensive literature on transcriptional regulation in this important model organism. AVAILABILITY: http://www.flyreg.org/ CONTACT: cbergman@gen.cam.ac.uk.

Binding Sites↗

Genomic database resources for Dictyostelium discoideum.

Dictyostelium is an attractive model system for the study of mechanisms basic to cellular function or complex multicellular developmental processes. Recent advances in Dictyostelium genomics have generated a wide spectrum of resources. However, much of the current genomic sequence information is still not currently available through GenBank or related databases. Thus, many investigators are unaware that extensive sequence data from Dictyostelium has been compiled, or of its availability and access. Here, we discuss progress in Dictyostelium genomics and gene annotation, and highlight the primary portals for sequence access, manipulation and analysis (http://genome.imb-jena.de/dictyostelium/; http://dictygenome.bcm.tmc.edu/; http://www.sanger. ac.uk/Projects/D_discoideum/; http://www.csm.biol. tsukuba.ac.jp/cDNAproject.html).

Amino Acid Sequence↗

The Vertebrate Genome Annotation (Vega) database.

The Vertebrate Genome Annotation (Vega) database (http://vega.sanger.ac.uk) has been designed to be a community resource for browsing manual annotation of finished sequences from a variety of vertebrate genomes. Its core database is based on an Ensembl-style schema, extended to incorporate curation-specific metadata. In collaboration with the genome sequencing centres, Vega attempts to present consistent high-quality annotation of the published human chromosome sequences. In addition, it is also possible to view various finished regions from other vertebrates, including mouse and zebrafish. Vega displays only manually annotated gene structures built using transcriptional evidence, which can be examined in the browser. Attempts have been made to standardize the annotation procedure across each vertebrate genome, which should aid comparative analysis of orthologues across the different finished regions.

Animals↗

The Ensembl computing architecture.

Ensembl is a software project to automatically annotate large eukaryotic genomes and release them freely into the public domain. The project currently automatically annotates 10 complete genomes. This makes very large demands on compute resources, due to the vast number of sequence comparisons that need to be executed. To circumvent the financial outlay often associated with classical supercomputing environments, farms of multiple, lower-cost machines have now become the norm and have been deployed successfully with this project. The architecture and design of farms containing hundreds of compute nodes is complex and nontrivial to implement. This study will define and explain some of the essential elements to consider when designing such systems. Server architecture and network infrastructure are discussed with a particular emphasis on solutions that worked and those that did not (often with fairly spectacular consequences). The aim of the study is to give the reader, who may be implementing a large-scale biocompute project, an insight into some of the pitfalls that may be waiting ahead.

Computational Biology↗

Implementation of a managed care model in an acute care setting.

The center of attention in recent nursing literature has been on the evolution of outcome-based care. This concept has emerged as an array of models that includes patient-focused care, case management, and managed care. All these approaches to redesigning patient care have a basic goal: to deliver quality care in a timely manner by using appropriate resources. The program at St. Margaret Mercy Healthcare Centers focused on the staff nurse as the multidisciplinary team member who would be responsible for coordinating patient care. This type of model, in which the pathway is initiated at the staff level, is called managed care. Managed care, as defined by Zander, is unit-based care that is organized to achieve specific patient outcomes within fiscally responsible time frames or lengths of stay (LOS) while using resources that are appropriate in amount and sequence to the specific case type and to the individual patient (1988a).

Clinical Protocols↗

Analyses of Five gallinacin genes and the Salmonella enterica serovar Enteritidis response in poultry.

Gallinacins in poultry are functional equivalents of mammalian beta-defensins, which constitute an integral component of the innate immune system. Salmonella enterica serovar Enteritidis is a gram-negative bacterium that negatively affects both human and animal health. To analyze the association of genetic variations of the gallinacin genes with the phenotypic response to S. enterica serovar Enteritidis, an F1 population of chickens was created by crossing four outbred broiler sires to dams of two highly inbred lines. The F1 chicks were evaluated for bacterial colonization after pathogenic S. enterica serovar Enteritidis inoculation and for circulating antibody levels after inoculation with S. enterica serovar Enteritidis bacterin vaccine. Five candidate genes were studied, including gallinacins 2, 3, 4, 5, and 7. Gene fragments were sequenced from the founder individuals of the resource population, and a mean of 13.2 single-nucleotide polymorphisms (SNP) per kilobase was identified. One allele-defining SNP per gene was utilized to test for statistical associations of sire alleles with progeny response to S. enterica serovar Enteritidis. Among the five gallinacin genes evaluated, the Gal3 and Gal7 SNPs in broiler sires were found to be associated with antibody production after S. enterica serovar Enteritidis vaccination. Utilization of these SNPs as molecular markers for the response to S. enterica serovar Enteritidis may result in the enhancement of the immune response in poultry.

Animals↗

Semantic similarity measures as tools for exploring the gene ontology.

Many bioinformatics resources hold data in the form of sequences. Often this sequence data is associated with a large amount of annotation. In many cases this data has been hard to model, and has been represented as scientific natural language, which is not readily computationally amenable. The development of the Gene Ontology provides us with a more accessible representation of some of this data. However it is not clear how this data can best be searched, or queried. Recently we have adapted information content based measures for use with the Gene Ontology (GO). In this paper we present detailed investigation of the properties of these measures, and examine various properties of GO, which may have implications for its future design.

Classification↗

Resources for genetic variation studies.

The rapid growth of genome-wide diversity databases, as well as ongoing large-scale resequencing projects targeting genes and other functional components of our genome, provide valuable resources of natural variation at the DNA sequence level. In this review, we briefly summarize the wealth of data on DNA polymorphisms in humans, the distribution of this diversity in the genome as well as among individuals, and the consequence of recombination on its organization. These data provide a set of powerful tools that can be used to better understand inherited phenotypic variation in humans. We discuss the implications for the design of studies investigating correlations between genotypes and phenotypes, both at the fundamental level of genome function and regulation, and for the mapping of disease genes.

Biomedical Research↗

YAC/BAC contig spanning the MHC class III region of cattle.

A contig of the class III region of the bovine major histocompatibility complex (MHC) was established from bacterial and yeast artificial chromosomes using PCR and BAC-end sequencing. The marker content of individual clones was determined by gene and BAC-end specific PCR, and the location of genes and BAC-ends was confirmed analyzing somatic hybrid cells. A comparative analysis indicated that the content and order of MHC class III genes is strongly conserved between cattle and other mammalian species. Fluorescence in situ hybridization localized the bovine class III region to BTA23q21-->q22. The results show that the collection of sequenced BAC-ends is a powerful resource for generating high-resolution comparative chromosome maps.

Animals↗

Point mutation analysis of archived cytogenetic slide DNA.

Archived Giemsa-stained cytogenetic slide repositories represent valuable DNA resources for medical, scientific, and forensic studies. Sequencing readily identified a Charcot-Marie-Tooth disease point mutation in a 209-bp PCR amplified product. With optimal PCR primers and amplification conditions, our protocol quickly and reliably isolated sufficient DNA for at least 12 independent PCR amplification reactions for forensic and medical applications from single slides up to 5 years old.

Alleles↗

New technologies and DNA resources for high throughput biology.

The rapid increase in DNA sequencing information is opening up new opportunities in genetics. The current methods for processing and analysing genetic data are, however, slow and labour intensive. The next wave of genetic analysis will rely on the analysis of DNA variation from large population based cohorts. These studies will provide important new data on population and disease genetics and have the potential to make a significant impact on our current healthcare practices. In order for these studies to deliver, we need to develop a new generation of ultra-rapid DNA technologies which will allow us to generate, capture and efficiently exploit these new data. This chapter describes the recent advances in DNA sequencing and genotyping technologies that will lead to 100-1000-fold increases in our ability to produce the DNA data we need to explore and exploit the new genetic opportunities to the full.

DNA Mutational Analysis↗

Cataloging the relationships between proteins: a review of interaction databases.

By organizing and making widely accessible the increasing amounts of data from high-throughput analyses, protein interaction databases have become an integral resource for the biological community in relating sequence data with higher-order function. To provide a sense of the use and applicability of these databases, we describe each of the major comprehensive interaction databases as well as some of the more specialized ones. Content description, search/browse functionalities, and data presentation are discussed. A succinct explanation of database contents helps the user quickly identify whether the database contains applicable information to their research interest. Broad levels of search/browse functions as well as descriptions/examples allow users to quickly find and access pertinent data. At this point, clear presentation of search results as well as the primary content is necessary. Many databases display information graphically or divided into smaller digestible parts over a number of tabbed/linked pages. In addition, cross-linking between the databases promotes interconnectivity of the data and is an added layer of relational data for the user. Overall, although these protein interaction databases are under continual improvement, their current state shows that much time and effort has gone into organizing and presenting these large sets of data-describing protein interactions.

Database Management Systems↗

[Coronary angioplasty in subgroups at risk: patients treated with bypass].

Percutaneous coronary intervention represents an established method to obtain revascularization in patients suffering from obstructive coronary artery disease. Despite the predictability of procedural results and the favorable clinical outcome shown in large series, there are still areas in which the clinical benefits of intervention are less impressive and a lot remains to be done to improve both procedural and clinical outcomes. Among these areas, one may well include the subgroup of patients with previous bypass surgery, who have been consistently shown to be affected by a high rate of periprocedural complications as well as late recurrence. The risk profile of these patients is examined under two main perspectives: the burden of more severe baseline clinical conditions (age, ventricular function, severity of coronary disease) and the negative impact of graft atherosclerosis. The basic assumption of this article is that a variable combination of these characteristics identifies subsets with increasing risk of complications and/or recurrence. For this reason, results of percutaneous revascularization in these patients may still represent a technical as well as a clinical challenge. In particular, the long debated relationship between composition of atherosclerotic plaque in saphenous vein graft, distal embolization, periprocedural myocardial damage and early and late adverse events represents a negative sequence that currently available pharmacologic and interventional resources cannot consistently antagonize. Added to this, are the unresolved issues of diffuse degeneration and chronic total occlusion of saphenous vein grafts, in which no therapeutic approach, alone or in combination, has substantially modified the poor outcome of these lesions. The use of glycoprotein IIb/IIIa receptor antagonists is also discussed, in the light of available data derived from large clinical trials, casting doubts on the efficacy of these otherwise essential pharmacologic agents. Lastly, the setting of acute myocardial infarction represents the clinical scenario in which the adverse effects of the combination of clinical and angiographic characteristics are clearly appreciated. Both coronary intervention and cardiac surgery represent fields of rapidly growing knowledge and technology. It is likely that in the near future we will witness major changes in the clinical management of these patients, thanks to the increasing utilization of arterial conduits, the widespread use of local drug delivery, the availability of new percutaneous devices, and the development of integrated pharmacologic and mechanical revascularization strategies.

Angioplasty, Balloon, Coronary↗