Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Large scale identification of genes involved in cell surface biosynthesis and architecture in Saccharomyces cerevisiae.

The sequenced yeast genome offers a unique resource for the analysis of eukaryotic cell function and enables genome-wide screens for genes involved in cellular processes. We have identified genes involved in cell surface assembly by screening transposon-mutagenized cells for altered sensitivity to calcofluor white, followed by supplementary screens to further characterize mutant phenotypes. The mutated genes were directly retrieved from genomic DNA and then matched uniquely to a gene in the yeast genome database. Eighty-two genes with apparent perturbation of the cell surface were identified, with mutations in 65 of them displaying at least one further cell surface phenotype in addition to their modified sensitivity to calcofluor. Fifty of these genes were previously known, 17 encoded proteins whose function could be anticipated through sequence homology or previously recognized phenotypes and 15 genes had no previously known phenotype.

Cell Membrane↗

Full-length cDNAs: more than just reaching the ends.

The development of functional genomic resources is essential to understand and utilize information generated from genome sequencing projects. Central to the development of this technology is the creation of high-quality cDNA resources and improved technologies for analyzing coding and noncoding mRNA sequences. The isolation and mapping of cDNAs is an entrée to characterizing the information that is of significant biological relevance in the genome of an organism. However, a bottleneck is often encountered when attempting to bring to full-length (or at least full-coding) a number of incomplete cDNAs in parallel, since this involves the nonsystematic, time consuming, and labor-intensive iterative screening of a number of cDNA libraries of variable quality and/or directed strategies to process individual clones (e.g., 5' rapid amplification of cDNA ends). Here, we review the current state of the art in cDNA library generation, as well as present an analysis of the different steps involved in cDNA library generation.

Automation↗

The carbohydrate sequence markup language (CabosML): an XML description of carbohydrate structures.

UNLABELLED: Bioinformatics resources for glycomics are very poor as compared with those for genomics and proteomics. The complexity of carbohydrate sequences makes it difficult to define a common language to represent them, and the development of bioinformatics tools for glycomics has not progressed. In this study, we developed a carbohydrate sequence markup language (CabosML), an XML description of carbohydrate structures. AVAILABILITY: The language definition (XML Schema) and an experimental database of carbohydrate structures using an XML database management system are available at http://www.phoenix.hydra.mki.co.jp/CabosDemo.html CONTACT: kikuchi@hydra.mki.co.jp.

Carbohydrate Sequence↗

Targeted Next-Generation Sequencing for Improved Clinical Outcomes in People Living With Rare Diseases in Global South: Protocol for a Systematic Review and Meta-Synthesis.

BACKGROUND: Rare diseases affect many individuals and pose major challenges in diagnosis and treatment, especially in Global South countries where health care resources are limited. Targeted next-generation sequencing (NGS) has significantly advanced diagnostic accuracy and clinical care for rare diseases globally; however, its implementation and impact within the Global South context remain insufficiently studied. OBJECTIVE: This study aims to evaluate the use, clinical benefits, challenges, and implementation outcomes of targeted NGS for diagnosing and managing rare diseases in Global South populations. Specifically, it seeks to quantify the diagnostic yield of NGS, examine its influence on subsequent clinical decision-making, and identify principal barriers to, and facilitators of, the implementation of targeted NGS approaches in these contexts. METHODS: This protocol follows the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. We will systematically search PubMed, Scopus, and Web of Science for studies published between 2005 and 2025 that report on the use of targeted NGS in Global South population with rare diseases. Two reviewers will independently perform study selection, data extraction, quality assessment, and evaluation of risk of bias by using QUADAS-2 for diagnostic accuracy studies and the risk of bias assessment tool for nonrandomized studies. Meta-analyses will be conducted to estimate pooled outcomes for diagnostic yield, with heterogeneity assessed using random effects models. Heterogeneity will be further examined through visual inspection of forest plots and by evaluating the chi-square test and I² statistic. RESULTS: The protocol has been registered with PROSPERO (CRD420251078455). Database search or screening, data extraction, and data synthesis are planned to commence in June 2026 and conclude by September 2026. Study findings will synthesize the diagnostic yield, clinical impact, and contextual determinants influencing the implementation of targeted NGS in Global South health care settings. CONCLUSIONS: This review will provide evidence on the application, advantages, limitations, and clinical outcomes of targeted NGS for individuals affected by rare diseases in countries of the Global South. The finding will identify priorities for capacity strengthening, policy development, and future genomic research. TRIAL REGISTRATION: PROSPERO CRD420251078455; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251078455. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/85150.

Rare Diseases↗

Database resources of the National Center for Biotechnology.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, PubMed, PubMed Central (PMC), LocusLink, the NCBITaxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR (e-PCR), Open Reading Frame (ORF) Finder, References Sequence (RefSeq), UniGene, HomoloGene, ProtEST, Database of Single Nucleotide Polymorphisms (dbSNP), Human/Mouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes and related tools, the Map Viewer, Model Maker (MM), Evidence Viewer (EV), Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD), and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

Analysis of sequence-tagged-connector strategies for DNA sequencing.

The BAC-end sequencing, or sequence-tagged-connector (STC), approach to genome sequencing involves sequencing the ends of BAC inserts to scatter sequence tags (STCs) randomly across the genome. Once any BAC or other large segment of DNA is sequenced to completion by conventional shotgun approaches, these STC tags can be used to identify a minimum tiling path of BAC clones overlapping the nucleation sequence for sequence extension. Here, we explore the properties of STC-sequencing strategies within a mathematical model of a random target with homologous repeats and imperfect sequencing technology to understand the consequences of varying various parameters on the incidence of problem clones and the cost of the sequencing project. Problem clones are defined as clones for which either (A) there is no identifiable overlapping STC to extend the sequence in a particular direction or (B) the identified STC with minimum overlap comes from a nonoverlapping clone, either owing to random false matches or repeat-family homology. Based on the minimum overlap, we estimate the number of clones to be entirely sequenced and, then, using cost estimates, identify the decision rule (the degree of sequence similarity required before a match is declared between an STC and a clone) to minimize overall sequencing cost. A method to optimize the overlap decision rule is highly desirable, because both the total cost and the number of problem clones are shown to be highly sensitive to this choice. For a target of 3 Gb containing approximately 800 Mb of repeats with 85%-90% identity, we expect <10 problem clones with 15 times coverage by 150-kb clones. We derive the optimal redundancy and insert sizes of clone libraries for sequencing genomes of various sizes, from microbial to human. We estimate that establishing the resource of STCs as a means of identifying minimally overlapping clones represents only 1%-3% of the total cost of sequencing the human genome, and, up to a point of diminishing returns, a larger STC resource is associated with a smaller total sequencing cost.

Genome, Human↗

Child safety seat use for infants with Pierre Robin sequence.

OBJECTIVE: To determine what child restraints would accommodate infants with Pierre Robin sequence who often require special attention in motor vehicle travel since microagnathia usually requires a prone position to keep the infant's airway open. RESEARCH DESIGN: Dynamic testing and clinical trial. SETTING: An Indiana children's hospital providing primary and tertiary care. PATIENTS: Four patients with Pierre Robin sequence are described to illustrate use of the modified infant car seat and the appropriateness of the car bed restraints for meeting requirements for prone positioning during travel. SELECTION PROCEDURES: Convenience sample. INTERVENTION: Selected restraints were loaned to families through a clinical setting until the patient was able to use a conventional child restraint. MEASUREMENTS AND RESULTS: Three child restraint systems were determined to accommodate the prone position necessary to keep the airway open for children with Pierre Robin sequence. Dynamic crash testing demonstrated the crashworthiness of an infant car seat modified to allow for prone positioning. Through a clinical trial, two car bed restraints were also found to provide safe prone positioning of infants. CONCLUSIONS: To enable safe transportation for infants with Pierre Robin sequence, health care providers can direct parents to appropriate resources for travel and can monitor the airway and oxygenation of the infant with Pierre Robin sequence before hospital discharge.

Equipment Design↗

Protein families and TRIBES in genome sequence space.

Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.

Algorithms↗

A comprehensive catalog of human KRAB-associated zinc finger genes: insights into the evolutionary history of a large family of transcriptional repressors.

Krüppel-type zinc finger (ZNF) motifs are prevalent components of transcription factor proteins in all eukaryotes. KRAB-ZNF proteins, in which a potent repressor domain is attached to a tandem array of DNA-binding zinc-finger motifs, are specific to tetrapod vertebrates and represent the largest class of ZNF proteins in mammals. To define the full repertoire of human KRAB-ZNF proteins, we searched the genome sequence for key motifs and then constructed and manually curated gene models incorporating those sequences. The resulting gene catalog contains 423 KRAB-ZNF protein-coding loci, yielding alternative transcripts that altogether predict at least 742 structurally distinct proteins. Active rounds of segmental duplication, involving single genes or larger regions and including both tandem and distributed duplication events, have driven the expansion of this mammalian gene family. Comparisons between the human genes and ZNF loci mined from the draft mouse, dog, and chimpanzee genomes not only identified 103 KRAB-ZNF genes that are conserved in mammals but also highlighted a substantial level of lineage-specific change; at least 136 KRAB-ZNF coding genes are primate specific, including many recent duplicates. KRAB-ZNF genes are widely expressed and clustered genes are typically not coregulated, indicating that paralogs have evolved to fill roles in many different biological processes. To facilitate further study, we have developed a Web-based public resource with access to gene models, sequences, and other data, including visualization tools to provide genomic context and interaction with other public data sets.

Computational Biology↗

Characterization of swine leukocyte antigen polymorphism by sequence-based and PCR-SSP methods in Meishan pigs.

Resource herds of swine leukocyte antigen (SLA)-characterized pigs are an important tool for the study of immune responses, disease resistance, and production traits. They are also valuable large animal models for biomedical research, such as transplantation. The Meishan breed of pig is an economically significant breed that is available at several research institutions in the United States. We have characterized the SLA polymorphism of the breeding stock in the herd maintained at the University of Illinois and developed a simple assay to SLA type individuals within that herd. We have used a reverse transcription-polymerase chain reaction (RT-PCR)-based SLA typing method to clone and DNA sequence 19 SLA alleles at three SLA class Ia (SLA-1, SLA-2, and SLA-3) and two SLA class II (SLA-DRB1 and SLA-DQB1) loci. Based on this sequence information, a rapid SLA typing assay was developed to discriminate each allele using PCR with sequence-specific primers (PCR-SSP). Using this method, we were able to characterize the entire Meishan breeding stock and identify four SLA haplotypes present in the herd. The combination of SLA typing by cloning and DNA sequencing with PCR-SSP is therefore a valuable tool for the characterization of SLA alleles and haplotypes in resource herds of pigs.

Animals↗

A human chromosome 7 yeast artificial chromosome (YAC) resource: construction, characterization, and screening.

The paradigm of sequence-tagged site (STS)-content mapping involves the systematic assignment of STSs to individual cloned DNA segments. To date, yeast artificial chromosomes (YACs) represent the most commonly employed cloning system for constructing STS maps of large genomic intervals, such as whole human chromosomes. For developing a complete YAC-based STS-content map of human chromosome 7, we wished to utilize a limited set of YAC clones that were highly enriched for chromosome 7 DNA. Toward that end, we have assembled a human chromosome 7 YAC resource that consists of three major components: (1) a newly constructed library derived from a human-hamster hybrid cell line containing chromosome 7 as its only human DNA; (2) a chromosome 7-enriched sublibrary derived from the CEPH mega-YAC collection by Alu-polymerase chain reaction (Alu-PCR)-based hybridization; and (3) a set of YACs isolated from several total genomic libraries by screening for > 125 chromosome 7 STSs. In particular, the hybrid cell line-derived YACs, which comprise the majority of the clones in the resource, have a relatively low chimera frequency (10-20%) based on mapping isolated insert ends to panels of human-hamster hybrid cell lines and analyzing individual clones by fluorescence in situ hybridization. An efficient strategy for polymerase chain reaction (PCR)-based screening of this YAC resource, which totals 4190 clones, has been developed and utilized to identify corresponding YACs for > 600 STSs. The results of this initial screening effort indicate that the human chromosome 7 YAC resource provides an average of 6.9 positive clones per STS, a level of redundancy that should support the assembly of large YAC contigs and the construction of a high-resolution STS-content map of the chromosome.

Animals↗

Integration of cytogenetic landmarks into the draft sequence of the human genome.

We have placed 7,600 cytogenetically defined landmarks on the draft sequence of the human genome to help with the characterization of genes altered by gross chromosomal aberrations that cause human disease. The landmarks are large-insert clones mapped to chromosome bands by fluorescence in situ hybridization. Each clone contains a sequence tag that is positioned on the genomic sequence. This genome-wide set of sequence-anchored clones allows structural and functional analyses of the genome. This resource represents the first comprehensive integration of cytogenetic, radiation hybrid, linkage and sequence maps of the human genome; provides an independent validation of the sequence map and framework for contig order and orientation; surveys the genome for large-scale duplications, which are likely to require special attention during sequence assembly; and allows a stringent assessment of sequence differences between the dark and light bands of chromosomes. It also provides insight into large-scale chromatin structure and the evolution of chromosomes and gene families and will accelerate our understanding of the molecular bases of human disease and cancer.

Chromosome Aberrations↗

Green sperm. Identification of male gamete promoters in Arabidopsis.

Previously, in an effort to better understand the male contribution to fertilization, we completed a maize (Zea mays) sperm expressed sequence tag project. Here, we used this resource to identify promoters that would direct gene expression in sperm cells. We used reverse transcription-polymerase chain reaction to identify probable sperm-specific transcripts in maize and then identified their best sequence matches in the Arabidopsis (Arabidopsis thaliana) genome. We tested five different Arabidopsis promoters for cell specificity, using an enhanced green fluorescent protein reporter gene. In pollen, the AtGEX1 (At5g55490) promoter is active in the sperm cells and not in the progenitor generative cell or in the vegetative cell, but it is also active in ovules, roots, and guard cells. The AtGEX2 (At5g49150) promoter is active only in the sperm cells and in the progenitor generative cell, but not in the vegetative cell or in other tissues. A third promoter, AtVEX1 (At5g62850) [corrected] was active in the vegetative cell during the later stages of pollen development; the other promoters tested (At1g66770 and At1g73350) did not function in pollen. Comparisons among GEX1 and GEX2 homologs from maize, rice (Oryza sativa), Arabidopsis, and poplar (Populus trichocarpa) revealed a core binding site for Dof transcription factors. The AtGEX1 and AtGEX2 promoters will be useful for manipulating gene expression in sperm cells, for localization and functional analyses of sperm proteins, and for imaging of sperm dynamics as they are transported in the pollen tube to the embryo sac.

Arabidopsis↗

Pseudogenes in metazoa: origin and features.

The complete genome sequences with their annotations are a considerable resource in biology, particularly in understanding the global structure of the genetic material at the molecular level. The reason why some eukaryotic genomes contain large quantities of apparently unnecessary DNA, namely pseudogenes, while others seem to invest in more efficient thinning processes or are equipped with protection systems against parasitic elements still remains a mystery. Several genome-wide surveys have been undertaken to identify pseudogenes in the completely sequenced genome, bringing to light some differences both in their amount and distribution. Since pseudogenes are important resources in evolutionary and comparative genomics - as 'molecular fossils' - in this paper, a survey on the origins, features, abundance and localisation of the different pseudogenes is reported. As an example of genes producing processed pseudogenes, some experimental data obtained in the authors' laboratories from the study of a nuclear gene coding for the mitochondrial transcription factor A (mtTFA), a key regulator of mitochondrial biogenesis, are also reported.

Animals↗

The dog genome: survey sequencing and comparative analysis.

A survey of the dog genome sequence (6.22 million sequence reads; 1.5x coverage) demonstrates the power of sample sequencing for comparative analysis of mammalian genomes and the generation of species-specific resources. More than 650 million base pairs (>25%) of dog sequence align uniquely to the human genome, including fragments of putative orthologs for 18,473 of 24,567 annotated human genes. Mutation rates, conserved synteny, repeat content, and phylogeny can be compared among human, mouse, and dog. A variety of polymorphic elements are identified that will be valuable for mapping the genetic basis of diseases and traits in the dog.

Animals↗

Database resources of the National Center for Biotechnology Information: 2002 update.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources that operate on the data in GenBank and a variety of other biological data made available through NCBI's web site. NCBI data retrieval resources include Entrez, PubMed, LocusLink and the Taxonomy Browser. Data analysis resources include BLAST, Electronic PCR, OrfFinder, RefSeq, UniGene, HomoloGene, Database of Single Nucleotide Polymorphisms (dbSNP), Human Genome Sequencing, Human MapViewer, Human inverted exclamation markVMouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB) and the Conserved Domain Database (CDD). Augmenting many of the web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at http://www.ncbi.nlm.nih.gov.

Amino Acid Sequence↗

Orphan transcripts in Arabidopsis thaliana: identification of several hundred previously unrecognized genes.

Expressed sequence tags (ESTs) represent a huge resource for the discovery of previously unknown genetic information and functional genome assignment. In this study we screened a collection of 178 292 ESTs from Arabidopsis thaliana by testing them against previously annotated genes of the Arabidopsis genome. We identified several hundreds of new transcripts that match the Arabidopsis genome at so far unassigned loci. The transcriptional activity of these loci was independently confirmed by comparison with the Salk Whole Genome Array Data. To a large extent, the newly identified transcriptionally active genomic regions do not encode 'classic' proteins, but instead generate non-coding RNAs and/or small peptide-coding RNAs of presently unknown biological function. More than 560 transcripts identified in this study are not represented by the Affymetrix GeneChip arrays currently widely used for expression profiling in A. thaliana. Our data strongly support the hypothesis that numerous previously unknown genes exist in the Arabidopsis genome.

Arabidopsis↗

An EST and STS-based YAC contig map of human chromosome 9q22.3.

We have isolated 48 yeast artificial chromosome (YAC) clones from a 4 cM/27 cR region of human chromosome 9q22.3 encompassed by the markers cen-D9S196-D9S173-tel. Within this region, we have assembled a 4.3-Mb YAC contig across the interval cen-FACC-D9S173-tel containing 42 clones. As a first step toward completing the detailed transcription map of the region, we have mapped 9 gene sequences and 10 expressed sequence tags. Fifteen polymorphic microsatellite repeat markers and 17 novel sequence-tagged sites from the region are also described. The mapping of polymorphic simple tandem repeat markers has permitted the integration of existing genetic and physical maps of the region. Together these maps provide a valuable resource for fine structure mapping and DNA sequencing across the region as well as for the identification of disease gene loci and the isolation of novel coding sequences.

Base Sequence↗