Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “high-throughput screen”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

The diagnostic potential of combined quantitative polymerase chain reaction and next-generation sequencing using the same primers for periprosthetic joint infection.

Next-generation sequencing (NGS) enables the detection of specific pathogens unidentifiable by conventional cultures, but its application in orthopedics remains inconsistent due to background contamination and irreproducible findings. This study evaluated the diagnostic performance of a novel workflow combining broad-range 16S rRNA gene quantitative PCR (qPCR) screening with downstream NGS, focusing on bacterial biomass thresholds. The qPCR assay demonstrated excellent intrarater reliability, with an intraclass correlation coefficient (ICC) of 0.961 (95% confidence interval, 0.881 to 0.997). Based on serially diluted positive controls, a quantitative threshold of 10⁵ CFU/mL was established as the minimum concentration required for the consistent detection of fastidious taxa, such as Escherichia coli. When evaluated against conventional cultures using 95 sonicate fluid and 276 pre/intraoperative tissue samples, the qPCR assay achieved a sensitivity of 80% and a specificity of 72%. Subsequent NGS sequencing of 26 clinical samples and 9 controls showed concordance in 4 of 6 culture-positive infected cases with NGS taxonomy, whereas the remaining discrepancies were likely attributable to culture-based phenotypic misidentification. Notably, among the qPCR-positive cases, three were culture-negative, including two hip prosthesis loosening cases exhibiting polymicrobial profiles, and one post-traumatic osteoarthritis case harboring low-level Staphylococcus. Crucially, this post-traumatic patient developed delayed periprosthetic joint infection (PJI) 2 years post-surgery, with cultures identifying Staphylococcus previously detected by the initial NGS analysis. Integrating qPCR screening with targeted NGS effectively refines pathogen identification, filters environmental artifacts, and overcomes the diagnostic limitations of culture-negative infections in orthopedic practice.IMPORTANCENext-generation sequencing (NGS) enables the detection of specific pathogens in clinical samples that are not identifiable by conventional methods. However, NGS applications in orthopedics have not been quantitatively evaluated, and findings have been inconsistent owing to contaminants and the presence of non-credible causative organisms. These factors primarily stem from the failure to evaluate low-biomass samples and the absence of proper controls, such as negative controls or mock community DNA samples. This study demonstrates that interpreting results from low-biomass samples requires careful consideration because NGS relies on relative bacterial abundances; distinguishing likely pathogens from contaminants is particularly challenging when bacterial loads are low. We demonstrated that combining NGS with quantitative PCR (qPCR) and applying a Cq cutoff can reduce false positives.

Humans↗

Genotypic Analysis and Clinical Findings of Sapovirus-Associated Acute Gastroenteritis in Mie Prefecture, Japan, 2010-2022.

Sapovirus (SaV) is one of the major viruses causing acute gastroenteritis. Of the 1981 fecal specimens collected through sentinel pediatric acute gastroenteritis pathogen surveillance in Mie Prefecture, Japan (2010-2022), 236 were positive for SaV, according to PCR screening. Whole or near-whole genome sequences were determined for 158 strains by next-generation sequencing. Genotype GI.1 was the most common of the nine SaV genotypes detected, followed by GII.3 and GII.1. Phylogenetic analysis showed that SaVs of these three genotypes separated into three different clusters depending on the year of detection, suggesting continuous genetic changes in the same genotype. Coinfections involving different SaV genotypes, as well as reinfections with SaV in the same individual, were observed in this study. The main clinical manifestations were diarrhea (68.4%) and vomiting (61.6%), with an increased rate of emesis, particularly in patients over 3 years of age. In addition, 18.1% of the children had fever. This study clarified the prevalence of viral genotypes as well as clinical findings of SaV-positive gastroenteritis in children, and revealed trends by age.

Humans↗

CRISPRessoSea: streamlined analysis and comparison of pooled amplicon CRISPR screens.

BACKGROUND: CRISPR genome editing enables precise modification of genomic targets but may also induce unintended edits at off-target sites with similar sequences. Pooled amplicon sequencing can assess on- and off-target editing across many samples, yet analyzing, aggregating, and visualizing results from multiple pooled experiments remains challenging. Tools to simplify and standardize these analyses are needed to provide reproducible and comparable interpretation of editing data. RESULTS: We developed CRISPRessoSea, a software package that processes, compares, and visualizes genome editing rates from pooled amplicon sequencing experiments. The tool provides standardized workflows for analyzing editing across multiple targets and samples, supports both nuclease- and base-editing modalities, and generates clear, data-rich summaries suitable for downstream interpretation. CONCLUSIONS: CRISPRessoSea facilitates reproducible, scalable analysis of CRISPR editing outcomes across diverse experimental designs, enabling more efficient and transparent assessment of genome editing specificity. The software is freely available at https://github.com/clementlab/CRISPRessoSea .

Software↗

Systematic Dissection of Key Driver Perturbation Signatures in Single Cells via ECCITE-seq.

CRISPR screens, such as expanded CRISPR-compatible cellular indexing of transcriptomes and epitopes by sequencing (ECCITE-seq), enable the simultaneous measurement of transcriptomes, gRNA identity, and cell-surface protein expression at single-cell resolution to systematically interrogate gene function. This platform provides a powerful and scalable experimental approach for validating disease-associated regulators identified by large-scale association studies and other computational methods, including network-based analyses of multi-omics data. Here, as an example application, we describe an ECCITE-seq framework to characterize the transcriptomic consequences of perturbing multiple neuronal key driver genes associated with Alzheimer's disease (AD) in human-induced pluripotent stem cell (hiPSC)-derived neurons. More broadly, by integrating customized pooled gRNA libraries with different CRISPR effectors across multiple cell types, this approach allows for the assessment of the regulatory impact of candidate genes implicated in development and disease processes.

Humans↗

Clinical and genomic characterization of Influenza A co-infection with SARS-CoV-2 and Influenza B: a respiratory surveillance study in Assam, India.

Influenza and SARS-CoV-2 are the primary contributors to seasonal respiratory infections and frequently co-circulate, creating significant health challenges. The present respiratory surveillance study was conducted in Dibrugarh, Assam, India from January 2025 to August 2025 to investigate the genomic characteristics of circulating viruses and identify potential co-infections. Overall, 4,948 respiratory samples were screened using multiplex real-time PCR, followed by subtyping of Influenza A and Influenza B. Next-generation sequencing (NGS) was performed in selected positives of SARS-CoV-2 and Influenza A. Genomic analysis included mutational profiling, phylogenetic analysis and N-glycosylation site prediction using bioinformatics tools. Two co-infection cases were detected: one involving Influenza A (H3N2) with SARS-CoV-2 (Omicron XFG lineage) and another involving Influenza A (H3N2) with Influenza B (Victoria lineage). Both patients experienced mild illness without hospitalisation. NGS revealed that the Influenza A (H3N2) viruses belonged to clade 3C.2a1b.2a.2a.3a.1 while SARS-CoV-2 sequence was classified under the Omicron XFG lineage. Mutational analysis of the HA gene showed several amino acid differences compared to the reference vaccine strain A/Darwin/6/2021. N-glycosylation analysis predicted conserved sites at positions 79, 181, 262, and 301 in all strains along with an additional predicted site at position 110 in both co-infection cases. Although the co-infection cases presented with mild clinical manifestations, the observed genomic variations indicate a potential role of co-infecting viruses in shaping viral evolution. Given the limited genomic data available from Northeast India, the study underscores the need for sustained large scale follow up and genomic surveillance to monitor emerging mutations and target future vaccine strategies.

Humans↗

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5 kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5 kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics↗

Identification of intragenic variants in pediatric patients with intellectual disability in Peru.

BACKGROUND: Intellectual disability in Latin America can reach a frequency of 12% of the population, these may include nutritional deficiencies, exposure to toxic or infectious agents, and the lack of universal neonatal screening programs. In 90% of patients with intellectual disability, the etiology can be attributed to variants in the genome. OBJECTIVE: to determine intragenic variants in patients with intellectual disability between 5 and 18 years old at Instituto Nacional de Salud del Niño. METHODS: It is a descriptive cross-sectional study with convenience sampling. A total of 124 children diagnosed with intellectual disability were selected based on psychological test results and availability for whole exome sequencing. In addition, a chromosomal analysis of 6.55 M was performed on ten patients with a negative result in sequencing. Relative and absolute frequencies and measures of central tendency and dispersion were determined according to their nature. In addition, multiple linear regression and Poisson regression were used to determine the association between some clinical characteristics and the probability of occurrence in patients with positive results. RESULTS: The median age of the patients was 6.3 (IQR = 5.95), males accounted for 57.3%, and 91.9% of the cases had mild intellectual disability. Exome sequencing determined the etiology in 30.6% of patients with intellectual disability, of which 52.6% were autosomal dominant inheritance. The most frequent genes found were MECP2, STXBP1 and LAMA2. A broad genotype-phenotype correlation was identified, highlighting the genetic heterogeneity of intellectual disability in this population. The presence of dermatologic lesions, dystonia, peripheral neurological disorders, and fourth finger flexion limitation were observed more frequently in patients with intellectual disability with "positive results". CONCLUSIONS: This study shows that one-third of patients with intellectual disability exhibit intragenic variants, highlighting the importance of genetic analysis for accurate diagnosis. The identification of genes such as MECP2, STXBP1, and LAMA2 underscores the genetic heterogeneity of intellectual disability in the studied population. These findings emphasize the need for genetic testing in clinical management and the implementation of early detection programs in Peru.

Humans↗

Novel and High-Throughput Method of Isolating Single Fetal Cells Using FACS for NIPT.

OBJECTIVE: To evaluate fluorescence activated cell sorting (FACS) as a method of single-cell isolation of rare circulating fetal cells from maternal blood for use in cell-based non-invasive prenatal testing (cbNIPT). METHOD: Blood samples (30 mL) were collected from 75 'low-risk' pregnant women (gestational age 10-15 weeks). Fetal cells were enriched and stained using magnetic activated cell sorting. Following enrichment, single fetal cells were sorted in individual PCR tubes by FACS. After cell lysis, verification of fetal cell origin was performed using short tandem repeat (STR) analysis with the GlobalFiler PCR Amplification kit. RESULTS: An average of 13.7 cells were sorted using FACS. STR analysis identified 8.2 fetal cells on average, representing 60.2% of the sorted cells. The four-step single-cell isolation procedure facilitated an overall enrichment of approximately 16-million-fold. One sample did not render any fetal cell, corresponding to 1.3% of the samples. CONCLUSION: FACS, which is typically used for segregation of large populations of cells, can be used for single-cell isolation of rare fetal cells in an automated setup. This not only helps in making cell isolation faster and high throughput but also provides fetal cells for a more comprehensive genetic analysis of the fetus.

Humans↗

Genomic mismatch scanning identifies human genomic DNA shared identical by descent.

Genomic mismatch scanning (GMS) is a high-throughput, high-resolution identity by descent mapping technique that enriches for genomic DNA fragments that are shared between related individuals. In GMS, DNA heteroduplexes are formed from restriction-digested genomic DNA fragments from two relatives. Mismatch-free DNA heteroduplexes, likely representing DNA shared identical by descent between the two individuals, are relatively purified by depleting the mismatch-containing heteroduplexes using the Escherichia coli mismatch repair proteins and exonuclease. Here, we demonstrate using quantitative microsatellite genotyping that, despite the complexity of the human genome, GMS can enrich the majority of restriction fragments that are identical by descent between two related humans. As the entire genome is selected in GMS, an extraordinarily dense set of markers (up to 200,000 markers) may be screened in parallel. The demonstration of the molecular enrichment of identical DNA fragments in the context of the whole human genome establishes conditions for the application of GMS to human genetics. This forms a frame-work for the further development of GMS as a hybridization-based mapping technique that utilizes DNA microarray technology to map the selected identical by descent DNA fragments.

DNA↗

Fluorescence-based viability assay for studies of reactive drug intermediates.

Studies of drug toxicity, toxicologic structure-function relationships, screening of idiosyncratic drug reactions, and a variety of cytotoxic events and cellular functions in immunology and cell biology require the sensitive and rapid processing of often large numbers of cell samples. This report describes the development of a high-sensitivity, high-throughput viability assay based on (a) the carboxyfluorescein derivative 2'-7'-biscarboxyethyl-5(6)-carboxyfluorescein (BCECF) as a vital dye, (b) instrumentation capable of processing multiple small (less than 100 cells) samples, and (c) a 96-well unidirectional vacuum filtration plate. Double staining of cultured peripheral blood mononuclear cells with BCECF and propidium iodide (PI) showed no overlap between PI+ (nonviable) and BCECF+ (viable) cells by flow cytometric analysis. Optimal conditions were developed for dye loading and minimizing physical cell damage and fluorescence quench during the assay procedure. The ratio of BCECF fluorescence to internal standard fluorescent particles was linear from 40 to greater than 20,000 cells with a signal:noise ratio of approximately 3 at 40 cells/well. Sulfamethoxazole hydroxylamine (SMX-HA) was used as a model toxic drug metabolite to explore the validity of the BCECF procedure. SMX-HA, but not its parent compound sulfamethoxazole, resulted in a dose dependent loss of cellular fluorescence and the parallel accumulation of PI+ nonviable cells. When compared to the currently used tetrazolium dye reduction viability assay, the BCECF method was 3-fold more sensitive, greater than 10-fold faster, and required 1/10-1/100 the cell numbers.

Cell Survival↗

Functional screening and single-cell cultivation of marine CO2-fixing bacteria via flow-mode Raman-activated cell sorting.

Most marine CO2-fixing microorganisms remain uncultivated due to strong culture bias and low throughput of conventional approaches, which fail to link in situ function with isolated strains and render slow-growing or low-abundance taxa virtually inaccessible. This study presents an integrated single-cell workflow that incorporates 13C-NaHCO3 labeling, high-throughput flow-mode Raman-activated cell sorting (RACS) and microwell cultivation for the isolation of active CO2-fixing bacteria from the Yellow Sea. Function-guided sorting was achieved by monitoring the 13C-induced Raman shifts of carotenoids (ν1 band: ∼1507 to ∼ 1503.78 cm-1 at 24 h). Genomic and physiological analyses identified Paraburkholderia aromaticivorans FR-4 as a novel facultative chemoautotrophic nitrite-oxidizing bacterium (NOB). Its genome encodes complete nitrite oxidation and Calvin cycle pathways, together with key carbon acquisition genes (carbonic anhydrase, bicarbonate transporter). FR-4 grows autotrophically using NO2- as the electron donor and CO2/HCO3- as the carbon source, confirming its ability to couple nitrite oxidation with carbon fixation, while retaining metabolic flexibility for heterotrophic growth. By directly linking in situ carbon-fixing activity, genotype, and phenotype, this workflow provides a targeted strategy for exploring elusive marine CO2-fixing bacteria and overcomes critical limitations of conventional cultivation.

Carbon-fixing↗

A simple and sensitive high-throughput assay for steroid agonists and antagonists.

We have developed a simple and highly sensitive tissue culture-based assay for the biological activity of steroids and synthetic steroidal compounds. A DNA cassette, containing a synthetic steroid-inducible promoter controlling the expression of a bacterial chloramphenicol acetyltransferase gene (GRE5-CAT), was inserted into an Epstein-Barr virus (EBV) episomal vector which replicates autonomously in primate and human cells. We then used this promoter/reporter system to generate two stably transfected human cell lines. In the cervical carcinoma cell line HeLa, which expresses high levels of glucocorticoid receptor, the GRE5 promoter is inducible over 100-fold by the synthetic glucocorticoid dexamethasone. In the breast carcinoma cell line T47D, which expresses progesterone and androgen receptors, the GRE5 promoter is inducible over 100-fold by either progesterone or dihydrotestosterone. In both cell lines basal expression of CAT activity is strictly dependent on the presence of steroid, so that very low levels of induction can be detected. Thus, the cell lines can be used to test for low levels of agonist activity in steroid antagonists. These cell lines can be used to screen compounds for steroid agonist or antagonist activity by testing extracts of cells grown in microtiter wells directly using a colorimetric CAT assay. This system should provide a sensitive and efficient method for screening and analysis of the activity of large numbers of natural or synthetic steroid agonists or antagonists.

Breast Neoplasms↗

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans↗

A pluripotent stem cell atlas of multilineage differentiation.

Human pluripotent stem cells offer a scalable platform to study genetic and signalling mechanisms governing cell lineage decisions during differentiation. Genome-wide and single-cell transcriptomics technologies likewise offer high-throughput analysis of heterogeneous cell differentiation states. While in vivo development has been extensively characterised using these technologies, there remains a need for comprehensive single-cell transcriptomic profiling of stem cell differentiation from pluripotency. Understanding gene expression changes governing differentiation in vitro is key to developing high fidelity differentiation protocols and understanding fundamental mechanisms of development. We generated a single-cell RNA sequencing time course to study the role of developmental signalling pathways on multilineage diversification from pluripotency in vitro. The combined dataset of over 60,000 cells spans cell types from a time course of differentiation across all germ layers, ranging from gastrulation cell states to progenitor and committed cell types. These data provide a diverse benchmarking reference point to compare against in vivo development and advance understanding of signalling regulation of differentiation, providing insights into protocol development, drug screening, and regenerative medicine applications.

Pluripotent Stem Cells↗

A homogeneous immunoassay based on AlphaLICA technology for detecting florfenicol residues in animal-derived foods.

Florfenicol (FF), a broad-spectrum amide antibiotic widely used in livestock, poultry, and aquaculture, poses potential threats to food safety and public health due to its residual accumulation. In this study, a novel homogeneous immunoassay based on Amplified Luminescent Proximity Homogeneous Assay (AlphaLICA) technology was developed for the first time for rapid screening of FF residues in milk and egg matrices. By covalently immobilizing the FF-BSA conjugate and goat anti-mouse IgG onto luminescent and photosensitive microspheres, respectively, the method achieved wash-free, homogeneous quantitative detection through a competitive immunoreaction. Under optimized conditions, the assay exhibited a linear range of 0.2-16.2 ng mL-1, with a limit of detection of 9.7 pg mL-1 and a limit of quantification of 183 pg mL-1. The intra- and inter-batch coefficients of variation ranged from 3.08% to 5.70% and 2.44% to 7.09%, respectively. Spike recovery rates in milk and egg matrices ranged from 93.18% to 107.17% (RSD &#x2264; 5.57%). Cross-reactivity with 11 other common antibiotics, including chloramphenicol and thiamphenicol, was below 0.1%, demonstrating excellent specificity. Comparative analysis with a commercial ELISA kit showed high consistency (r2 = 0.9332, p < 0.001). With high sensitivity, strong specificity, simple operation, and a detection time of only 10 min, this method provides a reliable technical platform for high-throughput, rapid monitoring of FF residues in milk and egg matrices.

Journal Article↗

Methylation-based droplet digital polymerase chain reaction shows high concordance with chronic lymphocytic leukemia IGHV somatic mutation status.

OBJECTIVE: Somatic hypermutation at immunoglobulin heavy chain variable (IGHV) genes, an established prognostic and predictive biomarker for chronic lymphocytic leukemia (CLL), is assessed by gene sequencing. We developed a single methylation-specific droplet digital polymerase chain reaction (methyl-ddPCR) to predict IGHV status in patients with CLL. METHODS: The CLL methylation array and IGHV data from the International Cancer Genome Consortium (ICGC) were used for biomarker discovery. Top-ranked candidate regions were manually screened for PCR primer and probe binding sites. A single methyl-ddPCR was evaluated on an internal cohort of CLLs with mutated (M), unmutated (U), and inconclusive IGHV results originally determined by next-generation sequencing (NGS). RESULTS: Analysis of ICGC data identified array probe cg23844018 as a candidate for the PCR. The corresponding CpG site showed high methylation levels in U-CLL and lower levels in M-CLL. On the internal cohort, a single optimal cutoff correctly classified 104 of 115 U- and M-CLLs (90.4%; area under the curve&#x2005;=&#x2005;0.96). The PCR data correlated with some prognostic fluorescence in situ hybridization and CLL subset groupings. Limited analysis suggests that the PCR may be able to stratify some patients with CLL who have inconclusive results on IGHV NGS testing. CONCLUSIONS: The methyl-ddPCR showed high concordance with CLL IGHV status in an internal cohort.

Humans↗

Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.

A rigorous analysis of the Merck-sponsored EST data with respect to known gene sequences increases the utility of the data set and helps refine methods for building a gene index. A highly curated human transcript data base was used as a reference data set of known genes. A detailed analysis of EST sequences derived from known genes was performed to assess the accuracy of EST sequence annotation. The EST data was screened to remove low-quality and low-complexity sequences. A set of high-quality ESTs similar to the transcript data base was identified using BLAST; this subset of ESTs was compared with the set of known genes using the Smith-Waterman algorithm. Error rates of several types were assessed based on a flexible match criterion defining sequence identity. The rate of lane-tracking errors is very low, approximately 0.5%. Insert size data is accurate within approximately 20%. Reversed clone and internal priming error rates are approximately 5% and 2.5%, respectively, contributing to the incorrect identification of reads as 3' ends of genes. Follow-up investigation reveals that a significant number of clones, miscategorized as reversed, represent overlapping genes on the opposite strand of entries in the transcript data base. Relevance of these results to the creation of a high-quality index to the human genome capable of supporting diverse genomic investigations is discussed.

Algorithms↗

Minisequencing: a specific tool for DNA analysis and diagnostics on oligonucleotide arrays.

We describe a method for multiplex detection of mutations in which the solid-phase minisequencing principle is applied to an oligonucleotide array format. The mutations are detected by extending immobilized primers that anneal to their template sequences immediately adjacent to the mutant nucleotide positions with single labeled dideoxynucleoside triphosphates using a DNA polymerase. The arrays were prepared by coupling one primer per mutation to be detected on a small glass area. Genomic fragments spanning nine disease mutations, which were selected as targets for the assay, were amplified in multiplex PCR reactions and used as templates for the minisequencing reactions on the primer array. The genotypes of homozygous and heterozygous genomic DNA samples were unequivocally defined at each analyzed nucleotide position by the highly specific primer extension reaction. In a comparison to hybridization with immobilized allele-specific probes in the same assay format, the power of discrimination between homozygous and heterozygous genotypes was one order of magnitude higher using the minisequencing method. Therefore, single-nucleotide primer extension is a promising principle for future high-throughput mutation detection and genotyping using high density DNA-chip technology.

DNA-Directed DNA Polymerase↗