Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Identification of candidate coding region single nucleotide polymorphisms in 165 human genes using assembled expressed sequence tags.

Using assembled expressed sequence tags (ESTs) from 50 different cDNA libraries, we have identified contigs that represent the complete coding sequences of 850 known human genes, and have scanned these for high quality sequence substitutions. We report the identification and characteristics of 201 candidate single nucleotide polymorphisms found in the coding sequences (cSNPs) of 165 of these genes. Using a conservative calculation, coding region nucleotide diversity (the average number of differences between any pair of chromosomes) was found to be 3 per 10,000 bp based on this data. This analysis reveals that assembled ESTs from multiple libraries may provide a rich source of comparative sequences to search for cSNPs in the human genome.

Amino Acid Substitution↗

The sequence of the human genome.

A 2.91-billion base pair (bp) consensus sequence of the euchromatic portion of the human genome was generated by the whole-genome shotgun sequencing method. The 14.8-billion bp DNA sequence was generated over 9 months from 27,271,853 high-quality sequence reads (5.11-fold coverage of the genome) from both ends of plasmid clones made from the DNA of five individuals. Two assembly strategies-a whole-genome assembly and a regional chromosome assembly-were used, each combining sequence data from Celera and the publicly funded genome effort. The public data were shredded into 550-bp segments to create a 2.9-fold coverage of those genome regions that had been sequenced, without including biases inherent in the cloning and assembly procedure used by the publicly funded group. This brought the effective coverage in the assemblies to eightfold, reducing the number and size of gaps in the final assembly over what would be obtained with 5.11-fold coverage. The two assembly strategies yielded very similar results that largely agree with independent mapping data. The assemblies effectively cover the euchromatic regions of the human chromosomes. More than 90% of the genome is in scaffold assemblies of 100,000 bp or more, and 25% of the genome is in scaffolds of 10 million bp or larger. Analysis of the genome sequence revealed 26,588 protein-encoding transcripts for which there was strong corroborating evidence and an additional approximately 12,000 computationally derived genes with mouse matches or other weak supporting evidence. Although gene-dense clusters are obvious, almost half the genes are dispersed in low G+C sequence separated by large tracts of apparently noncoding sequence. Only 1.1% of the genome is spanned by exons, whereas 24% is in introns, with 75% of the genome being intergenic DNA. Duplications of segmental blocks, ranging in size up to chromosomal lengths, are abundant throughout the genome and reveal a complex evolutionary history. Comparative genomic analysis indicates vertebrate expansions of genes associated with neuronal function, with tissue-specific developmental regulation, and with the hemostasis and immune systems. DNA sequence comparisons between the consensus sequence and publicly funded genome data provided locations of 2.1 million single-nucleotide polymorphisms (SNPs). A random pair of human haploid genomes differed at a rate of 1 bp per 1250 on average, but there was marked heterogeneity in the level of polymorphism across the genome. Less than 1% of all SNPs resulted in variation in proteins, but the task of determining which SNPs have functional consequences remains an open challenge.

Algorithms↗

"Beijing Region" (3pter-D3S3397) of the human genome: complete sequence and analysis.

The goal of the Human Genome Project (HGP) is to determine a complete and high-quality sequence of the human genome. China, as one of the six member states, takes a region between 3pter and D3S3397 of the human chromosome 3 as its share of this historic project, referred as "Beijing Region". The complete sequence of this region comprises of 17.4 megabasepairs (Mb) with an average GC content of 42% and an average recombination rate of 2.14 cM/Mb. Within Beijing Region, 122 known and 20 novel genes are identified, as well as 42607 single nucleotide polymorphisms (SNPs). Comprehensive analyses also reveal: (i) gene density and GC-content of Beijing Region are in agreement with human cytogenetic maps, i.e. G-minus bands are GC-rich and of a high gene density, whereas G-plus bands are GC-poor and of a relatively low gene density; (ii) the average recombination rate within Beijing Region is relatively high compared with other regions of chromosome 3, with the highest recombination rate of 6.06 cM/Mb in the subtelomeric area; (iii) it is most likely that a large gene, associated with the mammary gland, may reside in the 1.1 Mb gene-poor area near the telomere; (iv) many disease-related genes are genetically mapped to Beijing Region, including those associated with cancers and metabolic syndromes. All make Beijing Region an important target for in-depth molecular investigations with a purpose of medical applications.

Animals↗

Early developing embryos affect the gene expression patterns in the mouse oviduct.

Fertilization and development of mouse embryos occur in the ampullae of oviduct. We hypothesize that fetal-maternal communication exists in the preimplantation period, allowing optimal development of embryos. It is known that embryotrophic factors from oviduct affect the development of embryos. Although embryos affect their own transport in the oviduct, the mechanism of action is unknown. As a step toward understanding the action of embryos on oviductal physiology, we adopted suppression subtractive hybridization (SSH) to compare the gene expression in the mouse oviduct containing early embryos with that of oviduct containing oocytes. Ten to twelve 1-cell mouse embryos were transferred to one oviduct of a foster mother and similar number of oocytes were transferred to the contralateral oviduct. The animals were sacrificed after 48 h and their oviducts were excised for mRNA study. Using SSH, we screened out 250 putative positive clones from the subtracted embryo-containing oviduct library and 97 of them were screened positive by reverse dot-blot analysis. DNA sequence analysis identified genes that shared high homology with sequences in GenBank/EMBL database with unknown functions. Overall, 13 of the 90 high-quality sequences (14%) were homologous to 6 different genes previously described. Reverse Northern analysis confirmed that the expression of these genes were higher in the embryo-containing oviduct than in the oocyte-containing oviduct. About 12% of these clones (11/90) were novel. This article is the first to report identification of genes in the oviduct that are upregulated in the presence of embryos during the preimplantation period.

Animals↗

A sequence-ready BAC clone contig of human chromosome 10p15 spanning the loss of heterozygosity region in glioma.

Deletion of chromosome 10 is one of the most common chromosomal alterations in glioma. At 10p15, the telomeric region of the short arm of chromosome 10, loss of heterozygosity (LOH) has been frequently observed by microsatellite analysis, suggesting the presence of a tumor suppressor gene. We examined LOH in 34 gliomas on chromosome 10, and frequent LOH on 10p was detected on 10p15, in agreement with deletion mapping studies on chromosome 10. We then constructed a bacterial artificial chromosome (BAC) clone contig covering the critical region, which spanned the interval between D10S249 and D10S533 on 10p15. The map contained 68 BAC clones connected by 74 sequenced tag sites (STSs) and covered approximately 2.7 Mb, with one gap. A total of 74 STSs, including 6 microsatellite markers, 29 expressed sequenced tags (ESTs), and 39 BAC end STSs, were physically arranged. Twenty-eight ESTs were mapped in the interval between D10S249 and D10S559 (approximately 1200 kb), and another EST was mapped in the interval between D10S559 and D10S533 (approximately 1300 kb). This sequence-ready BAC clone contig map will be a basic resource for high-quality sequencing and positional cloning of the putative tumor suppressor gene at 10p15 in glioma.

Base Sequence↗

ForestTreeDB: a database dedicated to the mining of tree transcriptomes.

ForestTreeDB is intended as a resource that centralizes large-scale expressed sequence tag (EST) sequencing results from several tree species (http://foresttree.org/ftdb). It currently encompasses 344,878 quality sequences from 68 libraries, from diverse organs of conifer and hybrid poplar trees. It utilizes the Nimbus data model to provide a hosting system for multiple projects, and uses object-relational mapping APIs in Java and Perl for data accesses within an Oracle database designed to be scalable, maintainable and extendable. Transcriptome builds or unigene sets occupy the focal point of the system. Several of the five current species-specific unigenes were used to design microarrays and SNP resources. The ForestTreeDB web application provides the means for multiple combination database queries. It presents the user with a list of discrete queries to retrieve and download large EST datasets or sequences from precompiled unigene assemblies. Functional annotation assignment is not trivial in conifers which are distantly related to angiosperm model plants. Optimal annotations are achieved through database queries that integrate results from several procedures based open-source tools. ForestTreeDB aims to facilitate sequence mining of coherent annotations in multiple species to support comparative genomic approaches. We plan to continuously enrich ForestTreeDB with other resources through collaborations with other genomic projects.

Databases, Nucleic Acid↗

Sequence analysis of the 2nd intron revealed common sequence motifs providing the means for a unique sequencing based typing protocol of the HLA-A locus.

We here present a sequencing strategy for the HLA-A locus which is generally applicable for all HLA class I genes. The typing strategy is based on a group-specific amplification according to the serologically defined antigens. The PCR products carry the typing-relevant polymorphic regions of the 2nd and 3rd exon including the 2nd intron. The sequencing primers were designed to match conserved sequence motifs in the 2nd intron allowing a nested sequencing approach in 3' and 5' direction. These conserved regions were identified after sequence compilation of the 2nd intron of 143 clinical samples and 48 cell lines mostly from the 9th and 10th IHWC representing all serologically defined groups of alleles. This strategy allowed the use of only one 5' and one 3' sequencing primer regardless of the amplified allele. Therefore, it was possible to use dye terminator as well as dye primer sequencing chemistry. The amplification strategy allowed the separation of the haplotypes in almost all cases. Thus, an assignment of heterozygous positions requiring high sequencing quality was not necessary, allowing the application of Sequenase as well as TaqPolymerase as sequencing enzyme. Concerning the resolution of heterozygosity it is obvious that this approach is superior to a typing system using a single pair of generic primers followed by direct sequencing, since the latter technique is not capable of defining the cis/trans linkage of polymorphic sequences and, hence, cannot exclude the presence of unknown alleles.

Alleles↗

Optimization of coupled PCR amplification and cycle sequencing of cloned and genomic DNA.

We describe optimization of a coupled amplification and cycle sequencing (CAS) method for rapid characterization of cloned or genomic DNA. Our modification of this method, termed coupled PCR amplification and cycle sequencing (CPACS), utilizes commercially available reagents, does not require template purification and produces high-quality sequence ladders from nanogram quantities of complex genomic DNA. The reactions have been streamlined to permit automation. Finally, we show that the technique can be applied more efficiently in conjunction with the AutoTrans 350 Direct Transfer Electrophoresis System and 33P-labeled sequencing primers.

Autoanalysis↗

Assessing Hardy-Weinberg equilibrium in T2T-aligned 1000 genomes project.

Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. These motivate a re-examination of HWE in sex-aware analyses. Using the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequencing data from 2,490 individuals in the 1000 Genomes Project, we examined genome-wide sex-specific deviations from HWE across five super-populations. Our analyses were restricted to bi-allelic SNPs with non-missing genotypes and minor allele frequency (MAF) &#x2265;5% in both sexes of the five super-populations. We applied an allele-based framework to quantify both the magnitude and direction of Hardy-Weinberg disequilibrium (HWD), followed by a second-order omnibus meta-analysis that combined HWD results across populations and sexes. At a genome-wide significance threshold of p&#x2009;<&#x2009;5e-8, 0.9% of autosomal SNPs exhibited significant deviations from HWE. The majority of these deviations were associated with genomic features indicative of poor sequence quality. Restricting the analysis to reliable genomic regions substantially reduced the number of signals, yielding 255 autosomal SNPs and one non-pseudoautosomal chromosome X SNP. Among these, 140 autosomal SNPs displayed significant heterogeneity across populations but not across sexes. Notably, eight SNPs within a 15-bp region on chromosome 14q31.3 showed excess heterozygosity in both sexes of the African super-population (AFR). Finally, we developed a multivariate predictor of HWD based on sequence features, providing a practical tool that can be integrated into existing quality control pipelines for whole genome sequencing studies.

Journal Article↗

ESTAnnotator: A tool for high throughput EST annotation.

In high throughput sequence analysis, it is often necessary to combine the results of contemporary bioinformatics tools, because no individual tool alone computes all the requested information. ESTAnnotator is a tool for the high throughput annotation of expressed sequence tags (ESTs) by automatically running a collection of bioinformatics applications. In the first step, a quality check is performed and repeats, vector parts and low quality sequences are masked. Then successive steps of database searching and EST clustering are performed. Already known transcripts present within mRNA and genomic DNA reference databases are identified. Subsequently, tools for the clustering of anonymous ESTs, and for further database searches at the protein level, are applied. Finally, the outputs of each individual tool are gathered and the relevant results presented in a descriptive summary. ESTAnnotator was already successfully applied for the systematic identification and characterisation of novel human genes involved in cartilage/bone formation, growth, differentiation and homeostasis. ESTAnnotator is available at http://genome.dkfz-heidelberg.de, contact: genome@dkfz.de.

Cartilage↗

Gene expression and specificity in the mature zone of the lobster olfactory organ.

The lobster olfactory organ is an important model for investigating many aspects of the olfactory system. To facilitate study of the molecular basis of olfaction in lobsters, we made a subtracted cDNA library from the mature zone of the olfactory organ of Homarus americanus, the American lobster. Sequencing of the 5'-end of 5,184 cDNA clones produced 2,389 distinct high-quality sequences consisting of 1,944 singlets and 445 contigs. Matches to known sequences corresponded with the types of cells present in the olfactory organ, including specific markers of olfactory sensory neurons, auxiliary cells, secretory cells of the aesthetasc tegumental gland, and epithelial cells. The wealth of neuronal mRNAs represented among the sequences reflected the preponderance of neurons in the tissue. The sequences identified candidate genes responsible for known functions and suggested new functions not previously recognized in the olfactory organ. A cDNA microarray was designed and tested by assessing mRNA abundance differences between two of the lobster's major chemosensory structures: the mature zone of the olfactory organ and the dactyl of the walking legs, a taste organ. The 115 differences detected again emphasized the abundance of neurons in the olfactory organ, especially a cluster of mRNAs encoding cytoskeletal-associated proteins and cell adhesion molecules such as 14-3-3zeta, actins, tubulins, trophinin, Fax, Yel077cp, suppressor of profilin 2, and gelsolin.

Animals↗

cDNA2Genome: a tool for mapping and annotating cDNAs.

BACKGROUND: In the last years several high-throughput cDNA sequencing projects have been funded worldwide with the aim of identifying and characterizing the structure of complete novel human transcripts. However some of these cDNAs are error prone due to frameshifts and stop codon errors caused by low sequence quality, or to cloning of truncated inserts, among other reasons. Therefore, accurate CDS prediction from these sequences first require the identification of potentially problematic cDNAs in order to speed up the posterior annotation process. RESULTS: cDNA2Genome is an application for the automatic high-throughput mapping and characterization of cDNAs. It utilizes current annotation data and the most up to date databases, especially in the case of ESTs and mRNAs in conjunction with a vast number of approaches to gene prediction in order to perform a comprehensive assessment of the cDNA exon-intron structure. The final result of cDNA2Genome is an XML file containing all relevant information obtained in the process. This XML output can easily be used for further analysis such us program pipelines, or the integration of results into databases. The web interface to cDNA2Genome also presents this data in HTML, where the annotation is additionally shown in a graphical form. cDNA2Genome has been implemented under the W3H task framework which allows the combination of bioinformatics tools in tailor-made analysis task flows as well as the sequential or parallel computation of many sequences for large-scale analysis. CONCLUSIONS: cDNA2Genome represents a new versatile and easily extensible approach to the automated mapping and annotation of human cDNAs. The underlying approach allows sequential or parallel computation of sequences for high-throughput analysis of cDNAs.

Chromosome Mapping↗

Genomic exploration of the hemiascomycetous yeasts: 2. Data generation and processing.

The generation of sequencing data for the hemiascomycetous yeast random sequence tag project was performed using the procedures established at GENOSCOPE. These procedures include a series of protocols for the sequencing reactions, using infra-red labelled primers, performed on both ends of the plasmid inserts in the same reaction tube, and their analysis on automated DNA sequencers. They also include a package of computer programs aimed at detecting potential assignation errors, selecting good quality sequences and estimating their useful length.

Ascomycota↗

Organization of the Bacillus subtilis 168 chromosome between kdg and the attachment site of the SP beta prophage: use of Long Accurate PCR and yeast artificial chromosomes for sequencing.

Within the Bacillus subtilis genome sequencing project, the region between lysA and ilvA was assigned to our laboratory. In this report we present the sequence of the last 36 kb of this region, between the kdg operon and the attachment site of the SP beta prophage. A two-step strategy was used for the sequencing. In the first step, total chromosomal DNA was cloned in phage M13-based vectors and the clones carrying inserts from the target region were identified by hybridization with a cognate yeast artificial chromosome (YAC) from our collection. Sequencing of the clones allowed us to establish a number of contigs. In the second step the contigs were mapped by Long Accurate (LA) PCR and the remaining gaps closed by sequencing of the PCR products. The level of sequence inaccuracy due to LA PCR errors appeared to be about 1 in 10,000, which does not affect significantly the final sequence quality. This two-step strategy is efficient and we suggest that it can be applied to sequencing of longer chromosomal regions. The 36 kb sequence contains 38 coding sequences (CDSs), 19 of which encode unknown proteins. Seven genetic loci already mapped in this region, xpt, metB, ilvA, ilvD, thyB, dfrA and degR were identified. Eleven CDSs were found to display significant similarities to known proteins from the data banks, suggesting possible functions for some of the novel genes: cspD may encode a cold shock protein; bcsA, the first bacterial homologue of chalcone synthase; exol, a 5' to 3' exonuclease, similar to that of DNA polymerase I of Escherichia coli; and bsaA, a stress-response-associated protein. The protein encoded by yplP has homology with the transcriptional NifA-like regulators. The arrangement of the genes relative to possible promoters and terminators suggests 19 potential transcription units.

Amino Acid Sequence↗

LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA.

To compare entire genomes from different species, biologists increasingly need alignment methods that are efficient enough to handle long sequences, and accurate enough to correctly align the conserved biological features between distant species. We present LAGAN, a system for rapid global alignment of two homologous genomic sequences, and Multi-LAGAN, a system for multiple global alignment of genomic sequences. We tested our systems on a data set consisting of greater than 12 Mb of high-quality sequence from 12 vertebrate species. All the sequence was derived from the genomic region orthologous to an approximately 1.5-Mb region on human chromosome 7q31.3. We found that both LAGAN and Multi-LAGAN compare favorably with other leading alignment methods in correctly aligning protein-coding exons, especially between distant homologs such as human and chicken, or human and fugu. Multi-LAGAN produced the most accurate alignments, while requiring just 75 minutes on a personal computer to obtain the multiple alignment of all 12 sequences. Multi-LAGAN is a practical method for generating multiple alignments of long genomic sequences at any evolutionary distance. Our systems are publicly available at http://lagan.stanford.edu.

Animals↗

Real-world emergence of nirsevimab resistance in breakthrough infections with respiratory syncytial virus-B: a multicentre observational study in France.

BACKGROUND: Respiratory syncytial virus (RSV) is a leading cause of lower respiratory tract infection in infants. Nirsevimab, a long-acting monoclonal antibody targeting a conserved epitope on the prefusion F protein (site &#x3a6;), has shown high efficacy in clinical trials and early real-world studies. Although widespread resistance has not been reported, concerns remain about the emergence of escape variants, particularly among RSV-B viruses. During the 2024-25 RSV season in France, RSV-B predominated, providing a unique opportunity to examine breakthrough infections with RSV-B and resistance at a large scale. The study aimed to characterise RSV escape from nirsevimab using genotypic and phenotypic methods. METHODS: This POLYRES-2 project was a multicentre, national, observational study conducted in hospital settings (inpatients and outpatients) across France during the 2024-25 RSV season. We included infants aged 1 year or under with a RT-PCR-confirmed RSV infection in routine care, regardless of whether they had received nirsevimab. Infants were identified through hospital virology laboratory databases. Each participating centre was requested to include a balanced number of nirsevimab-exposed and non-exposed infected infants throughout the study period. Clinical data were retrieved from electronic medical records. We compared RSV susceptibility to nirsevimab in infants who received nirsevimab with that in nirsevimab-naive infants. Respiratory samples were sequenced for full-length RSV genomes. To ensure reliability, phylogenetic and mutational analyses were restricted to high-quality sequences with greater than or equal to 90% genome coverage and complete reads across the nirsevimab-binding site. Clinical RSV isolates were tested for neutralisation by nirsevimab. We analysed F candidate substitutions using a fusion inhibition assay. The primary outcomes were presence of resistance-associated substitutions (RASs) in the RSV F protein (site &#x3a6;) and phenotypic resistance to nirsevimab. FINDINGS: Among 1023 RSV-infected infants, 858 (83&#xb7;9%) had full-length RSV genome sequences: 419 (48&#xb7;8%) from nirsevimab-treated breakthrough infections (212 [50&#xb7;6%] RSV-A, 207 [49&#xb7;4%] RSV-B) and 439 (51&#xb7;2%) from nirsevimab-naive infants (192 [43&#xb7;7%] RSV-A, 247 [56&#xb7;3%] RSV-B). RASs were identified in two of 195 RSV-A breakthrough infections (1&#xb7;0%) and in 23 of 184 RSV-B breakthrough infections (12&#xb7;5%). In RSV-A, the only RAS was F:K209E, conferring intermediate resistance. In RSV-B, resistance was more frequent and diverse than in RSV-A: 12 of 23 (52.2%) resistant viruses carried a substitution at residue 208 (F:N208D, F:N208I, F:N208K, F:N208S, or F:N208Y). Additional novel substitutions, including F:I64V/F:K65E, F:K68I, F:L204S, and F:P205S, also mediated resistance. Notably, a resistant RSV-B variant (F:N208S) was detected almost 1 year after prophylaxis. No resistant RSV was detected in nirsevimab-naive infants. INTERPRETATION: Resistance to nirsevimab in RSV-B can emerge in real-world settings, affecting around 12% of breakthrough infections and showing greater diversity than previously recognised, although the clinical impact remains constrained by available evidence. Detection of resistant variants long after prophylaxis highlights the need for extended genomic surveillance. Integration of clinical and virological data will be essential to sustain the long-term effectiveness of RSV monoclonal antibody programmes. FUNDING: This study was supported by a grant from the Agence Nationale de Recherche sur le Sida et les h&#xe9;patites virales - Maladies Infectieuses Emergentes and the French Ministry of Health and Prevention.

Humans↗

Comparative bioinformatic analysis of genes expressed in common bean (Phaseolus vulgaris L.) seedlings.

To rapidly and cost-effectively generate gene expression data, we developed an annotated unigene database of common bean (Phaseolus vulgaris L.). In this study, 3 cDNA libraries were constructed from the bean breeding line SEL1308, 1 from young leaf and 2 from seedlings inoculated or not inoculated with the fungal pathogen Colletotrichum lindemuthianum (Sacc. & Magnus) Briosi & Cavara, which causes anthracnose in common bean. To this date, 5255 single-pass sequences have been included in the database after selection based on sequence quality. These ESTs were trimmed and clustered using the computer programs Phred and CAP3 to form a unigene collection of 3126 unique sequences. Within clusters, 318 single nucleotide polymorphisms (SNPs) and 68 insertions-deletions (indels) were found, indicating the presence of paralogous gene families in our database. Each unigene sequence was analyzed for possible function using their similarity to known genes represented in the GenBank database and classified into 14 categories. Only 314 unigenes showed significant similarities to Phaseolus genomic sequences and P. vulgaris ESTs, which indicates that 90% (2818 unigenes) of our database represent newly discovered common bean genes. In addition, 12% (387 unigenes) were shown to be specific to common bean. This study represents a first step towards the discovery of novel genes in beans and a valuable source of molecular markers for expressed gene tagging and mapping.

Computational Biology↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗