Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Complementary DNA sequencing: expressed sequence tags and human genome project.

Automated partial DNA sequencing was conducted on more than 600 randomly selected human brain complementary DNA (cDNA) clones to generate expressed sequence tags (ESTs). ESTs have applications in the discovery of new human genes, mapping of the human genome, and identification of coding regions in genomic sequences. Of the sequences generated, 337 represent new genes, including 48 with significant similarity to genes from other organisms, such as a yeast RNA polymerase II subunit; Drosophila kinesin, Notch, and Enhancer of split; and a murine tyrosine kinase receptor. Forty-six ESTs were mapped to chromosomes after amplification by the polymerase chain reaction. This fast approach to cDNA characterization will facilitate the tagging of most human genes in a few years at a fraction of the cost of complete genomic sequencing, provide new genetic markers, and serve as a resource in diverse biological research fields.

Amino Acid Sequence↗

Pash: efficient genome-scale sequence anchoring by Positional Hashing.

Pash is a computer program for efficient, parallel, all-against-all comparison of very long DNA sequences. Pash implements Positional Hashing, a novel parallelizable method for sequence comparison based on k-mer representation of sequences. The Positional Hashing method breaks the comparison problem in a unique way that avoids the quadratic penalty encountered with other sensitive methods and confers inherent low-level parallelism. Furthermore, Positional Hashing allows one to readily and predictably trade between sensitivity and speed. In a simulated comparison task, anchoring computationally mutated reads onto a genome, the sensitivity of Pash was equal to or greater than that of BLAST and BLAT, with Pash outperforming these programs as the reads became shorter and less similar to the genome. Using modest computing resources, we employed Pash for two large-scale sequence comparison tasks: comparison of three mammalian genomes, and anchoring millions of chimpanzee whole-genome shotgun sequencing reads onto the human genome. The results of these comparisons by Pash agree with those computed by other methods that use more than an order of magnitude more computing resources. These results confirm the sensitivity of Positional Hashing.

Animals↗

Human BAC ends.

The Human BAC Ends database includes all non-redundant human BAC end sequences (BESs) generated by The Institute for Genomic Research (TIGR), the University of Washington (UW) and California Institute of Technology (CalTech). It incorporates the available BAC mapping data from different genome centers and the annotation results of each end sequence for the contents of repeats, ESTs and STS markers. For each BAC end the database integrates the sequence, the phred quality scores, the map and the annotation, and provides links to sites of the library information, the reports of GenBank, dbGSS and GDB, and other relevant data. The database is freely accessible via the web and supports sequence or clone searches and anonymous FTP. The relevant sites and resources are described at http://www.tigr.org/ tdb/humgen/bac_end_search/bac_end_intro.html

Chromosome Mapping↗

Creation of a minimal tiling path of genomic clones for Drosophila: provision of a common resource.

On the basis of shotgun subclone libraries used in the sequencing of the Drosophila melanogaster genome, a minimal tiling path of subclones across much of the genome was determined. About 320,000 shotgun clones for chromosomes X(12-20), 2R, 2L, 3R, and 4 were available from the Berkeley Drosophila Genome Project. The clone inserts have an average length of 3.4 kb and are amenable to standard PCR amplification. The resulting tiling path covers 86.2% of chromosome X(12-20), 86.2% of chromosomal arm 2R, 79.0% of 2L, 89.6% of 3R, and 80.5% of chromosome 4. In total, the 25,135 clones represent 76.7 Mb--equivalent to about 67% of the genome--and would be suitable for producing a microarray on a single slide.

Animals↗

Technical advances: genome-wide cDNA-AFLP analysis of the Arabidopsis transcriptome.

cDNA-AFLP, a technology historically used to identify small numbers of differentially expressed genes, was adapted as a genome-wide transcript profiling method. mRNA levels were assayed in a diverse range of tissues from Arabidopsis thaliana plants grown under a variety of environmental conditions. The resulting cDNA-AFLP fragments were sequenced. By linking cDNA-AFLP fragments to their corresponding mRNAs via these sequences, a database was generated that contained quantitative expression information for up to two-thirds of gene loci in A. thaliana, ecotype Ws. Using this resource, the expression levels of genes, including those with high nucleotide sequence similarity, could be determined in a high-throughput manner merely by comparing cDNA-AFLP profiles with the database. The lengths of cDNA-AFLP fragments inferred from their electrophoretic mobilities correlated well with actual fragment lengths determined by sequencing. In addition, the concentrations of AFLP fragments from single cDNAs were highly correlated, illustrating the validity of cDNA-AFLP as a quantitative, genome-wide, transcript profiling method. cDNA-AFLP profiles were also qualitatively consistent with mRNA profiles obtained from parallel microarray analysis, and with data from previous studies.

Arabidopsis↗

The GSA Family in 2025: A Broadened Sharing Platform for Multi-omics and Multimodal Data.

The Genome Sequence Archive family (GSA family) provides a comprehensive suite of database resources for archiving, retrieving, and sharing multi-omics data for the global academic and industrial communities. It currently comprises four distinct database members: the Genome Sequence Archive (GSA, https://ngdc.cncb.ac.cn/gsa), the Genome Sequence Archive for Human (GSA-Human, https://ngdc.cncb.ac.cn/gsa-human), the Open Archive for Miscellaneous Data (OMIX, https://ngdc.cncb.ac.cn/omix), and the Open Biomedical Imaging Archive (OBIA, https://ngdc.cncb.ac.cn/obia). Compared to its 2021 version, the GSA family has expanded significantly by introducing a new repository, the OBIA, and by comprehensively upgrading the existing databases. Notable enhancements to the existing members include broadening the range of accepted data types, strengthening quality control systems, improving the data retrieval system, and refining data-sharing management mechanisms.

Humans↗

Information for the Coordinates of Exons (ICE): a human splice sites database.

We present a comprehensive database, Information for the Coordinates of Exons (ICE), of genomic splice sites (SSs) for 10,803 human genes. ICE contains 91,846 pairs of donor acceptor sites, supported by the alignment of "full-length" human mRNAs (including transcript variants) on human genomic sequences. ICE represents the largest collection of human SSs known to date and provides a significant resource to both molecular biologists and bioinformaticians alike. A user can visualize and extract genomic sequences around SSs of the donor acceptor pairs and can also visualize the primary structure of individual genes. We list in this article the 22 most frequently found canonical and noncanonical splice sites. The top four most represented donor acceptor pairs (GT-AG, GC-AG, AT-AC, and GT-GG) accounted for 99.16% of our data set. In addition, we calculated the SS matrix models for the three most common donor acceptor pairs. The database is focused on providing SSs and surrounding sequence information, associated SS and sequence characteristics, and relation to overall transcript structure. It allows targeted search and presents evidence for the gene structure.

Computational Biology↗

Conservation of intrinsic disorder in protein domains and families: I. A database of conserved predicted disordered regions.

Many protein regions have been shown to be intrinsically disordered, lacking unique structure under physiological conditions. These intrinsically disordered regions are not only very common in proteomes, but also crucial to the function of many proteins, especially those involved in signaling, recognition, and regulation. The goal of this work was to identify the prevalence, characteristics, and functions of conserved disordered regions within protein domains and families. A database was created to store the amino acid sequences of nearly one million proteins and their domain matches from the InterPro database, a resource integrating eight different protein family and domain databases. Disorder prediction was performed on these protein sequences. Regions of sequence corresponding to domains were aligned using a multiple sequence alignment tool. From this initial information, regions of conserved predicted disorder were found within the domains. The methodology for this search consisted of finding regions of consecutive positions in the multiple sequence alignments in which a 90% or more of the sequences were predicted to be disordered. This procedure was constrained to find such regions of conserved disorder prediction that were at least 20 amino acids in length. The results of this work included 3,653 regions of conserved disorder prediction, found within 2,898 distinct InterPro entries. Most regions of conserved predicted disorder detected were short, with less than 10% of those found exceeding 30 residues in length.

Amino Acid Sequence↗

Society and the human genome. Sir Frederick Gowland Hopkins Memorial Lecture.

In June 2000, the draft sequence of the human genome was announced. It is, and will be for some years, incomplete, but the vast majority is now available. Currently about a third is finished (including two complete chromosomes); the rest has good coverage, but not long-range continuity. First-pass analysis indicates, among other things, fewer genes than expected: about 40000 now looks a likely number. This uncertainty illustrates the difficulty of interpretation: the sequence is not an end in itself, but a resource to be continually reanalysed as our biological understanding increases. That is the scientific reason for releasing it promptly, fully and freely. The social reasons for doing so are even more compelling.

Databases as Topic↗

Paper2sequences: retrieval of sequences listed in a publication.

Our web-based tool simplifies the often laborious procedure of retrieving a set of biosequences in a publication or webpage. As a front-end to the Bioperl toolkit, it accepts as an input a list of identifiers. They are specified in an ASCII table (copy-pasted from the publication's PDF or HTML page) and give rise to queries in multiple databases for the protein/nucleic acid data specified. Currently, GenBank, PIR (Protein Information Resource) and Swiss-Prot are supported. For any sequence accession code listed, the database can be specified and, if retrieval fails, automatic lookup for the same code in other databases can be requested. Sequence length information (if specified) and heuristic rules are used to drive the lookup if multiple protein coding sequences (CDS) are part of a single accession. Warnings are issued in cases of ambiguities and inconsistencies. An advanced option enables the user to format the output in whatever format they wish.

Amino Acid Sequence↗

Differential gene expression in egg cells and zygotes suggests that the transcriptome is restructed before the first zygotic division in tobacco.

We applied suppression subtractive hybridization and mirror orientation selection to compare gene expression profiles of isolated Nicotiana tabacum cv SR1 zygotes and egg cells. Our results revealed that many differentially expressed genes in zygotes were transcribed de novo after fertilization. Some of these genes are critical to zygote polarity and pattern formation during early embryogenesis. This suggests that the transcriptome is restructed in zygote and that the maternal-to-zygotic transition happens before the first zygotic division, which is much earlier in higher plants than in animals. The expressed sequence tags used in this study provide a valuable resource for future research on fertilization and early embryogenesis.

Body Patterning↗

Genome-wide comparative phylogenetic analysis of the rice and Arabidopsis Dof gene families.

BACKGROUND: Dof proteins are a family of plant-specific transcription factors that contain a particular class of zinc-finger DNA-binding domain. Members of this family have been found to play diverse roles in gene regulation of processes restricted to the plants. The completed genome sequences of rice and Arabidopsis constitute a valuable resource for comparative genomic analyses, since they are representatives of the two major evolutionary lineages within the angiosperms. In this framework, the identification of phylogenetic relationships among Dof proteins in these species is a fundamental step to unravel functionality of new and yet uncharacterised genes belonging to this group. RESULTS: We identified 30 different Dof genes in the rice Oryza sativa genome and performed a phylogenetic analysis of a complete collection of the 36-reported Arabidopsis thaliana and the rice Dof transcription factors identified herein. This analysis led to a classification into four major clusters of orthologous genes and showed gene loss and duplication events in Arabidopsis and rice, that occurred before and after the last common ancestor of the two species. CONCLUSIONS: According to our analysis, the Dof gene family in angiosperms is organized in four major clusters of orthologous genes or subfamilies. The proposed clusters of orthology and their further analysis suggest the existence of monocot specific genes and invite to explore their functionality in relation to the distinct physiological characteristics of these evolutionary groups.

Amino Acid Sequence↗

Efficient linkage of 10 loci in the proximal region of the mouse X chromosome.

Interspecific Mus species crosses were used to construct a multilocus genetic map of the mouse X chromosome that extends for more than 50 cM. In these studies, we established the segregation of eight loci in more than 200 backcross progeny from crosses of M. musculus and M. spretus with a common inbred strain (C57BL/6JRos). Genetic divergence at the level of the nucleotide sequences makes these crosses a useful cumulative genetic resource for mapping additional genes defined by genomic or cDNA probes in a highly efficient manner. We have therefore devised a mapping strategy that uses a subset of these backcrosses that are recombinant between successive anchor loci to both localize and order an additional set of six genes without necessarily resorting to an analysis of the entire backcross series. Using this approach, we have defined the linkage of cytochrome b245 beta-chain (Cybb), synapsin (Syn-1), and two members of the X-linked lymphocyte-regulated gene family (Xlr-1, Xlr-2), as well as DXSmh141 and DXSmh172, two loci defined by random genomic probes. All six loci have been localized to the proximal portion of the mouse X chromosome and their order has been defined as Cybb, Otc, Syn-1/Timp, DXSmh141/Xlr-1, DXSmh172, Hprt, Xlr-2, Cf-9. Gene order was established by minimizing multiple recombination events across the region spanning an estimated 20 cM of the proximal X chromosome. The possible significance of the Xlr loci is discussed with respect to other X-chromosome loci that regulate the immune response.

Animals↗

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof.

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

Arabidopsis↗

Evolutionary history of coastal tiger beetles in Japan based on a comparative phylogeography of four species.

To reveal the phylogeographical patterns of four species of coastal tiger beetles in Japan (Lophyridia angulata, Abroscelis anchoralis, Cicindela lewisii and Chaetodera laetescripta), we conducted phylogenetic and nested clade analysis (NCA) using the mitochondrial DNA sequences of two loci (COI and 16S rRNA), with specimens sampled from Japan and neighbouring countries. Abroscelis anchoralis and L. angulata have similar disjunct distributions in Japan. The NCA indicated past fragmentation involving three isolated areas of A. anchoralis. In contrast, local populations of L. angulata in Japan shared the same haplotype, indicating recent vicariance. Co-occurrence of haplotypes from several divergent clades in Japanese populations of Ch. laetescripta suggested ancient vicariance and subsequent intermixing of local populations. The tree topology of C. lewisii, with shallow branches and little geographical segregation of haplotypes between Japan and Korea or within Japan, suggested that the Japanese population was segregated from the Korean population only recently. Restricted gene flow, with isolation by distance, was inferred for various geographical associations of haplotypes for coastal tiger beetles in the NCA. Based on these phylogeographical patterns, coupled with a molecular clock approach, the evolutionary history of four species of coastal tiger beetles was deduced, with the additional consideration of the competitive relationships among those species. We also discuss the conservation of highly localized A. anchoralis populations in Japan, using the concept of evolutionarily significant units.

Animals↗

Gene-associated single nucleotide polymorphism discovery in perennial ryegrass (Lolium perenne L.).

Molecular genetic marker development in perennial ryegrass has largely been dependent on anonymous sequence variation. The availability of a large-scale EST resource permits the development of functionally-associated genetic markers based on SNP variation in candidate genes. Genic SNP loci and associated haplotypes are suitable for implementation in molecular breeding of outbreeding forage species. Strategies for in vitro SNP discovery through amplicon cloning and sequencing have been designed and implemented. Putative SNPs were identified within and between the parents of the F(1)(NA(6) x AU(6)) genetic mapping family and were validated among progeny individuals. Proof-of-concept for the process was obtained using the drought tolerance-associated LpASRa2 gene. SNP haplotype structures were determined and correlated with predicted amino acid changes. Gene-length LD was evaluated across diverse germplasm collections. A survey of SNP variation across 100 candidate genes revealed a high frequency of SNP incidence (c. 1 per 54 bp), with similar proportions in exons and introns. A proportion (c. 50%) of the validated genic SNPs were assigned to the F(1)(NA(6) x AU(6)) genetic map, showing high levels of coincidence with previously mapped RFLP loci. The perennial ryegrass SNP resource will enable genetic map integration, detailed LD studies and selection of superior allele content during varietal development.

Breeding↗

A new family of powerful multivariate statistical sequence analysis techniques.

A novel multivariate statistical approach is presented for extracting and exploiting intrinsic information present in our ever-growing sequence data banks. The information extraction from the sequences avoids the pitfalls of intersequence alignment by analyzing secondary invariant functions derived from the sequences in the data bank rather than the sequences themselves. Such typical invariant function is a 20 x 20 histogram of occurrences of amino acid pairs in a given sequence or fragment thereof. To illustrate the potential of the approach an analysis of 10,000 protein sequences from the National Biomedical Research Foundation Protein Identification Resource is presented, whose analysis already reveals great biological detail. For example, zeta-hemoglobin is found to lie close to amphibian and fish chi-hemoglobin which, in turn, is an important clue to the physiological function of this mammalian early embryonic hemoglobin. The multivariate statistical framework presented unifies such apparently unrelated issues as phylogenetic comparisons between a set of sequences and distance matrices between the constituents of the biological sequences. The Multivariate Statistical Sequence Analysis (MSSA) principles can be used for a wide spectrum of sequence analysis problems such as: assignment of family memberships to new sequences, validation of new incoming sequences to be entered into the database, prediction of structure from sequence, discrimination of coding from non-coding DNA regions, and automatic generation of an atlas of protein or DNA sequences. The MSSA techniques represent a self-contained approach to learning continuously and automatically from the growing stream of new sequences. The MSSA approach is particularly likely to play a significant role in major sequencing efforts such as the human genome project.

Amino Acid Sequence↗

Comparative genomics.

The genomes from three mammals (human, mouse, and rat), two worms, and several yeasts have been sequenced, and more genomes will be completed in the near future for comparison with those of the major model organisms. Scientists have used various methods to align and compare the sequenced genomes to address critical issues in genome function and evolution. This review covers some of the major new insights about gene content, gene regulation, and the fraction of mammalian genomes that are under purifying selection and presumed functional. We review the evolutionary processes that shape genomes, with particular attention to variation in rates within genomes and along different lineages. Internet resources for accessing and analyzing the treasure trove of sequence alignments and annotations are reviewed, and we discuss critical problems to address in new bioinformatic developments in comparative genomics.

Animals↗