Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Sequence databases and homology searching using the World Wide Web.

RNA, DNA and protein sequence data produced in laboratories around the world may be freely accessed by other scientists using the Internet. Originally, these sequence databases were compiled from published literature, but now researchers may deposit sequence data directly, using the World Wide Web. Almost 50,000 protein sequences and over 600,000 nucleotide sequences are available, forming a significant scientific resource. By comparing novel sequence data with sequence entries and structural data showing significant homology, it is possible to predict the function and three-dimensional structure of novel proteins.

Amino Acid Sequence↗

Genomic approaches to the genetics of alcoholism.

When studying complex diseases such as alcoholism that develop as a result of numerous genetic and environmental factors, researchers can use the sequence data that have become available both for the human and for animal genomes. For these analyses, investigators are being aided by efforts to identify and characterize functionally relevant DNA sequences in the entire genomic DNA sequence--a process called annotation. Various bioinformatics and annotation tools can help in this enterprise. These include four primary approaches: (1) precomputed, annotated public Web sites that provide a plethora of information; (2) in-house analyses from which users can choose the appropriate analyses for their purposes; (3) Web-based annotation systems that analyze a user's DNA sequence; and (4) private resources that provide access to annotated genomic sequences at cost. In addition to careful study of the DNA sequence for clues about function, expression studies of mRNA levels using gene chips provide information about the activity levels of thousands of genes that may vary in different tissues, different animals and people, or under different environmental conditions.

Alcoholism↗

Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.

The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using the ktup = 2 sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by setting ktup = 1 allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA with ktup = 1 on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.

Algorithms↗

Discovery of a large number of previously unrecognized mitochondrial pseudogenes in fish genomes.

Nuclear inserted copies of mitochondrial origin (numts) vary widely among eukaryotes, with human and plant genomes harboring the largest repertoires. Numts were previously thought to be absent from fish species, but the recent release of three fish nuclear genome sequences provides the resource to obtain a more comprehensive insight into the extent of mtDNA transfer in fishes. From the sequence analyses of the genomes of Fugu rubripes, Tetraodon nigroviridis, and Danio rerio, we have identified 2, 5, and 10 recent numt integrations, respectively, which integrated into those genomes less than 0.6 million years (Myr) ago. Such results contradict the hypothesis of absence or rarity of numts in fishes, as (i) the ratio of numts to the total size of the nuclear genome in T. nigroviridis was superior to the ratio observed in several higher vertebrate species (e.g., chicken, mouse, and rat), and only surpassed by humans, and (ii) the mtDNA coverage transferred to the nuclear genome of D. rerio is exceeded only by human and mouse, within the whole range of eukaryotic genomes surveyed for numts. Additionally, 335, 336, and 471 old numts (>12.5 Myr) were detected in F. rubripes, T. nigroviridis, and D. rerio, respectively. Surprisingly, old numts are inserted preferentially into known or predicted genes, as inferred for recent numts in human. However, because in fish genomes such integrations are old, they are likely to represent evolutionary successes and they may be considered a potential important evolutionary mechanism for the enhancement of genomic coding regions.

Animals↗

The Universal Protein Resource (UniProt).

The ability to store and interconnect all available information on proteins is crucial to modern biological research. Accordingly, the Universal Protein Resource (UniProt) plays an increasingly important role by providing a stable, comprehensive, freely accessible central resource on protein sequences and functional annotation. UniProt is produced by the UniProt Consortium, formed in 2002 by the European Bioinformatics Institute (EBI), the Protein Information Resource (PIR) and the Swiss Institute of Bioinformatics (SIB). The core activities include manual curation of protein sequences assisted by computational analysis, sequence archiving, development of a user-friendly UniProt web site and the provision of additional value-added information through cross-references to other databases. UniProt is comprised of three major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase and the UniProt Reference Clusters. An additional component consisting of metagenomic and environmental sequences has recently been added to UniProt to ensure availability of such sequences in a timely fashion. UniProt is updated and distributed on a bi-weekly basis and can be accessed online for searches or download at http://www.uniprot.org.

Amino Acid Sequence↗

Improving interoperability between microbial information and sequence databases.

BACKGROUND: Biological resources are essential tools for biomedical research. Their availability is promoted through on-line catalogues. Common Access to Biological Resources and Information (CABRI) is a service for distribution of biological resources and related data collected by 28 European culture collections. Linking this information to bioinformatics databanks can make the collections' holdings more visible after a search in molecular biology databanks and vice-versa. Identification of links to sequence databases can be useful, but annotation and indexing problems, together with compilation errors, immediately arise. In this paper, we present our efforts for the identification of cross-references between CABRI catalogues and the EMBL Data Library and related results. RESULTS: An SRS site with both EMBL and CABRI catalogues has been set up. Ad-hoc changes in indexing scripts allowed to achieve homogeneous index keys and SRS link features have been used to identify links between databases. After manual checking and comparison with an alternative procedure, about 67,500 valid cross-references were identified, added to the EMBL Data Library and are now distributed with it. HTML links can be established from EMBL to CABRI network service. Procedures can be executed whenever needed. CONCLUSION: Links between EMBL and CABRI catalogues constitute an improved access to micro-organisms of certified quality and can produce positive effects on biomedical research. Further links between CABRI catalogues and other bioinformatics databases can now easily be defined by using these cross-references. Linking genetic information onto natural resources information may stand model for the integration of other databases containing empirical data on these materials.

Base Sequence↗

A relational database of protein structures designed for flexible enquiries about conformation.

A relational database of protein structure has been developed to enable rapid and flexible enquiries about the occurrence of many aspects of protein architecture. The coordinates of 294 proteins from the Brookhaven Data Bank have been processed by standard computer programs to generate many additional terms that quantify aspects of protein structure. These terms include solvent accessibility, main-chain and side-chain dihedral angles, and secondary structure. In a relational database, the information is stored in tables with columns holding the different terms and rows holding the different entries for the terms. The different relational base tables store the information about the protein coordinate set, the different chains in the protein, the amino acid residues and ligands, the atomic coordinates, the salt bridges, the hydrogen bonds, the disulphide bridges and the close tertiary contacts. The database was established under ORACLE management system. Enquiries are constructed in ORACLE using SQL (structured query language) which is simple to use and alleviates the need for extensive computer programs. A single table can be searched for entries that meet various criteria, e.g. all protein solved to better than a given resolution. The power of the database occurs when several tables, or the entries in a single table, are cross-correlated. For example the dihedral angles of proline in the fourth position in an alpha-helix in high resolution structures can be rapidly obtained. The structural database provides a powerful tool to obtain empirical rules about protein conformation. This database of protein structures is part of a joint project between Birkbeck College and Leeds University to establish an integrated data resource of protein sequences and structures (ISIS) that encodes the complex patterns of residues and coordinates that define protein conformation. The entire data resource (ISIS) will provide a system to guide all areas of protein modelling including structure prediction, site-directed mutagenesis and de novo protein design. The availability of ISIS is described in the paper.

Computer Simulation↗

An analysis of the Sargasso Sea resource and the consequences for database composition.

BACKGROUND: The environmental sequencing of the Sargasso Sea has introduced a huge new resource of genomic information. Unlike the protein sequences held in the current searchable databases, the Sargasso Sea sequences originate from a single marine environment and have been sequenced from species that are not easily obtainable by laboratory cultivation. The resource also contains very many fragments of whole protein sequences, a side effect of the shotgun sequencing method.These sequences form a significant addendum to the current searchable databases but also present us with some intrinsic difficulties. While it is important to know whether it is possible to assign function to these sequences with the current methods and whether they will increase our capacity to explore sequence space, it is also interesting to know how current bioinformatics techniques will deal with the new sequences in the resource. RESULTS: The Sargasso Sea sequences seem to introduce a bias that decreases the potential of current methods to propose structure and function for new proteins. In particular the high proportion of sequence fragments in the resource seems to result in poor quality multiple alignments. CONCLUSION: These observations suggest that the new sequences should be used with care, especially if the information is to be used in large scale analyses. On a positive note, the results may just spark improvements in computational and experimental methods to take into account the fragments generated by environmental sequencing techniques.

Amino Acid Sequence↗

The Peptaibol Database: a database for sequences and structures of naturally occurring peptaibols.

The Peptaibol Database is a sequence and structure resource for the unusual class of peptides known as peptaibols. These peptides exhibit antibiotic and membrane channel-forming activities. The database includes sequence, biological source and bibliographical data for the naturally occurring peptaibols. Information is also collated for the growing number of peptaibol 3D structures determined by either crystallography or NMR spectroscopy. The database can be obtained as a whole or can be queried by name, group, sequence motif, biological origin and/or literature reference. The Peptaibol Database can be freely accessed at http://www.cryst.bbk.ac.uk/peptaibol.

Anti-Bacterial Agents↗

Simplified hepatitis C virus genotyping by heteroduplex mobility analysis.

Heteroduplex mobility analysis (HMA) was used to genotype hepatitis C viruses (HCV) with PCR fragments derived from the 5' untranslated region (5'-UTR) or the NS5b region. HCV 5'-UTR fragments were amplified from 296 serum samples by use of a combined reverse transcription-PCR assay, and the genotypes of isolates were determined by sequencing. HCV genotype distributions in Australia were 39% for genotype 1a, 15% for 1b, 3% for 1a/b, <1% for 2a/c, 5% for 2b, 34% for 3a, <1% for 3b, and 1% for 4, and 1% of patients were infected with more than one genotype. Pairwise HMA of subtypes 1a, 1b, 2a/c, 2b, 3a, 3b, 4a, and 6a demonstrated that five distinct heteroduplex patterns were formed between the eight subtypes. A reference panel that contained a representative of each pattern (1a, 2b, 3a, 4a, and 6a) was used for genotyping. The pattern of heteroduplexes formed when a test isolate was mixed with the five reference isolates was correlated with the genotype, as determined by sequencing. Genotypes determined by HMA correlated exactly with sequencing results within the groups 1, 2, 3a, 3b/4, and 6. HMA was also used to simplify the identification of mixed infection with two HCV genotypes. In further studies, with amplicons from the NS5b region, HMA classified isolates into their respective subtypes, and the heteroduplex mobility ratio correlated closely with nucleotide sequence variation at the isolate, subtype, and genotype levels. HMA provides an adaptable, inexpensive, and rapid method of genotyping HCV that requires fewer resources than DNA sequencing.

5' Untranslated Regions↗

Structure of the mitochondrial control region of the Eurasian otter (Lutra lutra; Carnivora, Mustelidae): patterns of genetic heterogeneity and implications for conservation of the species in Italy.

In this study we determined the complete sequence of the mitochondrial DNA (mtDNA) control region of the Eurasian otter (Lutra lutra). We then compared these new sequences with orthologues of nine carnivores belonging to six families (Mustelidae, Mephitidae, Canidae, Hyaenidae, Ursidae, and Felidae). The comparative analyses identified all the conserved regions previously found in mammals. The Eurasian otter and seven other species have a single location with tandem repeats in the right domain, while the spotted hyena (Hyaenidae) and the tiger (Felidae) have repeated sequences in both the right and left domains. To assess the degree of genetic heterogeneity of the Eurasian otter in Italy we sequenced two fragments of the gene and analyzed length polymorphisms of repeated sequences and heteroplasmy in 32 specimens. The study includes 23 museum specimens collected in northern, central, and southern Italy; most of these specimens are from extinct populations, while the southern Italian samples belong to the sole extant Italian population of the Eurasian otter. The study also includes all the captive-reared animals living in the colony "Centro Lontra, Caramanico Terme" (Pescara, central Italy). The colony is maintained for reintroduction of the species. We found a low level of genetic polymorphism; a single haplotype is dominant, but our data indicate the presence in central and southern Italy of two slightly divergent haplotypes. One haplotype belongs to an extinct population, the other is present in the single extant Italian population. Analyses of length polymorphisms and heteroplasmy indicate that the autochthonous Italian samples are characterized by a distinct array of repeated sequences from captive-reared animals.

Animals↗

Mitochondrial genome-derived microsatellites reveal genetic diversity and population structure in Callery pear populations.

Callery pear (Pyrus calleryana Decne.; PC) possesses many desirable characteristics valued in managed landscapes. This has driven the release of numerous cultivars, including both hybrids and selections derived from native populations. The extensive planting of PC cultivars in managed areas has contributed to the widespread occurrence of invasive individuals across a broad range of habitats in the eastern United States (US). Self-incompatibility, tolerance to various environmental conditions, pathogen and pest resistance, intraspecific hybridization among the cultivars, possible interspecific hybridization with other Pyrus species, and seed dispersal by various vertebrates have contributed to the spread and persistence of PC across diverse environments. Because effective and environmentally appropriate management options remain limited, improved understanding of PC genetics may help inform management strategies. Previous studies have characterized PC diversity using nuclear genomic short sequence repeats (gSSRs), however, neither a mitochondrial genome resource nor mitochondrial short sequence repeats (mtSSRs) have been developed for this purpose. Here, we assembled a mitochondrial genome of 485,892 bp and used five mtSSRs to characterize mitochondrial&#xa0;diversity and population structure among accessions from the species' native range in Asia (n&#x2009;=&#x2009;72), southeastern US escapees (SNesc; n&#x2009;=&#x2009;90), Tennessee escapees (TNesc; n&#x2009;=&#x2009;90), and US-released commercial cultivars (UScult; n&#x2009;=&#x2009;69 representing 14 unique cultivars). We found a high genetic diversity (He&#x2009;=&#x2009;0.728) and evidence of genetic structure in PC. In distance-based and multivariate analyses, UScult occupied an intermediate position between the Asian populations and the US escapees. The observed mitochondrial diversity among samples assigned to PC cultivars is consistent with a complex genetic landscape and may reflect distinct maternal lineages, cultivar-labeling or record-keeping discrepancies, and/or technical variation. This study underscores the need for broader genomic investigations using authenticated cultivar reference material and high-resolution nuclear markers to resolve cultivar ancestry, validate true-to-name identity, and inform species management.

Genetic Variation↗

Multi-species comparative mapping in silico using the COMPASS strategy.

MOTIVATION: The completion of human and mouse genome sequences provides a valuable resource for decoding other mammalian genomes. The comparative mapping by annotation and sequence similarity (COMPASS) strategy takes advantage of the resource and has been used in several genome-mapping projects. It uses existing comparative genome maps based on conserved regions to predict map locations of a sequence. An automated multiple-species COMPASS tool can facilitate in the genome sequencing effort and comparative genomics study of other mammalian species. RESULTS: The prerequisite of COMPASS is a comparative map table between the reference genome and the predicting genome. We have built and collected comparative maps among five species including human, cattle, pig, mouse and rat. Cattle-human and pig-human comparative maps were built based on the positions of orthologous markers and the conserved synteny groups between human and cattle and human and pig genomes, respectively. Mouse-human and rat-human comparative maps were based on the conserved sequence segments between the two genomes. With a match to human genome sequences, the approximate location of a query sequence can be predicted in cattle, pig, mouse and rat genomes based on the position of the match relatively to the orthologous markers or the conserved segments. AVAILABILITY: The COMPASS-tool and databases are available at http://titan.biotec.uiuc.edu/COMPASS/

Algorithms↗

The Eukaryotic Promoter Database EPD.

The Eukaryotic Promoter Database (EPD) is an annotated non-redundant collection of experimentally characterised eukaryotic POL II promoters. The underlying definition of a promoter is that of a transcription initiation site. All information presented in EPD results from an independent evaluation of primary experimental data shown in the biological literature. Sequences flanking transcription initiation sites are indirectly given by pointers to EMBL sequences. The annotation part of a promoter entry includes description of the promoter-defining evidence, cross-references to other databases, and bibliographic references. Being designed as a resource for comparative sequence analysis, EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets. The database is available through the World Wide Web at URL http://cmpteam4.unil.ch

Animals↗

The ASTRAL Compendium in 2004.

The ASTRAL Compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. Partially derived from the SCOP database of protein structure domains, it includes sequences for each domain and other resources useful for studying these sequences and domain structures. The current release of ASTRAL contains 54,745 domains, more than three times as many as the initial release 4 years ago. ASTRAL has undergone major transformations in the past 2 years. In addition to several complete updates each year, ASTRAL is now updated on a weekly basis with preliminary classifications of domains from newly released PDB structures. These classifications are available as a stand-alone database, as well as integrated into other ASTRAL databases such as representative subsets. To enhance the utility of ASTRAL to structural biologists, all SCOP domains are now made available as PDB-style coordinate files as well as sequences. In addition to sequences and representative subsets based on SCOP domains, sequences and subsets based on PDB chains are newly included in ASTRAL. Several search tools have been added to ASTRAL to facilitate retrieval of data by individual users and automated methods. ASTRAL may be accessed at http://astral.stanford. edu/.

Animals↗

LGICdb: a manually curated sequence database after the genomes.

Ligand-gated ion channels form transmembrane ionic pores controlled by the binding of chemicals. The LGICdb aims to be a non-redundant, manually curated resource offering access to the large number of subunits composing extracellularly activated ligand-gated ion channels, such as nicotinic, ATP, GABA and glutamate ionotropic receptors. Composed of more than 500 human curated entries, the XML native database has been relocated in 2004 to the EBI. Its facilities have been enhanced with a new search system, customized multiple sequence alignments and manipulation of protein structures (http://www.ebi.ac.uk/compneur-srv/LGICdb/). Despite the vast improvement of general sequence resources, the LGICdb still provide sequences unavailable elsewhere.

Databases, Protein↗

A comparative study on earthworm hemoglobins: an amino acid sequence comparison of monomer globin chains of two species, Pontodrilus matsushimensis and Pheretima communissima that belong to the family Megascolecidae.

The monomer subunits of giant extracellular hemoglobins from earthworms Pontodrilus matsushimensis and Pheretima communissima that belong to the family Megascolecidae, Oligochaeta, were purified by a reversed-phase column, Resource RPC, and sequenced. The complete amino acid sequences of the two monomeric globin chains were determined: 141 amino acid residues with a molecular weight of 16,366 Da for Pontodrilus matsushimensis and 140 amino acid residues with a molecular weight of 16,000 for Pheretima communissima, respectively. The Pontodrilus matsushimensis monomer globin has three cysteine residues, and the two located at positions 2 and 131 are conserved as those observed in all annelids and contribute to form a disulfide-bonded interchain. The third cysteine residue at position 73 is the first evidence for the annelid monomer globin subunits. The physiological functions of the third cysteine residue, however, are still unknown. The monomer sequences of the two species were aligned with those of five known sequences from annelids, including a polychaete, Tylorrhynchus heterochaetus, and four oligochaetes, Pheretima hilgendorfi, Pheretima sieboldi, Lumbricus terrestris and Tubifex tubifex. Using computer analysis, a 87.9% identity of the amino acid sequences between two monomeric subunits of Pheretima communissima and Pheretima hilgendorfi hemoglobins showed the highest degree of sequence similarity. A molecular phylogenetic tree of seven species of annelids has constructed, suggesting that the divergence times among the three species of Pheretima and between Pheretima and Pontodrilus were 50 to 100 and about 209 million years ago, respectively.

Amino Acid Sequence↗

A front-end desaturase from Chlamydomonas reinhardtii produces pinolenic and coniferonic acids by omega13 desaturation in methylotrophic yeast and tobacco.

Pinolenic acid (PA; 18:3Delta(5,9,12)) and coniferonic acid (CA; 18:4Delta(5,9,12,15)) are Delta(5)-unsaturated bis-methylene-interrupted fatty acids (Delta(5)-UBIFAs) commonly found in pine seed oil. They are assumed to be synthesized from linoleic acid (LA; 18:2Delta(9,12)) and alpha-linolenic acid (ALA; 18:3Delta(9,12,15)), respectively, by Delta(5)-desaturation. A unicellular green microalga Chlamydomonas reinhardtii also accumulates PA and CA in a betain lipid. The expressed sequence tag (EST) resource of C. reinhardtii led to the isolation of a cDNA clone that encoded a putative fatty acid desaturase named as CrDES containing a cytochrome b5 domain at the N-terminus. When the coding sequence was expressed heterologously in the methylotrophic yeast Pichia pastoris, PA and CA were newly detected and comparable amounts of LA and ALA were reduced, demonstrating that CrDES has Delta(5)-desaturase activity for both LA and ALA. CrDES expressed in the yeast showed Delta(5)-desaturase activity on 18:1Delta(9) but not 18:1Delta(11). Unexpectedly, CrDES also showed Delta(7)-desaturase activity on 20:2Delta(11,14) and 20:3Delta(11,14,17) to produce 20:3Delta(7,11,14) and 20:4Delta(7,11,14,17), respectively. Since both the Delta(5) bond in C18 and the Delta(7) bond in C20 fatty acids are 'omega13' double bonds, these results indicate that CrDES has omega13 desaturase activity for omega9 unsaturated C18/C20 fatty acids, in contrast to the previously reported front-end desaturases. In order to evaluate the activity of CrDES in higher plants, transgenic tobacco plants expressing CrDES were created. PA and CA accumulated in the leaves of transgenic plants. The highest combined yield of PA and CA was 44.7% of total fatty acids, suggesting that PA and CA can be produced in higher plants on a large scale.

Amino Acid Sequence↗