Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Assembly and analysis of the mouse immunoglobulin kappa gene sequence.

The mechanisms regulating V gene usage leading to the immunoglobulin (Ig) repertoire have been of interest for many years but are only partially defined. To gain insight into these processes, we have assembled the nucleotide sequence of the Mus musculus Igkappa locus using data recently made available from genome-wide sequencing efforts. We found the locus to be 3.21 Mb in length and mapped all known functional, pseudo- and relic V gene segments onto the sequence, along with known regulatory elements. We corrected errors in former gene assignments, positions and orientations and identified a novel Vkappa4 gene segment. This assembly allowed the establishment of a unified nomenclature for the V genes based on their relative positions similar to the nomenclature system adopted for the human Ig loci. The 5' boundary of the locus is defined by the presence of the tumor-associated calcium-signal transducer-2 gene located 19 kb upstream of Vkappa24-140, the most distal V gene. No non- Vkappa genes were found in the sequence of the locus. Detailed analysis of the sequences 0.5 kb upstream, within, and 0.5 kb downstream of each potentially functional V gene revealed interesting patterns of statistically significant clustering of transcription factor consensus binding sites, generally specific to a particular family. We found E boxes were clustered not only in promoter regions, but also nearby recombination signal sequences. Family members of Vkappa4/5 genes exhibit a conserved pattern of octamer sites in their downstream regions, as well as Ebf sites in their introns, and Lef-1 sites in their upstream regions. We discuss potential functional implications of these findings in the context of possible combinatorial mechanisms for targeting V genes for rearrangement. The assembled sequence and its analyses are available as a resource to the scientific community.

Animals↗

Expression profiling of the Leishmania life cycle: cDNA arrays identify developmentally regulated genes present but not annotated in the genome.

As genomic sequencing of Leishmania nears completion, functional analyses that provide a global genetic perspective on biological processes are important. Despite polycistronic transcription, RNA transcript abundance can be measured using microarrays. To provide a resource to evaluate cDNA arrays, we undertook 5' expressed sequence tag analysis of 2183 full-length randomly selected cDNAs from Leishmania major promastigote (days 3, 7, 10 of culture in vitro), and lesion-derived amastigote libraries. PCR-amplified inserts from 1830 of these cDNA representing 1001 unique genes were spotted onto microarrays, and compared internally with PCR-amplified open reading frames (ORFs) from 904 genes representing 842 unique genes annotated in the L. major genome. Microarrays were screened with RNA from procyclic, metacyclic and amastigote populations of L. major. Redundant clones on the array gave highly reproducible results, providing confidence in identification of stage-specific gene expression. Four hundred and thirty unique (i.e. non-redundant) stage-specific genes were identified. A higher percentage of stage-specific gene expression was observed in amastigotes ( approximately 35%) compared to metacyclics ( approximately 12%) for both cDNAs and ORFs, but cDNAs provided a richer source of regulated genes than currently annotated ORFs from the Leishmania genome. In mapping cDNAs onto the Leishmania genome, we noted that approximately 42% aligned to regions not recognised as genes using current predictive annotation tools. These genes are highly represented in our stage-specific genes, and therefore represent important drug targets and vaccine candidates. Careful annotation of cDNAs onto the Leishmania genome will be important before producing the next generation of oligonucleotide arrays based on annotated genes of the genomic sequencing project.

Animals↗

The discovery and confirmation of single nucleotide polymorphisms in the human p53R2 gene by EST database analysis.

The human expressed sequence tag (EST) database provides a wealth of resources, which can be used to rapidly screen for potential polymorphisms in proteins of physiological interest. The human p53R2 gene, a recently identified ribonucleotide reductase, plays an important role in DNA repair and is involved in the pathway of p53 activity in response to the presence of DNA damage. On the basis of the alignment of human EST sequences, we identified three candidate polymorphisms at nt 2752, 2759 and 4696 in the 3'-untranslated region of the p53R2 gene. The presence of these polymorphisms was confirmed in a Caucasian population (n = 82) by allele-specific PCR and PCR/restriction fragment length polymorphism analyses. The rare allele frequency at position 4696 (15.5%) is higher than either rare allele frequency at position 2752 or 2759 (6 and 6%). Our results suggest that the human EST data may serve as a valuable source for the rapid identification of genetic variation.

3' Untranslated Regions↗

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024 bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial↗

Isolation of rhizobia from Ontario soils that are effective at fixing nitrogen with common bean (Phaseolus vulgaris).

UNLABELLED: Common bean (Phaseolus vulgaris) is an important crop in Canada and globally. Like other legumes, common bean establishes symbiotic interactions with nitrogen-fixing bacteria called rhizobia. However, nitrogen fixation by rhizobia in association with common bean is often suboptimal, constraining its productivity and necessitating the application of nitrogen fertilizer. To support the development of high-performing, locally adapted rhizobial inoculants for Ontario common bean growers, we isolated 216 common bean-nodulating rhizobia from southern Ontario soils using a nodule trapping approach with four common bean cultivars. Whole genome sequencing followed by phylogenomic analyses of all rhizobial isolates revealed substantial diversity, assigning them to 11 Rhizobium species, including two novel species. Nearly all isolates belong to the symbiovar phaseoli, spanning the nodC γ-a, γ-b, and α alleles, with four isolates belonging to the symbiovar gallica. Soil origin had a significant impact on the species-level community composition recovered during the nodule trapping experiments. In contrast, host trapping cultivar had only a minor influence on the recovered Rhizobium population. Greenhouse assays demonstrated that one of the novel Rhizobium species exhibited the highest average symbiotic effectiveness, although high-quality isolates were found across multiple species. Together, these results revealed a diverse and genomically variable Rhizobium community capable of forming effective symbioses with common bean in southern Ontario soils. Importantly, our genome-sequenced Rhizobium collection will serve as a valuable resource for identifying competitive and high-quality strains for the development of inoculants tailored to Ontario common bean production. IMPORTANCE: Common bean is a globally important food crop, yet its productivity is often limited by suboptimal nitrogen fixation, forcing growers to rely on synthetic fertilizers. Consequently, identifying high‑performing, locally adapted inoculant strains is essential for reducing dependence on synthetic nitrogen fertilizers and improving the sustainability of temperate agroecosystems. Our study provides a genome‑sequenced collection of common bean-nodulating Rhizobium from southern Ontario, revealing substantial species and genomic diversity across sampling locations. Greenhouse studies allowed us to identify multiple isolates that consistently fix nitrogen with, and enhance the growth of, common bean plants. Our findings highlight strong biogeographical structuring of the effective and competitive subpopulations of rhizobial communities and demonstrate that Ontario soils already harbor strains with high symbiotic potential. In addition, our Rhizobium collection represents a foundational resource to support future inoculant development and enables future work on the ecology, evolution, and applied optimization of legume-rhizobium symbioses.

Nanopore↗

Removing near-neighbour redundancy from large protein sequence collections.

MOTIVATION: To maximize the chances of biological discovery, homology searching must use an up-to-date collection of sequences. However, the available sequence databases are growing rapidly and are partially redundant in content. This leads to increasing strain on CPU resources and decreasing density of first-hand annotation. RESULTS: These problems are addressed by clustering closely similar sequences to yield a covering of sequence space by a representative subset of sequences. No pair of sequences in the representative set has >90% mutual sequence identity. The representative set is derived by an exhaustive search for close similarities in the sequence database in which the need for explicit sequence alignment is significantly reduced by applying deca- and pentapeptide composition filters. The algorithm was applied to the union of the Swissprot, Swissnew, Trembl, Tremblnew, Genbank, PIR, Wormpep and PDB databases. The all-against-all comparison required to generate a representative set at 90% sequence identity was accomplished in 2 days CPU time, and the removal of fragments and close similarities yielded a size reduction of 46%, from 260 000 unique sequences to 140 000 representative sequences. The practical implications are (i) faster homology searches using, for example, Fasta or Blast, and (ii) unified annotation for all sequences clustered around a representative. As tens of thousands of sequence searches are performed daily world-wide, appropriate use of the non-redundant database can lead to major savings in computer resources, without loss of efficacy. AVAILABILITY: A regularly updated non-redundant protein sequence database (nrdb90), a server for homology searches against nrdb90, and a Perl script (nrdb90.pl) implementing the algorithm are available for academic use from http://www.embl-ebi.ac. uk/holm/nrdb90. CONTACT: holm@embl-ebi.ac.uk

Algorithms↗

Complete set of eleven region-specific microdissection libraries for human chromosome 2.

The construction and characterization of 11 region-specific libraries for the entire human chromosome 2 have been completed, including four libraries for the short arm and six libraries for the long arm, plus a library for the centromere region. These libraries were constructed using the chromosome microdissection and microcloning technology. Eight libraries have been described previously. This paper presents the final three libraries: 2q21-q22 (designated 2Q5 library), 2q11-q14 (2Q6). and 2p11.1-q11.1 (2CEN). The sizes of the dissected regions ranged between 20 and 30 Mb, with the centromere region of about 4 Mb. All these libraries are large, potentially comprising hundreds of thousands of recombinant microclones. Between 77% and 97% of the microclones were shown to derive from respective dissected regions. From 26 to 66 unique sequence microlones were isolated and characterized in detail for each library. The microclones have short inserts, ranging between 50 and 600 bp, with a mean of about 200 bp. The short inserts can be conveniently sequenced as STSs to provide high density probes for the dissected region. A plasmid sub-library containing at least 20,000 microclones, and usually more, has been prepared from each library and deposited to ATCC for general distribution. The libraries have been used effectively in constructing high resolution physical maps and for contig assembly, as well as in positional cloning of disease genes assigned to the dissected region. Comparing to other chromosomes with detailed mapping information and densely populated probes, chromosome 2 remains largely under-exploited. The availability of a complete set of region-specific libraries and unique sequence microclones from the libraries should provide valuable resources for genome analysis, high resolution physical mapping, region-specific cDNA isolation, and positional cloning for chromosome 2.

Centromere↗

A Web-based classification system of DNA-binding protein families.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The family of DNA-binding proteins is one of the most populated and studied amongst the various genomes of bacteria, archaea and eukaryotes and the Web-based system presented here is an approach to their classification. The DnaProt resource is an annotated and searchable collection of protein sequences for the families of DNA-binding proteins. The database contains 3238 full-length sequences (retrieved from the SWISS-PROT database, release 38) that include, at least, a DNA-binding domain. Sequence entries are organized into families defined by PROSITE patterns, PRINTS motifs and de novo excised signatures. Combining global similarities and functional motifs into a single classification scheme, DNA-binding proteins are classified into 33 unique classes, which helps to reveal comprehensive family relationships. To maximize family information retrieval, DnaProt contains a collection of multiple alignments for each DNA-binding family while the recognized motifs can be used as diagnostically functional fingerprints. All available structural class representatives have been referenced. The resource was developed as a Web-based management system for online free access of customized data sets. Entries are fully hyperlinked to facilitate easy retrieval of the original records from the source databases while functional and phylogenetic annotation will be applied to newly sequenced genomes. The database is freely available for online search of a library containing specific patterns of the identified DNA-binding protein classes and retrieval of individual entries from our WWW server (http://kronos.biol.uoa.gr/~mariak/dbDNA.html).

Amino Acid Motifs↗

Bioinformatics Resources for In Silico Proteome Analysis.

In the growing field of proteomics, tools for the in silico analysis of proteins and even of whole proteomes are of crucial importance to make best use of the accumulating amount of data. To utilise this data for healthcare and drug development, first the characteristics of proteomes of entire species-mainly the human-have to be understood, before secondly differentiation between individuals can be surveyed. Specialised databases about nucleic acid sequences, protein sequences, protein tertiary structure, genome analysis, and proteome analysis represent useful resources for analysis, characterisation, and classification of protein sequences. Different from most proteomics tools focusing on similarity searches, structure analysis and prediction, detection of specific regions, alignments, data mining, 2D PAGE analysis, or protein modelling, respectively, comprehensive databases like the proteome analysis database benefit from the information stored in different databases and make use of different protein analysis tools to provide computational analysis of whole proteomes.

Journal Article↗

Cloning of a Twist orthologue from Enchytraeus coronatus (Annelida, Oligochaeta).

Enchytraeus coronatus is a small soil living oligochaete that can be maintained in culture with ease. Embryos are laid into cocoons, where they develop directly into a hatching worm within about two weeks. E. coronatus shows a simple morphology. Its transparency allows the microscopic analysis of developmental processes. To facilitate future studies on the development of specific tissues, like the mesoderm, we established molecular techniques and resources, like cDNA libraries, to allow the cloning of genes that are potentially relevant during development of the oligochaete. In this paper, we present first results using these new resources and describe the cloning of the Twist orthologue from E. coronatus. The Enchytraeus Twist protein (EcTwist) harbors all characteristic sequence features common for a true Twist orthologue. We believe that the resources described herein will facilitate phylogenetic studies on the molecular level, which will help to understand lophotrochozoan evolution.

Amino Acid Sequence↗

The NEIBank project for ocular genomics: data-mining gene expression in human and rodent eye tissues.

NEIBank is a project to gather and organize genomic resources for eye research. The first phase of this project covers the construction and sequence analysis of cDNA libraries from human and animal model eye tissues to develop an overview of the repertoire of genes expressed in the eye and a resource of cDNA clones for further studies. The sequence data are grouped and identified using the tools of bioinformatics and the results are displayed through a web site where they can be interrogated by keyword search, chromosome location, by Blast (sequence comparison) or by alignment on completed genomes. Many novel proteins and novel splice forms of known genes have already emerged from analysis of the accumulating data. This review provides an overview of the current state of the database for human eye tissues, with specific comparisons to some parallel data from mouse and rat, and with illustrative examples of the kinds of insights and discoveries these data can produce. One of the major themes that emerges is that at the molecular level human eye tissues have significant differences from those of rodents, encompassing species specific genes, alternative splice forms and great variation in levels of gene expression. These point to specific adaptations and mechanisms in the human eye and emphasize that care needs to be taken in the application of appropriate animal model systems.

Amino Acid Sequence↗

Mass spectrometry and the age of the proteome.

Mass spectrometry has become an important technique to correlate proteins to their genes. This has been achieved, in part, by improvements in ionization and mass analysis techniques concurrently with large-scale DNA sequencing of whole genomes. Genome sequence information has provided a convenient and powerful resource for protein identification using data produced by matrix-assisted laser desorption/ionization time-of-flight (MALDI/TOF) and tandem mass spectrometers. Both of these approaches have been applied to the identification of electrophoretically separated protein mixtures. New methods for the direct identification of proteins in mixtures using a combination of enzymatic proteolysis, liquid chromatographic separation, tandem mass spectrometry and computer algorithms which match peptide tandem mass spectra to sequences in the database are also emerging. This tutorial review describes the principles of ionization and mass analysis for peptide and protein analysis and then focuses on current methods employing MALDI and electrospray ionization for protein identification and sequencing. Database searching approaches to identify proteins using data produced by MALDI/TOF and tandem mass spectrometry are also discussed.

Amino Acid Sequence↗

Bioinformatics and the malaria genome: facilitating access and exploitation of sequence information.

The torrent of sequence information unleashed by the various genome sequencing projects, including that of Plasmodium falciparum, will lead to an unprecedented increase in the data available for research purposes. The scientific community is struggling to develop ways to assimilate this information and ensure that it is fully analysed in a way that enables rapid development of new therapeutic and diagnostic advances. This is particularly so for the field of tropical medicine where many of the scientists have had limited training in the area of Bioinformatics and may be further hampered by poor access to the sequence data. A number of collections of malaria genome sequence are available, each with their own advantages and disadvantages, however further improvements in these information resources are needed. In particular, there would be great benefit in integrating genomic sequence and functional genomics results with the large amount of pre-existing knowledge related to parasite biology and immunological interactions with the host. Attempts to achieve this include the PlasmoDB database, and the lessons learned in this effort could be of great utility to other organism-specific databases.

Animals↗

Conservation and evolution of microsatellite loci in primate taxa.

Microsatellites are promising genetic markers for the study of demographic structure and phylogenetic history in populations. However, little information exists on the molecular nature of the repeats and their flanking sequences of a same microsatellite in a large range of species. In this study, we report polymorphism and consensus sequences of eight microsatellite loci using human primers in 20 primate species. The results show size polymorphism in almost all species and microsatellites. These loci are therefore useful markers for population genetic studies between populations of the same species. Insertion/deletion events are frequent in the flanking regions, the majority concerning several contiguous bases. This is in contrast with the more usual single base pair events in non-coding regions. The ranges of allele lengths in non-human primates often show no overlap with that of human, usually due to the deletion/insertion events in the flanking sequences, producing smaller allele lengths rather than smaller numbers of repeats. The use of length of PCR product will bias the inter-species interpretation reducing the number of observable alleles and treating as the same allele very divergent molecular sequences. Caution should be used when employing microsatellites in cross-species comparisons in which the species under study are separated by significant amounts of evolutionary time: in such cases allele comparison cannot be based on lengths alone.

Animals↗

Analysis of gene expression in rose petals using expressed sequence tags.

Single-pass sequences were obtained from the 5'-ends of a total of 1794 rose petal cDNA clones. Cluster analysis identified 242 groups of sequences and 635 singletons indicating that the database represents a total of 877 genes. Putative functions could be assigned to 1151 of the transcripts. Expression analysis indicated that transcripts of several of the genes identified accumulated specifically in petals and stamens. The cDNA library and expressed sequence tag database described here represent a valuable resource for future research aimed at improving economically important rose characteristics such as flower form, longevity and scent.

Contig Mapping↗

Phylogeography of the endangered darkling beetle species of Pimelia endemic to Gran Canaria (Canary Islands).

Phylogenetic and geographical nested clade analysis (NCA) methods were applied to mitochondrial DNA sequences of Pimelia darkling beetles (Coleoptera, Tenebrionidae) endemic to Gran Canaria, an island in the Canary archipelago. The three species P. granulicollis, P. estevezi and P. sparsa occur on the island, the latter with three recognized subspecies. Another species, P. fernandezlopezi (endemic to the island of La Gomera) is a close relative of P. granulicollis based on partial Cytochrome Oxidase I mtDNA sequences obtained in a previous study. Some of these beetles are endangered, so phylogeographical structure within species and populations can help to define conservation priorities. A total of about 700 bp of Cytochrome Oxidase II were examined in 18 populations and up to 75 individuals excluding outgroups. Among them, 22 haplotypes were exclusive to P. granulicollis and P. estevezi and 31 were from P. sparsa. Phylogenetic analysis points to the paraphyly of Gran Canarian Pimelia, as the La Gomera P. fernandezlopezi haplotypes are included in them, and reciprocal monophyly of two species groups: one constituted by P. granulicollis, P. estevezi and P. fernandezlopezi (subgenus Aphanaspis), and the other by P. sparsa'sensu lato'. The two species groups show a remarkably high mtDNA divergence. Within P. sparsa, different analyses all reveal a common result, i.e. conflict between current subspecific taxonomic designations and evolutionary units, while P. estevezi and P. fernandezlopezi are very close to P. granulicollis measured at the mtDNA level. Geographical NCA identifies several cases of nonrandom associations between haplotypes and geography that may be caused by allopatric fragmentation of populations with some cases of restriction of gene flow or range expansion. Analyses of molecular variance and geographical NCA allow definition of evolutionary units for conservation purposes in both species-groups and suggest scenarios in which vicariance caused by geological history of the island may have shaped the pattern of the mitochondrial genetic diversity of these beetles.

Animals↗

Global population structure and taxonomy of the wandering albatross species complex.

A recent taxonomic revision of wandering albatross elevated each of the four subspecies to species. We used mitochondrial DNA and nine microsatellite markers to study the phylogenetic relationships of three species (Diomedea antipodensis, D. exulans and D. gibsoni) in the wandering albatross complex. A small number of samples from a fourth species, D. dabbenena, were analysed using mitochondrial DNA only. Mitochondrial DNA sequence analyses indicated the presence of three distinct groups within the wandering albatross complex: D. exulans, D. dabbenena and D. antipodensis/D. gibsoni. Although no fixed differences were found between D. antipodensis and D. gibsoni, a significant difference in the frequency of a single restriction site was detected using random fragment length polymorphism. Microsatellite analyses using nine variable loci, showed that D. exulans, D. antipodensis and D. gibsoni were genetically differentiated. Despite the widespread distribution of D. exulans, we did not detect any genetic differentiation among populations breeding on different island groups. The lower level of genetic differentiation between D. antipodensis and D. gibsoni should be reclassified as D. antipodensis. Within the context of the current taxonomy, these combined data support three species: D. dabbenena, D. exulans and D. antipodensis.

Animals↗

PANDIT: an evolution-centric database of protein and associated nucleotide domains with inferred trees.

PANDIT is a database of homologous sequence alignments accompanied by estimates of their corresponding phylogenetic trees. It provides a valuable resource to those studying phylogenetic methodology and the evolution of coding-DNA and protein sequences. Currently in version 17.0, PANDIT comprises 7738 families of homologous protein domains; for each family, DNA and corresponding amino acid sequence multiple alignments are available together with high quality phylogenetic tree estimates. Recent improvements include expanded methods for phylogenetic tree inference, assessment of alignment quality and a redesigned web interface, available at the URL http://www.ebi.ac.uk/goldman-srv/pandit.

Databases, Nucleic Acid↗