Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Comparative genomics of Staphylococcus aureus musculoskeletal isolates.

Much of the research aimed at defining the pathogenesis of Staphylococcus aureus has been done with a limited number of strains, most notably the 8325-4 derivative RN6390. Several lines of evidence indicate that this strain is unique by comparison to clinical isolates of S. aureus. Based on this, we have focused our efforts on two clinical isolates (UAMS-1 and UAMS-601), both of which are hypervirulent in our animal models of musculoskeletal infection. In this study, we used comparative genomic hybridization to assess the genome content of these two isolates relative to RN6390 and each of seven sequenced S. aureus isolates. Our comparisons were done by using an amplicon-based microarray from the Pathogen Functional Genomics Resource Center and an Affymetrix GeneChip that collectively represent the genomes of all seven sequenced strains. Our results confirmed that UAMS-1 and UAMS-601 share specific attributes that distinguish them from RN6390. Potentially important differences included the presence of cna and the absence of isaB, sarT, sarU, and sasG in the UAMS isolates. Among the sequenced strains, the UAMS isolates were most closely related to the dominant European clone EMRSA-16. In contrast, RN6390, NCTC 8325, and COL formed a distinct cluster that, by comparison to the other four sequenced strains (Mu50, N315, MW2, and SANGER-476), was the most distantly related to the UAMS isolates and EMRSA-16.

Adhesins, Bacterial↗

VKCDB: voltage-gated potassium channel database.

BACKGROUND: The family of voltage-gated potassium channels comprises a functionally diverse group of membrane proteins. They help maintain and regulate the potassium ion-based component of the membrane potential and are thus central to many critical physiological processes. VKCDB (Voltage-gated potassium [K] Channel DataBase) is a database of structural and functional data on these channels. It is designed as a resource for research on the molecular basis of voltage-gated potassium channel function. DESCRIPTION: Voltage-gated potassium channel sequences were identified by using BLASTP to search GENBANK and SWISSPROT. Annotations for all voltage-gated potassium channels were selectively parsed and integrated into VKCDB. Electrophysiological and pharmacological data for the channels were collected from published journal articles. Transmembrane domain predictions by TMHMM and PHD are included for each VKCDB entry. Multiple sequence alignments of conserved domains of channels of the four Kv families and the KCNQ family are also included. Currently VKCDB contains 346 channel entries. It can be browsed and searched using a set of functionally relevant categories. Protein sequences can also be searched using a local BLAST engine. CONCLUSIONS: VKCDB is a resource for comparative studies of voltage-gated potassium channels. The methods used to construct VKCDB are general; they can be used to create specialized databases for other protein families. VKCDB is accessible at http://vkcdb.biology.ualberta.ca.

Animals↗

MIPS Arabidopsis thaliana Database (MAtDB): an integrated biological knowledge resource based on the first complete plant genome.

Arabidopsis thaliana is the first plant for which the complete genome has been sequenced and published. Annotation of complex eukaryotic genomes requires more than the assignment of genetic elements to the sequence. Besides completing the list of genes, we need to discover their cellular roles, their regulation and their interactions in order to understand the workings of the whole plant. The MIPS Arabidopsis thaliana Database (MAtDB; http://mips.gsf.de/proj/thal/db) started out as a repository for genome sequence data in the European Scientists Sequencing Arabidopsis (ESSA) project and the Arabidopsis Genome Initiative. Our aim is to transform MAtDB into an integrated biological knowledge resource by integrating diverse data, tools, query and visualization capabilities and by creating a comprehensive resource for Arabidopsis as a reference model for other species, including crop plants.

Arabidopsis↗

Grouping and identification of sequence tags (GRIST): bioinformatics tools for the NEIBank database.

NEIBank is a project to develop and organize genomics and bioinformatics resources for the eye. As part of this effort, tools have been developed for bioinformatics analysis and web based display of data from expressed sequence tag (EST) analyses. EST sequences are identified and formed into groups or clusters representing related transcripts from the same gene. This is carried out by a rules-based procedure called GRIST (GRouping and Identification of Sequence Tags) that uses sequence match parameters derived from BLAST programs. Linked procedures are used to eliminate non-mRNA contaminants. All data are assembled in a relational database and assembled for display as web pages with annotations and links to other informatics resources. Genome projects generate huge amounts of data that need to be classified and organized to become easily accessible to the research community. GRIST provides a useful tool for assembling and displaying the results of EST analyses. The NEIBank web site contains a growing set of pages cataloging the known transcriptional repertoire of eye tissues, derived from new NEIBank cDNA libraries and from eye-related data deposited in the dbEST section of GenBank.

Animals↗

Attentional load and implicit sequence learning.

A widely employed conceptualization of implicit learning hypothesizes that it makes minimal demands on attentional resources. This conjecture was investigated by comparing learning under single-task and dual-task conditions in the sequential reaction time (SRT) task. Participants learned probabilistic sequences, with dual-task participants additionally having to perform a counting task using stimuli that were targets in the SRT display. Both groups were then tested for sequence knowledge under single-task (Experiments 1 and 2) or dual-task (Experiment 3) conditions. Participants also completed a free generation task (Experiments 2 and 3) under inclusion or exclusion conditions to determine if sequence knowledge was conscious or unconscious in terms of its access to intentional control. The experiments revealed that the secondary task impaired sequence learning and that sequence knowledge was consciously accessible. These findings disconfirm both the notion that implicit learning is able to proceed normally under conditions of divided attention, and that the acquired knowledge is inaccessible to consciousness. A unitary framework for conceptualizing implicit and explicit learning is proposed.

Adolescent↗

Modeling and analysis of protein design under resource constraints.

The potency, or fitness, of a protein-based drug can be enhanced by changing the sequence of its underlying protein. We present a novel stochastic model for the sequence-fitness relation, and estimate its four parameters from industrial data. Using this model, we formulate and analyze two variants of the protein design problem. In the single-period design problem, the designer needs to decide under capacity constraints which set of sequences to screen in order to maximize the expected fitness of the best sequence in the set. In the more general two-period design problem, the designer can afford two screening rounds and needs to allocate resources optimally across the two periods to maximize the same objective function. Analytical and simulation results allow us to assess the utility of the proposed design strategies for various parameter regimes.

Animals↗

Clustering and analysis of protein families.

Various sequence-motif and sequence-cluster databases have been integrated into a new resource known as InterPro. Because the contributing databases have different clustering principles and scoring sensitivities, the combined assignments complement each other for grouping protein families and delineating domains. InterPro and new developments in the analysis of both the phylogenetic profiles of protein families and domain fusion events improve the prediction of specific functions for numerous proteins.

Amino Acid Motifs↗

The homeodomain resource: a prototype database for a large protein family.

The Homeodomain Resource is an annotated collection of non-redundant protein sequences, three-dimensional structures and genomic information for the homeodomain protein family. Release 2.0 contains 765 full-length homeodomain-containing sequences, 29 experimentally derived structures and 116 homeobox loci implicated in human genetic disorders. Entries are fully hyperlinked to facilitate easy retrieval of the original records from source databases. A simple search engine with a graphical user interface is provided to query the component databases and assemble customized data sets. A new feature for this release is the addition of more automated methods for database searching, maintenance and implementation of efficient data management. The Homeodomain Resource is freely available through the WWW at http://genome.nhgri.nih.gov/homeodomain

Amino Acid Sequence↗

Integrative approaches to determining Csl function.

While there is an ever-increasing amount of information regarding cellulose synthase catalytic subunits (CesA) and their role in the formation of the cell wall, the remainder of the enzymes that synthesize structural cell wall polysaccharides are unknown. The completion of the Arabidopsis genome and the wealth of the sequence information from other plant genome projects provide a rich resource for determining the identity of these enzymes. Arabidopsis contains six families of genes related to cellulose synthase, the cellulose synthase-like (Csl) genes. Our laboratory is taking a multidisciplinary approach to determine the function of the Csl genes, incorporating genomic, genetic and biochemical data. Information from expressed sequence tag (EST) projects has revealed the presence of Csl genes in all plant species with a significant number of ESTs. Certain Csl families appear to be missing from some species. For example, no examples of CslG ESTs have been found in rice or maize. Microarray data and reporter constructs are being used to determine the expression pattern of the CesA and Csl genes in Arabidopsis. Mutations and insertion events have been identified in a majority of the genes in the Arabidopsis CesA superfamily and are being characterized by phenotypic and biochemical analysis. While we cannot yet link the function of any of the Csl genes to their respective products, the expression and localization of these genes is consistent with the expected expression pattern of polysaccharide synthases that contribute to the primary cell wall.

Arabidopsis↗

Genetic structure and diversity in Oryza sativa L.

The population structure of domesticated species is influenced by the natural history of the populations of predomesticated ancestors, as well as by the breeding system and complexity of the breeding practices exercised by humans. Within Oryza sativa, there is an ancient and well-established divergence between the two major subspecies, indica and japonica, but finer levels of genetic structure are suggested by the breeding history. In this study, a sample of 234 accessions of rice was genotyped at 169 nuclear SSRs and two chloroplast loci. The data were analyzed to resolve the genetic structure and to interpret the evolutionary relationships between groups. Five distinct groups were detected, corresponding to indica, aus, aromatic, temperate japonica, and tropical japonica rices. Nuclear and chloroplast data support a closer evolutionary relationship between the indica and the aus and among the tropical japonica, temperate japonica, and aromatic groups. Group differences can be explained through contrasting demographic histories. With the availability of rice genome sequence, coupled with a large collection of publicly available genetic resources, it is of interest to develop a population-based framework for the molecular analysis of diversity in O. sativa.

Base Sequence↗

Breed classification of Lao People's Democratic Republic (Lao PDR) and Thai native chickens using synchrotron radiation-based Fourier transform infrared spectroscopy and genotyping by sequencing.

Lao PDR harbors substantial genetic diversity in native chicken populations, representing an important resource for sustainable production and long-term food security. This study aimed to classify five Lao native chicken breeds-Ou, Black Bone, Horn Chou, Yolk, and Chae-and to discriminate them from a Thai native breed, Leung Hang Khao (LK), using integrative genotype-based approaches. Blood samples were collected from 50 LK and Lao native chickens (32 Ou, 10 Black Bone, 9 Horn Chou, 121 Yolk, and 41 Chae). Genomic DNA was extracted and analyzed using synchrotron radiation-based Fourier-transform infrared (SR-FTIR) spectroscopy to characterize biochemical composition, while genotyping-by-sequencing (GBS) was employed to identify genome-wide single nucleotide polymorphisms (SNPs). SR-FTIR analysis revealed highly significant differences among breeds in nucleotide-associated functional groups, including thymine, adenine, guanine, cytosine, as well as DNA backbone and deoxyribose components (P < 0.001). Multivariate analyses demonstrated that principal component analysis (PCA) of SR-FTIR spectra effectively discriminated chicken breeds, while hierarchical cluster analysis (HCA) further resolved them into two major clusters with distinct sub-clusters, reflecting variation in DNA biochemical composition. In contrast, GBS analysis identified 1484 common SNPs; however, PCA based on SNP data showed limited resolution in clearly separating breeds, despite revealing similar clustering trends. Overall, the results highlight the strong discriminatory power of SR-FTIR spectroscopy for rapid and effective classification of native chicken breeds at the molecular level, outperforming SNP-based differentiation under the current marker density. This study provides novel insights into the application of synchrotron-based spectroscopic techniques in poultry genetics and contributes valuable baseline information for the conservation and utilization of Lao native chicken genetic resources.

Breed classification↗

Differential extraction and protein sequencing reveals major differences in patterns of primary cell wall proteins from plants.

The proteins of the primary cell walls of suspension cultured cells of five plant species, Arabidopsis, carrot, French bean, tomato, and tobacco, have been compared. The approach that has been adopted is differential extraction followed by SDS-polyacrylamide gel electrophoresis (PAGE), rather than two-dimensional gel analysis, to facilitate protein sequencing. Whole cells were washed sequentially with the following aqueous solutions, CaCl2, CDTA (cyclohexane diaminotetraacetic acid, DTT (dithiothreitol), NaCl, and borate. SDS-PAGE analysis showed consistent differences between species. From the 233 proteins that were selected for sequencing, 63% gave N-terminal data. This analysis shows that (i) patterns of proteins revealed by SDS-PAGE are strikingly different for all five species, (ii) a large number of these proteins cannot be identified by data base searches indicating that a significant proportion of wall proteins have not been previously described, (iii) the major proteins that can be identified belong to very different classes of proteins, (iv) the majority of proteins found in the extracellular growth media are absent from their respective cell wall extracts, and (v) the results of the extraction process are indicative of higher order structure. It appears that aspects of speciation reside in the complement of extracellular wall proteins. The data represent a protein resource for cell wall studies complementary to EST (expressed sequence tag) and DNA sequencing strategies.

Amino Acid Sequence↗

REDfly: a Regulatory Element Database for Drosophila.

Bioinformatics studies of transcriptional regulation in the metazoa are significantly hindered by the absence of readily available data on large numbers of transcriptional cis-regulatory modules (CRMs). Even the richly annotated Drosophila melanogaster genome lacks extensive CRM information. We therefore present here a database of Drosophila CRMs curated from the literature complete with both DNA sequence and a searchable description of the gene expression pattern regulated by each CRM. This resource should greatly facilitate the development of computational approaches to CRM discovery as well as bioinformatics analyses of regulatory sequence properties and evolution.

Animals↗

Motor sequence complexity and performing hand produce differential patterns of hemispheric lateralization.

Studies in brain damaged patients conclude that the left hemisphere is dominant for controlling heterogeneous sequences performed by either hand, presumably due to the cognitive resources involved in planning complex sequential movements. To determine if this lateralized effect is due to asymmetries in primary sensorimotor or association cortex, whole-brain functional magnetic resonance imaging was used to measure differences in volume of activation while healthy right-handed subjects performed repetitive (simple) or heterogeneous (complex) finger sequences using the right or left hand. Advanced planning, as evidenced by reaction time to the first key press, was greater for the complex than simple sequences and for the left than right hand. In addition to the expected greater contralateral activation in the sensorimotor cortex (SMC), greater left hemisphere activation was observed for left, relative to right, hand movements in the ipsilateral left superior parietal area and for complex, relative to simple, sequences in the left premotor and parietal cortex, left thalamus, and bilateral cerebellum. No such volumetric asymmetries were observed in the SMC. Whereas the overall MR signal intensity was greater in the left than right SMC, the extent of this asymmetry did not vary with hand or complexity level. In contrast, signal intensity in the parietal and premotor cortex was greater in the left than right hemisphere and for the complex than simple sequences. Signal intensity in the caudal anterior cerebellum was greater bilaterally for the complex than simple sequences. These findings suggest that activity in the SMC is associated with execution requirements shared by the simple and complex sequences independent of their differential cognitive requirements. In contrast, consistent with data in brain damaged patients, the left dorsal premotor and parietal areas are engaged when advanced planning is required to perform complex motor sequences that require selection of different effectors and abstract organization of the sequence, regardless of the performing hand.

Adult↗

Human BAC ends quality assessment and sequence analyses.

End sequences from bacterial artificial chromosomes (BACs) provide highly specific sequence markers in large-scale sequencing projects. To date, we have generated >300,000 end sequences from >186,000 human BAC clones with an average read length of >460 bp for a total of 141 Mb covering approximately 4.7% of the genome. Over 60% of the clones have BAC end sequences (BESs) from both ends representing more than fivefold coverage of the human genome by the paired-end clones. Our quality assessments and sequence analyses indicate that BESs from human BAC libraries developed at The California Institute of Technology (CalTech) and Roswell Park Cancer Institute have similar properties. The analyses have highlighted differences in insert size for different segments of the CalTech library. Problems with the fidelity of tracking of sequence data back to physical clones have been observed in some subsets of the overall BES dataset. The annotation results of BESs for the contents of available genomic sequences, sequence tagged sites, expressed sequence tags, protein encoding regions, and repeats indicate that this resource will be valuable in many areas of genome research.

Chromosome Mapping↗

Ancestral polymorphisms in genetic markers obscure detection of evolutionarily distinct populations in the endangered Florida grasshopper sparrow (Ammodramus savannarum floridanus).

Genetic analyses of bird subspecies designated as conservation units can address whether they represent units with independent evolutionary histories and provide insights into the evolutionary processes that determine the degree to which they are genetically distinct. Here we use mitochondrial DNA control region sequence and six microsatellite DNA loci to examine phylogeographical structure and genetic differentiation among five North American grasshopper sparrow (Ammodramus savannarum) populations representing three subspecies, including a population of the endangered Florida subspecies (A. s. floridanus). This federally listed taxon is of particular interest because it differs phenotypically from other subspecies in plumage and behaviour and has also undergone a drastic decline in population size over the past century. Despite this designation, we observed no phylogeographical structure among populations in either marker: mtDNA haplotypes and microsatellite genotypes from floridanus samples did not form clades that were phylogenetically distinct from variants found in other subspecies. However, there was low but significant differentiation between Florida and all other populations combined in both mtDNA (FST = 0.069) and in one measure of microsatellite differentiation (theta = 0.016), while the non-Florida populations were not different from each other. Based on analyses of mtDNA variation using a coalescent-based model, the effective sizes of these populations are large (approximately 80,000 females) and they have only recently diverged from each other (< 26,000 ybp). These populations are probably far from genetic equilibrium and therefore the lack of phylogenetic distinctiveness of the floridanus subspecies and minimal genetic differentiation is due most probably to retained ancestral polymorphism. Finally, levels of variation in Florida were similar to other populations supporting the idea that the drastic reduction in population size which has occurred within the last 100 years has not yet had an impact on levels of variation in floridanus. We argue that despite the lack of phylogenetic distinctiveness of floridanus genotypes the observed genetic differentiation and previously documented phenotypic differences justify continued designation of this subspecies as a protected population segment.

Animals↗

IMGT, the international ImMunoGeneTics database.

The international ImMunoGeneTics database (IMGT) (http://imgt.cines.fr), is a high quality integrated information system specializing in Immunoglobulins (IG), T cell Receptors (TR) and Major Histocompatibility Complex (MHC) of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB) with different interfaces (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ('IMGT Marie-Paule page') and interactive tools for sequence analysis (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene). IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on IMGT-ONTOLOGY. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and thera-peutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

Plant MPSS databases: signature-based transcriptional resources for analyses of mRNA and small RNA.

MPSS (massively parallel signature sequencing) is a sequencing-based technology that uses a unique method to quantify gene expression level, generating millions of short sequence tags per library. We have created a series of databases for four species (Arabidopsis, rice, grape and Magnaporthe grisea, the rice blast fungus). Our MPSS databases measure the expression level of most genes under defined conditions and provide information about potentially novel transcripts (antisense transcripts, alternative splice isoforms and regulatory intergenic transcripts). A modified version of MPSS has been used to perform deep profiling of small RNAs from Arabidopsis, and we have recently adapted our database to display these data. Interpretation of the small RNA MPSS data is facilitated by the inclusion of extensive repeat data in our genome viewer. All the data and the tools introduced in this article are available at http://mpss.udel.edu.

Arabidopsis↗