Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Library Automation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Automated DNA chip annotation tables at IFOM: the importance of synchronisation and cross-referencing of sequence databases.

The increasing popularity of DNA chip technology for the study of gene expression is producing, for each experiment, a sizable quantity of numerical data to analyse and an accompanying large number of gene identifiers that should be associated with the relevant biological annotation. We describe here a website at IFOM (FIRC Institute of Molecular Oncology) where we release regularly updated annotation tables for the most used Affymetrix oligonucleotide DNA chips and for the whole Research Genetics 46K clone collection for cDNA arrays. These tables are synchronised with every new release of the mouse and human UniGene databases (NCBI; National Center for Biotechnology Information), allowing fast and easy preliminary annotation of DNA array experiments. We also report some comparative evidence about the importance of biological database synchronisation and cross-references in the process of generating annotation tables for DNA chips.

Abstracting and Indexing↗

A comprehensive library of DNA-binding site matrices for 55 proteins applied to the complete Escherichia coli K-12 genome.

A major mode of gene regulation occurs via the binding of specific proteins to specific DNA sequences. The availability of complete bacterial genome sequences offers an unprecedented opportunity to describe networks of such interactions by correlating existing experimental data with computational predictions. Of the 240 candidate Escherichia coli DNA-binding proteins, about 55 have DNA-binding sites identified by DNA footprinting. We used these sites to construct recognition matrices, which we used to search for additional binding sites in the E. coli genomic sequence. Many of these matrices show a strong preference for non-coding DNA. Discrepancies are identified between matrices derived from natural sites and those derived from SELEX (Systematic Evolution of Ligands by Exponential enrichment) experiments. We have constructed a database of these proteins and binding sites, called DPInteract (available at http://arep.med.harvard.edu/dpinteract).

Bacterial Proteins↗

The covalent structure of cartilage collagen. Amino acid sequence of residues 552-661 of bovine alpha1(II) chains.

The covalent structure of the first 111 residues from the N-terminus of peptide alpha1(II)-CB10 from bovine nasal-cartilage collagen is presented. This region comprises residues 552-661 of the alpha1(II) chain. The sequence was determined by automated Edman degradation of peptide alpha1(II)-CB10 and of peptides produced by cleavage with trypsin and hydroxylamine. Comparison of this region of the alpha1(II) chain with the homologous segment of the alpha1(I) chain indicated a homology level of 85%, slightly higher than that of 81% reported for the N-terminal region of the alpha1(II) chain (Butler, Miller & Finch (1976) Biochemistry15, 3000-3006). The occurrence of two residues of glycosylated hydroxylysine was established at positions 564 and 603, the first present exclusively as galactosylhydroxylysine and the latter as a mixture of galactosylhydroxylysine and glucosylgalactosylhydroxylysine. Also, two residues at positions 648 and 657 were tentatively identified as glycosylated hydroxylysines. The amino acid sequences adjacent to the hydroxylysine residues so far identified in the alpha1(II) chain were compared with the homologous regions of the alpha1(I) and alpha2 chains, but no obvious prerequisite for hydroxylation could be seen. From comparison with the homologous sequence of the alpha1(I) chain, it appears that the alpha1(II)-chain sequence presented here contains three more amino acids than that reported for the alpha1(I) chain. This triplet would be interposed between residues 63 and 64 of the reported sequence of peptide alpha1(I)-CB7 from calf skin collagen. Data on the purification of the subpeptides and their amino acid compositions have been deposited as Supplementary Publication SUP 50087 (7 pages) at the British Library Lending Division, Boston Spa, Wetherby, West Yorkshire LS23 7BQ, U.K., from whom copies can be obtained on the terms indicated in Biochem. J. (1978) 169, 5.

Amino Acid Sequence↗

An efficient and generic extension to ITK to process arbitrary shaped regions of interest.

The paper describes a software method to extend ITK (Insight ToolKit, supported by the National Library of Medicine), leading to ITK++. This method, which is based on the extension of the iterator design pattern, allows the processing of regions of interest with arbitrary shapes, without modifying the existing ITK code. We experimentally evaluate this work by considering the practical case of the liver vessel segmentation from CT-scan images, where it is pertinent to constrain processings to the liver area. Experimental results clearly prove the interest of this work: for instance, the anisotropic filtering of this area is performed in only 16 s with our proposed solution, while it takes 52 s using the native ITK framework. A major advantage of this method is that only add-ons are performed: this facilitates the further evaluation of ITK++ while preserving the native ITK framework.

Algorithms↗

Primary structure of a constituent polypeptide chain (AIII) of the giant haemoglobin from the deep-sea tube worm Lamellibrachia. A possible H2S-binding site.

The deep-sea tube worm Lamellibrachia, belonging to the Phylum Vestimentifera, contains two giant extracellular haemoglobins, a 3000 kDa haemoglobin and a 440 kDa haemoglobin. The former consists of four haem-containing chains (AI-AIV) and two linker chains (AV and AVI) for the assembly of the haem-containing chains [Suzuki, Takagi & Ohta (1988) Biochem. J. 255, 541-545]. The tube-worm haemoglobins are believed to have a function of transporting sulphide (H2S) to internal bacterial symbionts, as well as of facilitating O2 transport [Arp & Childress (1983) Science 219, 295-297]. We have determined the complete amino acid sequence of Lamellibrachia chain AIII by automated or manual Edman sequencing. The chain is composed of 144 amino acid residues, has three cysteine residues at positions 3, 74 and 133, and has a molecular mass of 16,620 Da, including a haem group. The sequence showed significant homology (30-50% identity) with those of haem-containing chains of annelid giant haemoglobins. Two of the three cysteine residues are located at the positions where an intrachain disulphide bridge is formed in all annelid chains, but the remaining one (Cys-74) was located at a unique position, compared with annelid chains. Since the chain AIII was shown to have a reactive thiol group in the intact 3000 kDa molecule by preliminary experiments, the cysteine residue at position 74 appears to be one of the most probable candidates for the sulphide-binding sites. A phylogenetic tree was constructed from nine chains of annelid giant haemoglobins and one chain of vestimentiferan tube-worm haemoglobin now determined. The tree clearly showed that Lamellibrachia chain AIII belongs to the family of strain A of annelid giant haemoglobins, and that the two classes of Annelida, polychaete and oligochaete, and the vestimentiferan tube worm diverged at almost the same time. H.p.l.c. patterns of peptides (Figs. 4-7), amino acid compositions of peptides (Table 2) and amino acid sequences of intact protein and peptides (Table 3) have been deposited as Supplementary Publication SUP 50154 (13 pages) at the British Library Document Supply Centre, Boston Spa, Wetherby, West Yorkshire LS23 7BQ, U.K., from whom copies can be obtained on the terms indicated in Biochem. J. (1990) 265, 5.

Amino Acid Sequence↗

Semi-automated high throughput combinatorial solid-phase organic synthesis.

A semi-automated technique for massive parallel solid-phase organic synthesis based on a "split only" strategy is described. Two different types of purpose-oriented reaction vessels are used. The initial steps are performed in domino blocks, and the resin-bound intermediates then split into wells of a micro plate for the last combinatorial step. The domino block is a reaction block for manual and semi-automatic parallel solid-phase organic synthesis that simplifies liquid exchange and integrates common synthetic steps. The synthesis in micro plates does not use any filter for separation of resin beads from the supernatant liquid, and allows high throughput parallel synthesis on solid phase to be performed. This technique, documented on examples of diverse disubstituted benzenes, includes the use of gaseous cleavage in the last synthetic step and allows the synthesis of thousands of compounds per day in mg quantities.

Amino Acids↗

Activity classification using realistic data from wearable sensors.

Automatic classification of everyday activities can be used for promotion of health-enhancing physical activities and a healthier lifestyle. In this paper, methods used for classification of everyday activities like walking, running, and cycling are described. The aim of the study was to find out how to recognize activities, which sensors are useful and what kind of signal processing and classification is required. A large and realistic data library of sensor data was collected. Sixteen test persons took part in the data collection, resulting in approximately 31 h of annotated, 35-channel data recorded in an everyday environment. The test persons carried a set of wearable sensors while performing several activities during the 2-h measurement session. Classification results of three classifiers are shown: custom decision tree, automatically generated decision tree, and artificial neural network. The classification accuracies using leave-one-subject-out cross validation range from 58 to 97% for custom decision tree classifier, from 56 to 97% for automatically generated decision tree, and from 22 to 96% for artificial neural network. Total classification accuracy is 82 % for custom decision tree classifier, 86% for automatically generated decision tree, and 82% for artificial neural network.

Activities of Daily Living↗

Strategies for whole microbial genome sequencing and analysis.

The introduction of methods for automated DNA sequence analysis nearly a decade ago, together with more recent advances in the field of bioinformatics, have revolutionized biology and medicine and have ushered in a new era of genomic science, the study of genes and genomes. These new technologies have had an impact on many areas of research, including the association between genes and disease, in DNA-based diagnostics, and in the sequencing of genomes from human and other model organisms. The demonstration in 1995, that automated DNA sequencing methods could be used to decipher the entire genome sequence of a free-living organism, Haemophilus influenzae, was a milestone in both the genomics and microbial fields [1]. Since the first report of the complete sequence of H. influenzae, these methodologies have been adopted by laboratories around the world. The complete genomic sequence of five eubacterial species [1-5], one archaea [6], and the eukaryote, Saccharomyces cerevisiae [7], have been reported in the last 18 months. At the beginning of 1997 more than a dozen microbial genome projects are at or near completion, with many others in progress. It is likely that in the next few years we will see the complete sequence of perhaps as many as 30-40 microbial genomes. In this article, we will review methods for whole genome sequencing and analysis and examine how this information can be exploited to better understand microbial physiology and evolution.

Base Sequence↗

Registration of MR and CT images of the liver: comparison of voxel similarity and surface based registration algorithms.

The purpose of this work was to determine the feasibility and efficacy of retrospective registration of MR and CT images of the liver. The open-source ITK Insight Software package developed by the National Library of Medicine (USA) contains a multi-resolution, voxel-similarity-based registration algorithm which we selected as our baseline registration method. For comparison we implemented a multi-scale surface fitting technique based on the head-and-hat algorithm. Registration accuracy was assessed using the mean displacement of automatically selected point landmarks. The ITK voxel-similarity-based registration algorithm performed better than the surface-based approach with mean misregistration in the range of 7.7-8.4 mm for CT-CT registration, 8.2 mm for MR-MR registration, and 14.0-18.9 mm for MR-CT registration compared to mean misregistration from the surface-based technique in the range of 9.6-11.1 mm for CT-CT registration, 9.2-12.4 mm for MR-MR registration, and 15.2-19.0 mm for MR-CT registration.

Algorithms↗

Analysis of badger urine volatiles using gas chromatography-mass spectrometry and pattern recognition techniques.

The potential for badger urine to signal olfactory information relating to sex, age class and seasonality was investigated by performing GCMS headspace analysis followed by pattern recognition statistical analysis on 84 urine samples collected from different categories of animal. Approximately 300 compounds were identified using library searching, and GCMS peak areas were recorded for the 33 most common. PCA was performed on the normalised and standardised data from all badgers, through which significant seasonal trends and groupings of homologous series of compounds were detected. PCA was also performed on the three subgroups of adults in the spring, summer and autumn, and a level of sexual discrimination was possible during the latter two seasons. Malanobis distances on the scores of the first five principal components provided good discrimination for these three subgroups, but discrimination was poor when all samples were analysed together. This, combined with the initial results of the PCA, confirms that a strong seasonal trend is imposed upon the sexual trend in this dataset. Our initial analysis indicates that badger urine potentially contains olfactory cues relating to sex and season. The relevance of these findings to understanding olfactory communication in mammals is discussed.

Animals↗

Cloning and sequence analysis of the cDNA for arachidonate 12-lipoxygenase of porcine leukocytes.

The complete amino acid sequence of arachidonate 12-lipoxygenase (EC 1.13.11.31) of porcine leukocytes was deduced by cloning and sequence analysis of DNA complementary to its mRNA. The sequence was confirmed by automated Edman degradation of the N-terminal regions of the native enzyme and its proteolytic fragments. The cDNA had an open reading frame encoding 662 amino acid residues with a calculated molecular weight of 74,911. Amino acid residues 533-545, Cys-(Xaa)3-Cys-(Xaa)3-His-(Xaa)3-His, showed significant homology to the short cysteine- or histidine-containing sequences proposed as the metal-binding domains of transcription factors and various metal-containing proteins [Berg, J. M. (1986) Science 232, 485-487]. The amino acid sequence of 12-lipoxygenase exhibited 86% identity with human reticulocyte 15-lipoxygenase and showed 41% identity with human leukocyte 5-lipoxygenase. The 12-lipoxygenase cDNA recognized a 3.4-kilobase mRNA species in various porcine cell types, with the largest amount in leukocytes, followed by pituitary, lung, jejunum, and spleen.

Amino Acid Sequence↗

A learned comparative expression measure for affymetrix genechip DNA microarrays.

Perhaps the most common question that a microarray study can ask is, "Between two given biological conditions, which genes exhibit changed expression levels?" Existing methods for answering this question either generate a comparative measure based upon a static model, or take an indirect approach, first estimating absolute expression levels and then comparing the estimated levels to one another. We present a method for detecting changes in gene expression between two samples based on data from Affymetrix GeneChips. Using a library of over 200,000 known cases of differential expression, we create a learned comparative expression measure (LCEM) based on classification of probe-level data patterns as changed or unchanged. LCEM uses perfect match probe data only; mismatch probe values did not prove to be useful in this context. LCEM is particularly powerful in the case of small microarry studies, in which a regression-based method such as RMA cannot generalize, and in detecting small expression changes. At the levels of selectivity that are typical in microarray analysis, the LCEM shows a lower false discovery rate than either MAS5 or RMA trained from a single chip. When many chips are available to RMA, LCEM performs better on two out of the three data sets, and nearly as well on the third. Performance of the MAS5 log ratio statistic was notably bad on all datasets.

Algorithms↗

Modeling side-chain conformation for homologous proteins using an energy-based rotamer search.

We have developed a computational method for accurately predicting the conformation of side-chain atoms when building a protein structure from a known homologous structure. A library of rotamers is used to model the side-chains, allowing an average of five to six different conformations per residue. Local sites of adjacent side-chains are defined throughout the protein, and all combinations of side-chain rotamers are evaluated within each site using a molecular mechanics force field enhanced by the inclusion of a solvation term. At each site, the lowest energy combination of side-chains is identified and added onto the fixed protein backbone. A series of test cases using the refined X-ray structure of alpha-lytic protease has shown that: (1) the force field can correctly predict up to 90% of side-chain rotamers; (2) the assumption of side-chain rotamer geometry is usually a very good approximation; and (3) the complete combinatorial conformation search is able overcome local minima and identify the lowest energy rotamer set for the protein in the absence of a starting bias to the correct structure. Tests with several pairs of homologous proteins have shown that the algorithm is quite successful at predicting side-chain conformation even when the protein backbone used to generate side-chain positions deviates from the correct conformation. The root-mean-square (r.m.s.) deviation of predicted side-chain atoms rises from 1.31 A (average r.m.s.d. 0.73 A) in a test case with the correct backbone to only 2.68 A (1.95 A average r.m.s.d.) in a test case with < 35% homology. The high accuracy of this method suggests that it may be a useful automated tool for modeling protein structure.

Algorithms↗

High-throughput screening for potent and selective inhibitors of Plasmodium falciparum dihydroorotate dehydrogenase.

Plasmodium falciparum is the causative agent of the most serious and fatal malarial infections, and it has developed resistance to commonly employed chemotherapeutics. The de novo pyrimidine biosynthesis enzymes offer potential as targets for drug design, because, unlike the host, the parasite does not have pyrimidine salvage pathways. Dihydroorotate dehydrogenase (DHODH) is a flavin-dependent mitochondrial enzyme that catalyzes the fourth reaction in this essential pathway. Coenzyme Q (CoQ) is utilized as the oxidant. Potent and species-selective inhibitors of malarial DHODH were identified by high-throughput screening of a chemical library, which contained 220,000 drug-like molecules. These novel inhibitors represent a diverse range of chemical scaffolds, including a series of halogenated phenyl benzamide/naphthamides and urea-based compounds containing napthyl or quinolinyl substituents. Inhibitors in these classes with IC(50) values below 600 nm were purified by high pressure liquid chromatography, characterized by mass spectroscopy, and subjected to kinetic analysis against the parasite and human enzymes. The most active compound is a competitive inhibitor of CoQ with an IC(50) against malarial DHODH of 16 nm, and it is 12,500-fold less active against the human enzyme. Site-directed mutagenesis of residues in the CoQ-binding site significantly reduced inhibitor potency. The structural basis for the species selective enzyme inhibition is explained by the variable amino acid sequence in this binding site, making DHODH a particularly strong candidate for the development of new anti-malarial compounds.

Animals↗

Fold recognition by combining profile-profile alignment and support vector machine.

MOTIVATION: Currently, the most accurate fold-recognition method is to perform profile-profile alignments and estimate the statistical significances of those alignments by calculating Z-score or E-value. Although this scheme is reliable in recognizing relatively close homologs related at the family level, it has difficulty in finding the remote homologs that are related at the superfamily or fold level. RESULTS: In this paper, we present an alternative method to estimate the significance of the alignments. The alignment between a query protein and a template of length n in the fold library is transformed into a feature vector of length n + 1, which is then evaluated by support vector machine (SVM). The output from SVM is converted to a posterior probability that a query sequence is related to a template, given SVM output. Results show that a new method shows significantly better performance than PSI-BLAST and profile-profile alignment with Z-score scheme. While PSI-BLAST and Z-score scheme detect 16 and 20% of superfamily-related proteins, respectively, at 90% specificity, a new method detects 46% of these proteins, resulting in more than 2-fold increase in sensitivity. More significantly, at the fold level, a new method can detect 14% of remotely related proteins at 90% specificity, a remarkable result considering the fact that the other methods can detect almost none at the same level of specificity.

Algorithms↗

Model-based assessment of cardiovascular health from noninvasive measurements.

Cardiovascular health is currently assessed through a variety of hemodynamic parameters, many of which can only be determined by invasive measurement often requiring hospitalization. A noninvasive method of evaluating several of these parameters such as systemic vascular resistance (SVR), maximum left ventricular elasticity (E(LV)), end diastolic volume (VED), and cardiac output, is presented. The method has three elements: (1) a distributed model of the human cardiovascular system (Ozawa et aL, Ann. Biomed. Eng. 29:284-297, 2001) to generate a solution library that spans the anticipated range of parameter values, (2) a method for establishing the multidimensional relationship between features computed from the arterial blood pressure and/or flow traces (e.g., mean arterial pressure, pulse amplitude, mean flow velocity) and the critical hemodynamic parameters, and (3) a parameter estimation method that yields the best fit between measured and computed data. Sensitivity analyses were used to determine the critical parameters, and the influence of fixed model parameters. Using computer-generated brachial pressure and velocity profiles (which can be measured noninvasively), the error associated with this method was found to be less than 3% for SVR, and less than 10% for E(LV) and V(ED). Simulations were also performed to test the ability of the approach to predict changes in SVR and E(LV) from an initial base line state.

Algorithms↗

Functional annotation by identification of local surface similarities: a novel tool for structural genomics.

BACKGROUND: Protein function is often dependent on subsets of solvent-exposed residues that may exist in a similar three-dimensional configuration in non homologous proteins thus having different order and/or spacing in the sequence. Hence, functional annotation by means of sequence or fold similarity is not adequate for such cases. RESULTS: We describe a method for the function-related annotation of protein structures by means of the detection of local structural similarity with a library of annotated functional sites. An automatic procedure was used to annotate the function of local surface regions. Next, we employed a sequence-independent algorithm to compare exhaustively these functional patches with a larger collection of protein surface cavities. After tuning and validating the algorithm on a dataset of well annotated structures, we applied it to a list of protein structures that are classified as being of unknown function in the Protein Data Bank. By this strategy, we were able to provide functional clues to proteins that do not show any significant sequence or global structural similarity with proteins in the current databases. CONCLUSION: This method is able to spot structural similarities associated to function-related similarities, independently on sequence or fold resemblance, therefore is a valuable tool for the functional analysis of uncharacterized proteins. Results are available at http://cbm.bio.uniroma2.it/surface/structuralGenomics.html.

Algorithms↗

New softwares for automated microsatellite marker development.

Microsatellites are repeated small sequence motifs that are highly polymorphic and abundant in the genomes of eukaryotes. Often they are the molecular markers of choice. To aid the development of microsatellite markers we have developed a module that integrates a program for the detection of microsatellites (TROLL), with the sequence assembly and analysis software, the Staden Package. The module has easily adjustable parameters for microsatellite lengths and base pair quality control. Starting with large datasets of unassembled sequence data in the form of chromatograms and/or text data, it enables the creation of a compact database consisting of the processed and assembled microsatellite containing sequences. For the final phase of primer design, we developed a program that accepts the multi-sequence 'experiment file' format as input and produces a list of primer pairs for amplification of microsatellite markers. The program can take into account the quality values of consensus bases, improving success rate of primer pairs in PCR. The software is freely available and simple to install in both Windows and Unix-based operating systems. Here we demonstrate the software by developing primer pairs for 427 new candidate markers for peanut.

Arachis↗