Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data reproducibility”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Enzyme family classification by support vector machines.

One approach for facilitating protein function prediction is to classify proteins into functional families. Recent studies on the classification of G-protein coupled receptors and other proteins suggest that a statistical learning method, Support vector machines (SVM), may be potentially useful for protein classification into functional families. In this work, SVM is applied and tested on the classification of enzymes into functional families defined by the Enzyme Nomenclature Committee of IUBMB. SVM classification system for each family is trained from representative enzymes of that family and seed proteins of Pfam curated protein families. The classification accuracy for enzymes from 46 families and for non-enzymes is in the range of 50.0% to 95.7% and 79.0% to 100% respectively. The corresponding Matthews correlation coefficient is in the range of 54.1% to 96.1%. Moreover, 80.3% of the 8,291 correctly classified enzymes are uniquely classified into a specific enzyme family by using a scoring function, indicating that SVM may have certain level of unique prediction capability. Testing results also suggest that SVM in some cases is capable of classification of distantly related enzymes and homologous enzymes of different functions. Effort is being made to use a more comprehensive set of enzymes as training sets and to incorporate multi-class SVM classification systems to further enhance the unique prediction accuracy. Our results suggest the potential of SVM for enzyme family classification and for facilitating protein function prediction. Our software is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.

Amino Acid Sequence↗

A thin chip microsprayer system coupled to Fourier transform ion cyclotron resonance mass spectrometry for glycopeptide screening.

A thin polymer microchip was coupled with a Fourier transform ion cyclotron resonance (FTICR) 9.4 T mass spectrometer and the method was optimized in negative ion mode for glycopeptide screening. The interface between the polymer microchip and FTICR mass spectrometer consists of an in-laboratory conceived and designed mounting system that exhibits robust and controllable alignment of the chip toward the inlet of the mass spectrometer. The particular attribute of the polymer chip coupled to the FTICR mass spectrometer, to achieve an increase in ionization efficiency and sensitivity under the premise of high mass accuracy of detection, is highlighted by the large number of major and minor glycopeptide structures detected and identified in highly heterogeneous mixtures obtained from urine matrices. Glycoforms expressing various saccharide chain lengths ranging from tri- to dodecasaccharide, bearing up to three sialic acid moieties, could be detected and assigned based on the accuracy of the mass measurement (average mass deviation below 6 ppm) of their molecular ions. -Thin chipESI-FTICRMS is a potent novel system for glycomic screening of complex mixtures, as demonstrated for identification of singly sialylated O-glycosylated amino acids and peptides from urine matrices, and could be considered for general applicability in the glycoanalytical field.

Amino Acid Sequence↗

Inferences on the within-subject coefficient of variation.

The within-subject coefficient of variation (WCV) is widely used as a measure of precision and reproducibility of data in medical and biological science. In this paper, generalized confidence intervals and tests for a single WCV is developed using the concept of generalized pivots under the assumption of one-way random effect model. This approach is further extended to two-sample cases. The resulting procedures are easy to compute and have good properties in term of coverage probabilities and type-I error control at small sample sizes. The proposed methods are illustrated by a real life example.

Analysis of Variance↗

Sequencing the yeast genome: an international achievement.

The yeast genome is currently being sequenced by a Consortium of European laboratories, in collaboration with a wider international network of researchers. It is expected that within the next two years Saccharomyces cerevisiae will become the first eukaryotic organism to have been completely genetically mapped and sequenced. This article traces the sequencing enterprise from its beginnings, outlining the intentions, the organisation, and the achievements so far. The tasks which remain are discussed, emphasising the follow-on research into the evolution of primitive karyotypes, and, more particularly, into the nature of novel genes revealed during sequencing. The functional analysis of novel genes is attracting an ever wider community of yeast scientists, so that research which began with a decision to sequence a simple genome promises to remain a focus for international cooperation.

Chromosome Mapping↗

A structural basis for sequence comparisons. An evaluation of scoring methodologies.

A residue-exchange matrix has been derived that is suitable for comparison of amino acid sequences. This matrix is based on the tabulation of 207,795 amino acid replacements observed in 65 homologous sets of structurally aligned three-dimensional structures (235 proteins). The majority of the data is from structural comparisons where there is between 15 and 40% sequence identity. As a result, a scoring matrix such as the one devised here should provide a sensitive basis for the comparison of amino acid sequences and the search for homologous sequences in amino acid databases. In order to assess the value of this matrix we have made a comparative analysis with 12 other published scoring matrices that have been used for the alignment of protein amino acid sequences. We find that the matrix derived here is among the better performers in terms of alignment significance, detection of homologous sequences and the accuracy of alignments.

Amino Acid Sequence↗

Statistical alignment: computational properties, homology testing and goodness-of-fit.

The model of insertions and deletions in biological sequences, first formulated by Thorne, Kishino, and Felsenstein in 1991 (the TKF91 model), provides a basis for performing alignment within a statistical framework. Here we investigate this model.Firstly, we show how to accelerate the statistical alignment algorithms several orders of magnitude. The main innovations are to confine likelihood calculations to a band close to the similarity based alignment, to get good initial guesses of the evolutionary parameters and to apply an efficient numerical optimisation algorithm for finding the maximum likelihood estimate. In addition, the recursions originally presented by Thorne, Kishino and Felsenstein can be simplified. Two proteins, about 1500 amino acids long, can be analysed with this method in less than five seconds on a fast desktop computer, which makes this method practical for actual data analysis.Secondly, we propose a new homology test based on this model, where homology means that an ancestor to a sequence pair can be found finitely far back in time. This test has statistical advantages relative to the traditional shuffle test for proteins.Finally, we describe a goodness-of-fit test, that allows testing the proposed insertion-deletion (indel) process inherent to this model and find that real sequences (here globins) probably experience indels longer than one, contrary to what is assumed by the model.

Algorithms↗

Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.

Global surveys of genomes measure the usage of essential molecular parts, defined here as protein families, superfamilies or folds, in different organisms. Based on surveys of the first 20 completely sequenced genomes, we observe that the occurrence of these parts follows a power-law distribution. That is, the number of distinct parts (F) with a given genomic occurrence (V) decays as F=aV(-b), with a few parts occurring many times and most occurring infrequently. For a given organism, the distributions of families, superfamilies and folds are nearly identical, and this is reflected in the size of the decay exponent b. Moreover, the exponent varies between different organisms, with those of smaller genomes displaying a steeper decay (i.e. larger b). Clearly, the power law indicates a preference to duplicate genes that encode for molecular parts which are already common. Here, we present a minimal, but biologically meaningful model that accurately describes the observed power law. Although the model performs equally well for all three protein classes, we focus on the occurrence of folds in preference to families and superfamilies. This is because folds are comparatively insensitive to the effects of point mutations that can cause a family member to diverge beyond detectable similarity. In the model, genomes evolve through two basic operations: (i) duplication of existing genes; (ii) net flow of new genes. The flow term is closely related to the exponent b and can accommodate considerable gene loss; however, we demonstrate that the observed data is reproduced best with a net inflow, i.e. with more gene gain than loss. Moreover, we show that prokaryotes have much higher rates of gene acquisition than eukaryotes, probably reflecting lateral transfer. A further natural outcome from our model is an estimation of the fold composition of the initial genome, which potentially relates to the common ancestor for modern organisms. Supplementary material pertaining to this work is available from www.partslist.org/powerlaw.

Animals↗

A study of solvent polarity and hydrogen bonding effects on the nitrogen NMR shielding of isomeric tetrazoles and ab initio calculation of the nitrogen shielding of azole systems

High-precision nitrogen NMR shieldings, bulk susceptibility corrected, are reported for the N-methyl derivatives of the two existing isomeric tetrazoles (I, II) in a variety of solvents which represent a wide range of solvent properties from the point of view of polarity as well as hydrogen bond donor and acceptor strength. The observed range of solvent-induced nitrogen shielding variations of I and II is significant for the pyrrole-type nitrogens (N-Me), up to 9 ppm, and even more so for pyridine-type nitrogen atoms, where it can attain a value of 20 ppm. There is a clear distinction between the two types of nitrogen atoms in that the former exhibit a deshielding effect with increasing polarity of the medium while the latter experience an increase in the magnetic shielding of their nuclei. The latter effect is significantly augmented by solvent-to-solute hydrogen-bond formation where the pyridine-type nitrogens are involved directly. It is also quite diversified throughout the pyridine-type nitrogen atoms and seems to constitute a measure of relative basicity with respect to hydrogen-bond formation of the nitrogens concerned. This basicity seems to parallel that with respect to a full transfer of a proton, as can be reckoned from ab initio calculations of the relevant protonation energies reported in the present study. The experimental data for the tetrazoles in cyclohexane solutions are combined with those obtained in our earlier extensive studies on azole, diazole, and triazole ring systems, for a comparison with ab initio calculations of the nitrogen shieldings concerned. The latter were carried out using the coupled Hartree-Fock/GIAO/6-31++G** approach and geometry optimizations employing the same basis set; they show a good linear correlation with the experimental data and reproduce not only major changes but also most of the subtle variations in the experimental nitrogen shieldings of the azole systems as a whole. Copyright 1998 Academic Press.

Journal Article↗

A mathematical model for germinal centre kinetics and affinity maturation.

We present a mathematical model which reproduces experimental data on the germinal centre (GC) kinetics of the primed primary immune response and on affinity maturation observed during the reaction. We show that antigen masking by antibodies which are produced by emerging plasma cells can drive affinity maturation and provide a feedback mechanism by which the reaction is stable against variations in the initial antigen amount over several orders of magnitude. This provides a possible answer to the long-standing question of the role of antigen reduction in driving affinity maturation. By comparing model predictions with experimental results, we propose that the selection probability of centrocytes and the recycling probability of selected centrocytes are not constant but vary during the GC reaction with respect to time. It is shown that the efficiency of affinity maturation is highest if clones with an affinity for the antigen well above the average affinity in the GC leave the GC for either the memory or plasma cell pool. It is further shown that termination of somatic hypermutation several days before the end of the germinal centre reaction is beneficial for affinity maturation. The impact on affinity maturation of simultaneous initiation of memory cell formation and somatic hypermutation vs. delayed initiation of memory cell formation is discussed.

Antibody Affinity↗

Experimental models for the study of cardiovascular function and disease.

In the study of cardiovascular biology, both under conditions of health and disease, the investigator enjoys the availability of a vast range of experimental models ranging from man to a single molecule and beyond. There is also a vast spectrum of measurable indices of function and injury. This is particularly so in the case of myocardial ischemia, a disease which still contributes to the majority of deaths in the Western Hemisphere. Each experimental model, each species and each end-point has its own inherent advantages and disadvantages and appreciating these will help the investigator select the most appropriate study system for the particular question under investigation. This article endeavours to identify some of these strengths and weaknesses and reveals the frequently encountered paradox that the greater the amount and reproducibility of data the further removed is the model from clinical reality. Fortunately, however, an appreciation of this 'weakness' can often be exploited for the advancement of knowledge.

Animals↗

Is primary chemotherapy useful for all patients with primary invasive breast cancer?

Chemotherapy dose intensification in breast tumours is being evaluated in many multicentre trials, its indication being based on a clinical response in high-risk patients, thus selecting for tumours with rapid proliferation and low resistance. However, results from randomized trials are still pending. Clinical and pathological responses to therapy are valuable surrogate endpoints following primary chemotherapy. They will make it possible to distinguish at an early stage between patients who still retain an apoptotic response to chemotherapy and those patients whose disease will progress rapidly due to resistance mechanisms. For practical purposes, patients at risk and capable of responding represent the population of choice for primary systemic chemotherapy. Thus, by investigating mechanisms of response and resistance during the first courses of treatment we may target chemotherapy at those patients likely to benefit most from this treatment. A number of immunotherapy and vaccination trials are being conducted in many different centres. There is a lot of anecdotal evidence that cancer vaccines could help patients, but little yet in the way of solid, reproducible clinical data. Best responses to clinical testing would ideally be expected in early-stage disease because there is less tumour bulk and the patient's immune system is still able to respond. Patients with early breast cancer who are at high risk of recurrence and who have failed to respond to primary chemotherapy might be given the option of participating in adjuvant vaccination trials following the completion of local therapy.

Antineoplastic Agents↗

Neuropharmacological properties of electrophysiologically identified, visually responsive neurones of the posterior lateral suprasylvian area. A microiontophoretic study.

Extracellular recordings have been made from 118 electrophysiologically identified neurones lying in the posterior lateral suprasylvian area (PLLS and PLMS) of cats anaesthetized with Nembutal. Eighty-one cells were activated synaptically by the electrical stimulation of cortical and subcortical sites known to be the sources of monosynaptic projections to the lateral suprasylvian area; latencies to such activations have been measured. The locations and sizes of the receptive fields of 55 neurones were determined. The direction sensitivity and ocularity of these cells also were examined. The effects of various pharmacological agonists and antagonists have been observed on visual responsiveness and synaptic excitability. The excitatory effects of subcortical (dorsal lateral geniculate nucleus and pulvinar nuclear complex) electrical stimulation on the activity of suprasylvian neurones were reduced substantially by the iontophoretic administration of atropine. Antagonists of the receptors for the excitatory amino acids reduced the effectiveness, on the single cell evoked activity, of stimulation of the ipsilateral 17/18 border region and contralateral homotopic lateral suprasylvian area. Both classes of antagonist reduced the magnitude of neuronal responses to photic stimulation, and these response attenuations were additive when the antagonists were ejected concurrently. All of the pharmacological effects were reversible and reproducible. These data lend support to the proposition that acetylcholine and an excitatory amino acid are mediators of synaptic transmission of cortical visual processes in the lateral suprasylvian area.

Acetylcholine↗

A flow cytometric bromodeoxyuridine/DNA analysis for cytokinetics of the pancreatic cancer cell line Capan-2 after irradiation.

The cytokinetics of the human pancreatic adenocarcinoma cell line Capan-2 were analyzed by a bromodeoxyuridine (BrdUrd)/DNA analysis with flow cytometry (FCM). The reproducibility of data was shown to be superior to that achieved using a DNA analysis program with FCM. The perturbation of cytokinetics following irradiation was examined. Bivariate distribution was able to distinguish each cell phase cohort, even following as strong an irradiation as 10 Gy, a condition for which DNA analysis proved unusable. BrdUrd/DNA analysis showed the G2M phase population after irradiation with 2 Gy to be significantly greater (P < 0.005) than that of the untreated cells. This difference was not detectable by the DNA analysis program due to the large standard deviation (SD) of the data. BrdUrd/DNA analysis has a sufficient quantitative accuracy that strongly perturbed cytokinetics in patients treated with various therapeutic agents, for example, with combined therapies, can thus be analyzed by this approach.

Adenocarcinoma↗

The consistency of cardiac output measurement (CO2 rebreathe) in children during exercise.

Exercise cardiac output (Q) was determined using the CO2 rebreathing equilibrium method. Five repeat tests in 12 boys and two tests over a 4 month interval in 47 boys were performed. Regression equations to predict Q from VO2 were in close agreement with dye dilution studies in boys (Eriksson and Koch 1973). Group mean data were reproducible from trial to trial. The day-to-day variability of Q, with a coefficient of variation of 7-8%, was found to be higher than when the CO2 method has been applied in adults. This greater variability was related, in part, to a larger biological variation in children as depicted in such relatively simple measures as submaximal exercise heart rate. The larger variability was also related to inaccuracies in the methods of PaCO2 estimation in children. Estimation from end-tidal CO2 concentrations requires further research to establish a correction for the alveolar-arterial gradient during exercise in children. Estimation of the child's dead space in exercise, with subsequent derivation of PaCO2 from the Bohr equation, also could be improved. Nevertheless, Q estimates in children exercising above VO2 1.01 X min-1 showed a day-to-day and long term stability acceptable for use in research and clinical studies.

Adolescent↗

Quantitation of microbial metabolism.

Quantitation is a characteristic property of natural sciences and technologies and is the background for all kinetic and dynamic studies of microbial life. This presentation concentrates therefore on materials and methods as tools necessary to accomplish a sound, quantitative and mechanistic understanding of metabolism. Mathematical models are the software, bioreactors, actuators and analytical equipment are the hardware used. Experiments must be designed and performed in accordance with the relaxation times of the biosystem investigated; some of the respective consequences are discussed and commented in detail. Special emphasis is given to the required density, accuracy and reproducibility of data as well as their validation.

Bacteria↗

Improved technique for investigation of cell metabolism by 31P NMR spectroscopy.

31P NMR studies on microorganisms have been carried out with the cells embedded in agarose gel. The novel use of the gel for the NMR studies has advantages over the usual liquid suspensions in terms of improved reproducibility of data and cell viability, with no net loss of spectral quality. Polyphosphate formation in Escherichia coli was monitored continuously for up to 24 h and metabolic changes in yeast for 6 h. Changes of the intracellular pH during glycolysis in yeast were determined from the chemical shift of the internal Pi. NMR titration curves of Pi in the presence of Mg2+ indicate uncertainties in internal pH values estimated by this technique.

Escherichia coli↗

Two simple methods for the evaluation of topically active anti-inflammatory steroidal ointments.

Simple laboratory methods for quantitating the topical anti-inflammatory activity of steroidal ointments are described. One is of croton oil ear edema in rats and the other is a new method using homologous passive cutaneous anaphylaxis (PCA) in rats. In order to avoid problems such as the animals' licking and/or rubbing the ointment at the applied sites, which might result in oral uptake, each rat was housed individually and fitted with a plastic collar in the croton oil experiment. The sites of ointment application in the PCA experiment were covered with adhesive plaster. Optimal experimental conditions were as follows. In the former method, ointments were applied to the inside surface of the ear 5 min after the irritant treatment and anti-edematous activity was determined after 6 h. In the latter, ointments were applied 3 h before the antigenic challenge to the dorsal area of animals which had been passively sensitized by anti-serum, and inhibition of the increased permeability was determined 45 min after the challenge. These methods were found to be reliable with respect to sensitivity and reproducibility of data. Ointments of halcinonide, betamethasone-17-valerate, hydrocortisone-17-butyrate, fluocinonide, flumethasone-21-pivalate and beclomethasone-17,21-dipropionate were evaluated by these methods.

Administration, Topical↗