Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “biological data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Selective integration of multiple biological data for supervised network inference.

MOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request.

Algorithms↗

A case study of high-throughput biological data processing on parallel platforms.

MOTIVATION: Analysis of large biological data sets using a variety of parallel processor computer architectures is a common task in bioinformatics. The efficiency of the analysis can be significantly improved by properly handling redundancy present in these data combined with taking advantage of the unique features of these compute architectures. RESULTS: We describe a generalized approach to this analysis, but present specific results using the program CEPAR, an efficient implementation of the Combinatorial Extension algorithm in a massively parallel (PAR) mode for finding pairwise protein structure similarities and aligning protein structures from the Protein Data Bank. CEPAR design and implementation are described and results provided for the efficiency of the algorithm when run on a large number of processors. AVAILABILITY: Source code is available by contacting one of the authors.

Algorithms↗

A bayesian framework for parentage analysis: the value of genetic and other biological data.

We develop fractional allocation models and confidence statistics for parentage analysis in mating systems. The models can be used, for example, to estimate the paternities of candidate males when the genetic mother is known or to calculate the parentage of candidate parent pairs when neither is known. The models do not require two implicit assumptions made by previous models, assumptions that are potentially erroneous. First, we provide formulas to calculate the expected parentage, as opposed to using a maximum likelihood algorithm to calculate the most likely parentage. The expected parentage is superior as it does not assume a symmetrical probability distribution of parentage and therefore, unlike the most likely parentage, will be unbiased. Second, we provide a mathematical framework for incorporating additional biological data to estimate the prior probability distribution of parentage. This additional biological data might include behavioral observations during mating or morphological measurements known to correlate with parentage. The value of multiple sources of information is increased accuracy of the estimates. We show that when the prior probability of parentage is known, and the expected parentage is calculated, fractional allocation provides unbiased estimates of the variance in reproductive success, thereby correcting a problem that has previously plagued parentage analyses. We also develop formulas to calculate the confidence interval in the parentage estimates, thus enabling the assessment of precision. These confidence statistics have not previously been available for fractional models. We demonstrate our models with several biological examples based on data from two fish species that we study, coho salmon (Oncorhychus kisutch) and bluegill sunfish (Lepomis macrochirus). In coho, multiple males compete to fertilize a single female's eggs. We show how behavioral observations taken during spawning can be combined with genetic data to provide an accurate calculation of each male's paternity. In bluegill, multiple males and multiple females may mate in a single nest. For a nest, we calculate the fertilization success and the 95% confidence interval of each candidate parent pair.

Animals↗

PC/VAX or standalone PC-based general purpose biological data collection system.

A system, to collect, analyse and display biological data, is developed using IBM PC AT compatibles (PCs) or CED1401/1609 devices networked to a VAX environment. It can be operated in three separate modes: using the CED/IEEE/VAX network; using the PC/Ethernet/VAX network; as a standalone PC. The original system comprised CED 1401/1609 data collection devices running on the IEEE bus. This has been superseded by a PC-286 or better incorporating an analogue-to-digital convertor (DT2824-PGH) and a communications interface board (DEPCA) linked by thinwire Ethernet (ThinWire) running DECnet with their product network application software PCSA. This network has not only doubled the original throughput but has also removed the two major IEEE constraints: 4 m between devices and the physical linking of devices to the VAX. The PCs are logically linked to clustered VAXes on Ethernet, giving flexible networking supporting multiple ThinWire segments, each supporting a maximum of 30 PCs per 185 m segment length. As the enhanced design compliments the original, both may operate concurrently, appear similar in operation to the user and use the same analysis software, all of which help reduce the rate of system obsolescence.

Computer Communication Networks↗

On transforming biological data to Gaussian form.

Much of the statistical analysis of biological data depends on the assumption that the data are Gaussian (or normal). Some well-known procedures which use this assumption are (i) t-tests (ii) analysis of variance (iii) regression estimation and their attendant tests. If the data are not Gaussian, one can use nonparametric statistical techniques, if they exist, but they often require larger amounts of data to obtain equally precise results (see for example Lumsden and Mullen (7) for a discussion of this with regard to reference value estimation). If the data are not Gaussian a fruitful approach to their analyses lies in trying to find a transformation which will render tham Gaussian. The data thus transformed to a Gaussian form, can be analyzed validly using standard statistical techniques. The process of finding a good transformation of the data has often been an arbitrary and ad hoc one. The purpose of this article is to look at a particular technique for attempting to render nonGaussian data Gaussian, and to illustrate its applicability and breadth of use.

Animals↗

Genetic network analysis in light of massively parallel biological data acquisition.

Complementary DNA microarray and high density oligonucleotide arrays opened the opportunity for massively parallel biological data acquisition. Application of these technologies will shift the emphasis in biological research from primary data generation to complex quantitative data analysis. Reverse engineering of time-dependent gene-expression matrices is amongst the first complex tools to be developed. The success of reverse engineering will depend on the quantitative features of the genetic networks and the quality of information we can obtain from biological systems. This paper reviews how the (1) stochastic nature, (2) the effective size, and (3) the compartmentalization of genetic networks as well as (4) the information content of gene expression matrices will influence our ability to perform successful reverse engineering.

Computational Biology↗

Applying hybrid reasoning to mine for associative features in biological data.

We develop the means to mine for associative features in biological data. The hybrid reasoning schema for deterministic machine learning and its implementation via logic programming is presented. The methodology of mining for correlation between features is illustrated by the prediction tasks for protein secondary structure and phylogenetic profiles. The suggested methodology leads to a clearer approach to hierarchical classification of proteins and a novel way to represent evolutionary relationships. Comparative analysis of Jasmine and other statistical and deterministic systems (including Explanation-Based Learning and Inductive Logic Programming) are outlined. Advantages of using deterministic versus statistical data mining approaches for high-level exploration of correlation structure are analyzed.

Algorithms↗

An object oriented user interface for analysis of biological data.

In a previous paper we described a self-documented file and a collection of general purpose programs or tools that facilitates the management and analysis of biological data. The tools can be specified in a pipeline to accomplish a specific analysis task. However, we found that it was difficult for investigators to learn the UNIX command language for specifying pipelines, specify selection tasks through a command language, and visualize the data as they were transformed and rearranged. To alleviate these problems we developed an object-oriented user interface for the pipeline programs. The system consists of four major programs for visualization: Vedit, Vgraf, Vscan, and V spread. Vedit is a simple text editor, Vgraf is a flexible graphics program, Vscan facilitates scanning graphically through large files, and Vspread provides spreadsheet-like capabilities. To demonstrate how the visualization programs are used together to accomplish the needed analysis we describe two case studies and then discuss how well the system accomplished the goals of visualization, short learning curve, and user adaptability.

Biology↗

DDBJ in the stream of various biological data.

In the past year we at DDBJ (http://www.ddbj.nig. ac.jp) have made a steady increase in the number of data submissions with a 50.6% increment in the number of bases or 46.5% increment in the number of entries. Among them the genome data of man, ascidian and rice hold the top three. Our activity has extended to providing a tool that enables sequence retrieval using regular expressions, and to launching our SOAP server and web services to facilitate the acquisition of proper data and tools from a huge number of biological data resources on websites worldwide. We have also opened our public gene expression database, CIBEX.

Animals↗

Seven-channel digital telemetry system for monitoring and direct computer capturing of biological data.

A seven-channel telemetry system for collection and display of biological data is presented. The system can amplify bioelectrical signals in the range of 2 microV to 200 mV and has a bandwidth of 0.1-80 Hz. After multiplexing, the signals are digitized with a resolution of 8 bits. The data are frequency modulated directly on a VHF transmitter. After receiving the data on a VHF receiver, they are routed directly to the RS232 input connector on the PC. Thereby the advantage of direct communication between the transmitter and the PC can be utilized. Expensive analog equipment is avoided and display of the signals on the PC screen as well as signal analysis can be performed. The system has been tested and was found to be stable and highly reliable.

Amplifiers, Electronic↗

Mining biological data using self-organizing map.

This paper presents a novel method of mining biological data using a self-organizing map (SOM). After partitioning a set of protein sequences using SOM, conventional homology alignment is applied to each cluster to determine the conserved local motif (biological pattern) for the cluster. These local motifs are then regarded as rules for prediction and classification. In the application to the prediction of HIV protease cleavage sites in proteins, we found that the rules derived from this method are much more robust than those derived from the decision tree method.

Algorithms↗

Recombinant hirudin (HBW 023): biological data of ten patients with severe venous thrombo-embolism.

This study reports on the biological data of ten patients with acute venous thrombo-embolism. They were treated for 5 days with continuous intravenous infusion of a fixed dose (0.05 mg/kg/hr) of a recombinant hirudin (r-H HBW 023 Behringwerke, Germany). The plasma level of r-H (HBW 023), assessed by an anti-factor IIa amidolytic activity, was stable after Day 2 and showed considerable individual variations. It correlated with APTT ratio, suggesting that this test is a reliable tool to monitor therapy. In contrast, thrombin time was constantly over 120 sec (control 15 sec) and consequently was not a useful parameter. Prothrombin time showed a slight, but significant, prolongation, which was correlated with the increase of APTT ratio. There was no bleeding time prolongation, platelet count, or ATIII level decrease. Levels of thrombin-antithrombin III complexes, and D-dimers, which were high in all patients on admission, decreased during the course of the treatment but remained abnormal on Day 5, showing an ongoing hemostasis and fibrinolysis activation: this is consistent with the delayed, but only slightly decreased thrombin generation evidenced by thrombin generation test performed on Day 3. These results suggest that thrombin inhibition by rH-hirudin at this dosage is only partial, which allows the generation of traces of thrombin needed for the feed-back thrombin production generated by factor V and VIII activation.

Adult↗

OPSEG: a general routine for smoothing and interpolating discrete biological data.

The optimal segments technique is a new approach to smoothing and interpolating between small numbers of discrete biological data. This method balances the degree of smoothness against the expected error of the observed data. The OPSEG computer program searches for the set of smoothed data points which will match the overall difference between the smoothed and observed data to an a priori estimate of the measurement error. The smoothed curve is described as a series of linked individual line segments. This approach is useful for the analysis of biological signals such as plasma measurements of hormone and metabolite concentration and has been applied to the development of assay standard curves.

Biometry↗

A computer program for linear nonparametric and parametric identification of biological data.

A computer program package for parametric ad nonparametric linear system identification of both static and dynamic biological data, written for an LSI-11 minicomputer with 28 K of memory, is described. The program has 11 possible commands including an instructional help command. A user can perform nonparametric spectral analysis and estimation of autocorrelation and partial autocorrelation functions of univariate data and estimate nonparametrically the transfer function and possibly an associated noise series of bivariate data. In addition, the commands provide the user the means to derive a parametric autoregressive moving average model for univariate data, to derive a parametric transfer function and noise model for bivariate data, and to perform several model evaluation tests such as pole-zero cancellation, examination of residual whiteness and uncorrelatedness with the input. The program, consisting of a main program and driver subroutine as well as six overlay segments, may be run interactively or automatically.

Computers↗