Search PubMedSearch

SEARCH · Search PubMed

Results for “Programming Languages”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Parallel computation and FASTA: confronting the problem of parallel database search for a fast sequence comparison algorithm.

We have parallelized the FASTA algorithm for biological sequence comparison using Linda, a machine-independent parallel programming language. The resulting parallel program runs on a variety of different parallel machines. A straight-forward parallelization strategy works well if the amount of computation to be done is relatively large. When the amount of computation is reduced, however, disk I/O becomes a bottleneck which may prevent additional speed-up as the number of processors is increased. The paper describes the parallelization of FASTA, and uses FASTA to illustrate the I/O bottleneck problem that may arise when performing parallel database search with a fast sequence comparison algorithm. The paper also describes several program design strategies that can help with this problem. The paper discusses how this bottleneck is an example of a general problem that may occur when parallelizing, or otherwise speeding up, a time-consuming computation.

Algorithms

Harnessing networked workstations as a powerful parallel computer: a general paradigm illustrated using three programs for genetic linkage analysis.

It is widely accepted that parallel computers, which have the ability to execute different parts of a program simultaneously, will offer dramatic speed-up for many time-consuming biological computations. The paper describes how the use of the machine-independent parallel programming language, Linda, allows parallel programs to run on an institution's network of workstations. In this way, an institution can harness existing hardware, which is often either idle or vastly underutilized, as a powerful 'parallel machine' with supercomputing capability. The paper illustrates this very general paradigm by describing the use of Linda to parallelize three widely used programs for genetic linkage analysis, a mathematical technique used in gene mapping. The paper then discusses a number of technical, administrative and social issues that arise when creating such a computational resource.

Algorithms

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software

polars-bio-fast, scalable, and out-of-core operations on large genomic interval datasets.

MOTIVATION: Genomic studies very often rely on computationally intensive analyses of relationships between features, which are typically represented as intervals along a 1D coordinate system (such as positions on a chromosome). In this context, the Python programming language is extensively used for manipulating and analyzing data stored in a tabular form of rows and columns, called a DataFrame. Pandas is the most widely used Python DataFrame package and has been criticized for inefficiencies and scalability issues, which its modern alternative-Polars-aims to address with a native backend written in the Rust programming language. RESULTS: polars-bio is a Python library that enables fast, parallel and out-of-core operations on large genomic interval datasets. Its main components are implemented in Rust, using the Apache DataFusion query engine and Apache Arrow for efficient data representation. It is compatible with Polars and Pandas DataFrame formats. In a real-world comparison (107 versus 1.2×106 intervals), our library runs overlap queries 6.5×, nearest queries 15.5×, count_overlaps queries 38×, and coverage queries 15× faster than Bioframe. On equally sized synthetic sets (107 versus 107), the corresponding speedups are 1.6×, 5.5×, 6×, and 6×. In streaming mode, on real and synthetic interval pairs, our implementation uses 90× and 15× less memory for overlap, 4.5× and 6.5× less for nearest, 60× and 12× less for count_overlaps, and 34× and 7× less for coverage than Bioframe. Multi-threaded benchmarks show good scalability characteristics. To the best of our knowledge, polars-bio is the most efficient single-node library for genomic interval DataFrames in Python. AVAILABILITY AND IMPLEMENTATION: polars-bio is an open-source Python package distributed under the Apache License available for major platforms, including Linux, macOS, and Windows in the PyPI registry. The online documentation is https://biodatageeks.org/polars-bio/ and the source code is available on GitHub: https://github.com/biodatageeks/polars-bio and Zenodo: https://doi.org/10.5281/zenodo.16374290. are available at Bioinformatics online.

Software

Implementation of a movement paradigm using the Commodore 64 microcomputer.

Implementation of an alternating movement paradigm for monkeys was achieved using an inexpensive but versatile microcomputer, the Commodore 64. During task performance, the computer monitors one of three user selectable input signals (e.g. joint position) and continuously displays this signal as a moving cursor on a video monitor along with a user positioned target box. Other user defined parameters include in-target holding time and reinforcement ratio. The system also provides two sound cues to signal entry of the cursor into the target box and successful completion of a trial. Extensive use is made of the computer's intrinsic hardware features for implementation of movement paradigm functions. Use of external components is limited to digitizing and interface hardware. A two part software package consisting of a BASIC and a machine language program performs all task and hardware related functions. Acquisition and display of analog input signals, display of target positions, and delivery of auditory cues and applesauce rewards are all controlled by the machine language program. All user defined parameters are specified from the BASIC menu program. The specific programs described in this paper should be applicable to the control of tasks requiring alternation of a behavioral parameter between two target zones.

Animals

[Use of computers in the roentgenological diagnosis of gynecologic diseases].

The authors analyze the use of computers in x-ray examinations of gynecologic patients. X-ray signs of infertility, endometritis, tuberculosis, myoma, malignant tumors, etc. were formalized. A total of 131 patients were examined and a council of physicians for these cases was computer-simulated. Variants of computer-processed x-ray diagnoses are presented, their informativeness indexes ranging from 0 to 100%. Programmed processing may be realized via SM-4, SM-1420, IZOT-1016C, Electronika 100-25 computers. The FORTRAN program language was employed to make up the programs.

Diagnosis, Computer-Assisted

GENPRO: automatic generation of Prolog clause files for knowledge-based systems in the biomedical sciences.

With the increasing interest in using knowledge-based approaches for protein structure prediction and modelling, there is a requirement for general techniques to convert molecular biological data into structures that can be interpreted by artificial intelligence programming languages (e.g. Prolog). We describe here an interactive program that generates files in Prolog clausal form from the most commonly distributed protein structural data collections. The program is flexible and enables a variety of clause structures to be defined by the user through a general schema definition system. Our method can be extended to include other types of molecular biological database or those containing non-structural information, thus providing a uniform framework for handling the increasing volume of data available to knowledge-based systems in biomedicine.

Database Management Systems

Community-based computerized donor record systems.

Computer programs for management of donor information have been developed for the Champaign County Blood Bank, a division of the Regional Health Resource Center, Urbana, Illinois. The system provides the blood bank with reports from the donor files, incorporating the donor's last donation dates, ABO groups and Rh factors, memberships in assurance programs, and rare donor information to generate lists to meet either daily or emergency inventory needs. The system was designed to decrease time requirements for donor recruitment, improve donor base sampling, aid in support of special recruitment programs, and provide statistical profiles of the community donor base. Experiences in the development and use of the system indicate requirements for effective development of such systems include careful design, strict monitoring of performance, use of a versatile programming language, and incorporation of program modifications via staff-programmer interaction throughout implementation.

Blood Donors