Search PubMedSearch

SEARCH · Search PubMed

Results for “Programming Languages”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Realfreq: real-time base modification analysis for nanopore sequencing.

SUMMARY: Nanopore sequencers allow sequencing data to be accessed in real-time. This allows live analysis to be performed, while the sequencing is running, reducing the turnaround time of the results. We introduce realfreq, a framework for obtaining real-time base modification frequencies while a nanopore sequencer is in operation. Realfreq calculates and allows access to the real-time base modification frequency results while the sequencer is running. We demonstrate that the data analysis rate with realfreq on a laptop computer can keep up with the output data rate of a nanopore MinION sequencer, while a desktop computer can keep up with a single PromethION 2 solo flowcell. AVAILABILITY AND IMPLEMENTATION: Realfreq is a free and open-source application implemented in C programming language and shell scripts. The source code and the documentation for realfreq can be found at https://github.com/imsuneth/realfreq. The version used for the manuscript is also available at https://doi.org/10.5281/zenodo.15128668.

Nanopore Sequencing

A novel method for across-chromosome phasing without relative data.

MOTIVATION: Across-chromosome phasing identifies which haplotypes of different chromosomes come from the same parent. This differs from within-chromosome phasing, which uses linkage disequilibrium patterns to determine which alleles were co-inherited within each chromosome but does not match haplotypes across different chromosomes. While across-chromosome phasing can be conducted using genotypes from parents or close relatives, current methods perform poorly for samples of unrelated individuals. Here, we introduce a novel approach for across-chromosome phasing that employs a window-based SNP-similarity metric, eliminating the need for data from close relatives or detection of identical-by-descent haplotypes. RESULTS: Using UK Biobank offspring with both parents genotyped as a gold standard, we evaluated the performance of our method by phasing the offspring without using parental data. In genomic data with no within-chromosome phase errors, our algorithm achieved a mean across-chromosome phasing accuracy of 95%, with 53% of individuals phased perfectly. When data was pre-phased computationally using a standard within-chromosome phasing algorithm, mean accuracy for across-chromosome phasing dropped to 83.1%. Thus, our method is limited primarily by the accuracy of within-chromosome phasing accuracy and can approach near-perfect across-chromosome phasing accuracy as within-chromosome phasing accuracy improves. AVAILABILITY AND IMPLEMENTATION: The implementation was executed within a multi-node computational environment of University of Colorado Boulder Research Computing (Blanca Cluster: https://www.colorado.edu/rc/resources/blanca), employing parallelization techniques in the C programming language. The source code has been made publicly accessible online at https://github.com/emmanuelsapin/AcrossChromosomesPhasing, thereby facilitating reproducibility of the results for researchers with authorized access to the UK Biobank dataset.

Algorithms

An interpretable deep learning framework uncovers features governing CRISPR-Cas9 genome-editing efficiency.

MOTIVATION: CRISPR-Cas9 genome-editing efficiency is strongly influenced by the sequence composition and positional context of single-guide RNAs (sgRNAs). Although numerous deep learning-based models have been developed to predict Cas9 efficiency from sgRNA sequences, most operate as black boxes, offering limited insight into the sequence determinants underlying Cas9 activity. In addition, previous studies often overlook how the positional context of sequence motifs within sgRNAs influences their effects on Cas9 binding or cleavage. RESULTS: We introduce DeepCC9, an interpretable machine learning framework that combines explicit sequence feature extraction with a residual block-based deep architecture to improve interpretability and identify composition- and position-based motifs governing Cas9 genome-editing efficiency. We applied this method to multiple Cas9 variant datasets, achieving superior predictive performance compared with existing methods while enabling direct interpretation of sequence motifs and their positional effects. Our analysis uncovered 74 sequence motifs enriched or depleted at specific positions within sgRNAs and strongly associated with Cas9 efficiency, providing mechanistic insight into sequence features that influence guide performance. Together, these results establish DeepCC9 as a generalizable and interpretable framework for modeling sequence-function relationships and advancing the understanding of the sequence determinants underlying CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The authors have implemented their algorithm in the Python programming language (version 3.X), which is accessible using (https://zenodo.org/records/20073890).

Deep Learning

SNPannotator: automated functional annotation of genetic variants and linked proxies.

SUMMARY: Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation. AVAILABILITY AND IMPLEMENTATION: The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.

Software

Microcomputer programs for back translation of protein to DNA sequences and analysis of ambiguous DNA sequences.

Three computer programs are described which may be used to translate a DNA sequence into a protein sequence, back translate the protein sequence into an ambiguous DNA sequence, and then do pattern searching in the ambiguous sequence. The programs are written in the C programming language, have been compiled to run on a microcomputer under the CP/M 80 operating system, and may be copied in binary format through a modem. They are also to become available for the IBM/PC.

Amino Acid Sequence

Apple Macintosh programs for nucleic and protein sequence analyses.

This paper describes a package of programs for handling and analyzing nucleic acid and protein sequences using the Apple Macintosh microcomputer. There are three important features of these programs: first, because of the now classical Macintosh interface the programs can be easily used by persons with little or no computer experience. Second, it is possible to save all the data, written in an editable scrolling text window or drawn in a graphic window, as files that can be directly used either as word processing documents or as picture documents. Third, sequences can be easily exchanged with any other computer. The package is composed of thirteen programs, written in Pascal programming language.

Amino Acid Sequence

MIPS: a database for protein sequences, homology data and yeast genome information.

The MIPS group (Martinsried Institute for Protein Sequences) at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, collects, processes and distributes protein sequence data within the framework of the tripartite association of the PIR-International Protein Sequence Database (,). MIPS contributes nearly 50% of the data input to the PIR-International Protein Sequence Database. The database is distributed on CD-ROM together with PATCHX, an exhaustive supplement of unique, unverified protein sequences from external sources compiled by MIPS. Through its WWW server (http://www.mips.biochem.mpg.de/ ) MIPS permits internet access to sequence databases, homology data and to yeast genome information. (i) Sequence similarity results from the FASTA program () are stored in the FASTA database for all proteins from PIR-International and PATCHX. The database is dynamically maintained and permits instant access to FASTA results. (ii) Starting with FASTA database queries, proteins have been classified into families and superfamilies (PROT-FAM). (iii) The HPT (hashed position tree) data structure () developed at MIPS is a new approach for rapid sequence and pattern searching. (iv) MIPS provides access to the sequence and annotation of the complete yeast genome (), the functional classification of yeast genes (FunCat) and its graphical display, the 'Genome Browser' (). A CD-ROM based on the JAVA programming language providing dynamic interactive access to the yeast genome and the related protein sequences has been compiled and is available on request.

Academies and Institutes

Computer programs to assist in high resolution thermal denaturation and circular dichroism studies on nucleic acids.

Computer programs are described that direct the collection, processing, and graphical display of numerical data obtained from high resolution thermal denaturation (1-3) and circular dichroism (4) studies. Besides these specific applications, the programs may also be useful, either directly or as programming models, in other types of spectrophotometric studies employing computers, programming languages, or instruments similar to those described here (see Materials and Methods).

Circular Dichroism

Accurate detection of tandem repeats exposes ubiquitous reuse of biological sequences.

Tandem repetition is one of the major processes underlying genome evolution and phenotypic diversification. While newly formed tandem repeats are often easy to identify, it is more challenging to detect repeat copies as they diverge over evolutionary timescales. Existing programs for finding tandem repeats return markedly different results, and it is unclear which predictions are more correct and how much room remains for improvement. Here, we introduce DetectRepeats, a new method that uses empirical information about structural repeats to improve the accuracy of repeat detection. We show that DetectRepeats advances the state-of-the-art by finding highly divergent repeats with relatively few false positive detections. We apply DetectRepeats to genomes across the tree of life to discover an enrichment of detectable tandem repeats within different genes, genome regions, and taxa. Furthermore, we use phylogenetic reconciliation to determine that some tandem repeats continue to evolve through intra-repeat unit replacement. In this manner, tandem repeats serve as a renewable genetic resource offering a bountiful source of alternative genetic material. Our work unlocks the confident detection of ancient tandem repeats, opening a doorway to future discoveries. DetectRepeats is part of the DECIPHER package for the R programming language and available via Bioconductor.

Tandem Repeat Sequences

Epidemiologic modeling using a microcomputer spreadsheet package.

Epidemiologic modeling has provided both researchers and students with a means of studying complex disease processes as well as making intervention recommendations to decision makers. To develop more than the most elementary model, however, it has become necessary to be well versed in a computer programming language. While this has deterred many modelers in the past, with microcomputers it is now possible to develop even complex models without significant investment of time spent in learning a computer language. In addition to being affordable, many microcomputers offer "canned" spreadsheet packages which are readily adapted for epidemiologic modeling. To demonstrate this, two models were developed and run using a microcomputer spreadsheet package: 1) the classic Reed-Frost model, and 2) a modified Reed-Frost model with two intermixing subpopulations.

Animals

An object-oriented database for protein structure analysis.

An object-oriented database system has been developed which is being used to store protein structure data. The database can be queried using the logic programming language Prolog or the query language Daplex. Queries retrieve information by navigating through a network of objects which represent the primary, secondary and tertiary structures of proteins. Routines written in both Prolog and Daplex can integrate complex calculations with the retrieval of data from the database, and can also be stored in the database for sharing among users. Thus object-oriented databases are better suited to prototyping applications and answering complex queries about protein structure than relational databases. This system has been used to find loops of varying length and anchor positions when modelling homologous protein structures.

Amino Acid Sequence

Features of commercial computer software systems for medical examiners and coroners.

There are many ways of automating medical examiner and coroner offices, one of which is to purchase commercial software products specifically designed for death investigation. We surveyed four companies that offer such products and requested information regarding each company and its hardware, software, operating systems, peripheral devices, applications, networking options, programming language, querying capability, coding systems, prices, customer support, and number and size of offices using the product. Although the four products (CME2, ForenCIS, InQuest, and Medical Examiner's Software System) are similar in many respects and each can be installed on personal computers, there are differences among the products with regard to cost, applications, and the other features. Death investigators interested in office automation should explore these products to determine the usefulness of each in comparison with the others and in comparison with general-purpose, off-the-shelf databases and software adaptable to death investigation needs.

Computers

Numerical integration simulation programs for the microcomputer.

Programs for use with the Apple II Plus microcomputer that generate graphic simulations of various linear and Michaelis-Menten pharmacokinetic models are described. The programs numerically integrate sets of differential equations for appropriate pharmacokinetic models. Multiple oral (or intramuscular), intravenous bolus, or infusion doses (continuous or discontinuous) may be administered in any combination. Doses as well as pharmacokinetic parameters may be changed at the end of each simulated dosing interval. The programs can be easily modified by users familiar with the BASIC programming language and offer an economical approach to pharmacokinetic simulation.

Computers

The frequencies of HLA alleles and haplotypes and their distribution among donors and renal patients in the UNOS registry.

HLA allele and haplotype frequencies are used in transplantation, anthropology, forensic medicine, and studies of the associations between HLA factors and the immune response. The cost of determining these frequencies through family studies can be avoided by estimating them from population data. We have utilized the data in the UNOS donor registry and kidney transplant waiting list to estimate allele and haplotype frequencies for the HLA-A, -B, and -DR(B1) loci and report the allele and a portion of the haplotype data here. Using programs written in A Program Language (APL) we were able to perform all analyses on a personal computer. We have found that the distribution of haplotype frequencies varies among the races, with Caucasians having a greater number of both more common and extremely rare haplotypes. Despite the sizes of the groups studied, only one-third to two-thirds of the haplotypes theoretically possible were actually observed. Although the data confirm the well-known fact that the distributions of alleles and haplotypes varies among races, they also reveal that certain common haplotypes are shared among all racial groups and represent an opportunity for well-matched transplants between donors and recipients of different races.

Alleles

A Systematic Review of Spatial Epidemiological Modeling Approaches Applied During the COVID-19 Pandemic.

BACKGROUND: A wide range of epidemiological modeling approaches have been applied to the SARS-CoV-2 pandemic, which presents an opportunity to assess common approaches applied to specific research questions. Spatial models interrogate how heterogeneities and host movement dynamics influence local and regional patterns of disease, issues that were of great interest for understanding and controlling SARS-CoV-2. OBJECTIVE: Here we present a systematic review of spatial epidemiological modeling approaches of SARS-CoV-2. We describe common themes and highlight unique strategies, providing a foundation for researchers to devise spatial models most appropriate for future pathogens and epidemics. Our review also categorizes the research questions that were addressed with spatial models, highlights parameter estimation techniques, and describes the cyber infrastructure used for model development. METHODS: We conducted a systematic review using Web of Science and a standardized set of keywords, followed by thorough examination of abstracts and full texts to determine which studies met our inclusion criteria. To guide our description and comparisons of models, we developed a Geography, Population, Movement (GPM) framework that conceptualizes the interactions between three distinct subcomponents of any spatial model. The geographic model represents the physical arena in which the model is implemented, the intra-population model describes the transmission and disease processes that occur within distinct spatial units of the geography, and the movement model describes the algorithms that dictate how hosts move among spatial units within the geography. RESULTS: The search identified a total of 193 articles, of which 109 were included in our review. The most abundant intra-population modeling methods were agent-based (47.7%) and compartmental modeling (29.4%) approaches. Movement models ranged in complexity, with the most complex models implementing commuter movement among many points of interest in the geographic arena, which were sometimes parameterized by fine-scale mobility data. Geographic models ranged from describing microcosms, such as single classrooms, all the way up to multi-country models. Of the 63.3% of models studies that specified the programming language used, we detected ten different languages, with Matlab and Python being the most frequent, although only 30.6% of studies provided open-access code for their models. We also described eight specialized software systems that were used to construct agent-based or compartment models of COVID-19. CONCLUSIONS: Our review identified and characterized a variety of spatial modeling strategies and software that were usefully employed to address many relevant epidemiological questions for COVID-19. Future research is needed to quantitatively assess which modeling approaches are most appropriate in specific situations, to answer specific questions, or to apply to certain disease systems. Moreover, future cyberinfrastructure could help to modularize and standardize modeling approaches, which would increase transparency and reproducibility, and which would facilitate a detailed examination of which model attributes relate to model performance in a variety of contexts.

COVID-19

Scriptable access to the Caenorhabditis elegans genome sequence and other ACEDB databases.

Much of the world's genomic data are available to the community through networked databases that are accessed via Web interfaces. Although this paradigm provides browse-level access and has greatly facilitated linking between databases, it does not provide any convenient mechanism for programmatically fetching and integrating data from diverse databases. We have created a library and an application programming interface (API) named AcePerl that provides simple, direct access to ACEDB databases from the Perl programming language. With this library, programmers and computer-savvy biologists can write software to pose complex queries on local and remote ACEDB databases, retrieve the data, integrate the results, and move data objects from one database to another. In addition, a set of Web scripts running on top of AcePerl provides Web-based browsing of any local or remote ACEDB database. AcePerl and the AceBrowser Web browser run on Unix systems and are available under a license that allows for unrestricted use and redistribution. Both packages can be downloaded from URL. A Microsoft Windows port of AcePerl is in the planning stages.

Animals

Fuzzy control of mean arterial pressure in postsurgical patients with sodium nitroprusside infusion.

We developed a fuzzy control system to provide closed-loop control of mean arterial pressure (MAP) in postsurgical patients in a cardiac surgical intensive care unit setting by regulating sodium nitroprusside (SNP) infusion. The fuzzy controller, originally expert-system-based, was analytically converted to ten nonfuzzy control algorithms, which reduced execution time dramatically. The core of the control algorithms was a nonlinear proportional-integral (PI) controller whose proportional gain and integral gain adjusted continuously according to error and rate change of error of the process output. The gains became larger when process output was far from desired setpoint and smaller when process output was close to desired setpoint, resulting in more dynamic and stable control performance than the regular PI controller, especially when a linear process with time-delay or a nonlinear process was involved. The control algorithms, encoded in C programming language, were implemented to control MAP in patients. Preliminary clinical results showed that the average percentage of time in which MAP stayed between 90% and 110% of the MAP setpoint was 89.31%, with a standard deviation of 4.96%. These were calculated based on 12 patient trials, with total trial time of 95 and 13 min.

Algorithms

Use of commercial 'authoring systems' for medical education.

A recent development in computer-assisted medical instruction has been the introduction of 'authoring systems'. Authoring systems are computer programs which can allow an instructor to prepare computer-based medical instructional materials without the need to know programming languages or have more than minimal familiarity with the computer hardware. This report documents the use of a commercially available authoring system that was used to prepare a tutorial for medical student instruction. This lesson presented information about paediatric developmental disabilities in both a text and question-and-answer format. Significant improvement in knowledge was demonstrated by the pre- and post-test results of the study group compared to the control group. The control group consisted of students who did not view the tutorial but had been assigned to a paediatric developmental disabilities clinic. The medical students who viewed the tutorial generally had very favourable comments about the use of such a system for the presentation of new information.

Attitude of Health Personnel