Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article↗

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers↗

GeneLynx: a gene-centric portal to the human genome.

GeneLynx is a meta-database providing an extensive collection of hyperlinks to human gene-specific information in diverse databases available on the Internet. The GeneLynx project is based on the simple notion that given any gene-specific identifier (accession number, gene name, text, or sequence), scientists should be able to access a single location that provides a set of links to all the publicly available information pertinent to the specified human gene. GeneLynx was implemented as an extensible relational database with an intuitive and user-friendly Web interface. The data are automatically extracted from more than 40 external resources, using appropriate approaches to maximize coverage of the available data. Construction and curation of the system is mediated by a custom set of software tools. An indexing utility is provided to facilitate the establishment of hyperlinks in external databases. A unique feature of the GeneLynx system is a communal curation system for user-aided annotation. GeneLynx can be accessed freely at http://www.genelynx.org.

Database Management Systems↗

Computational comparison of human genomic sequence assemblies for a region of chromosome 4.

Much of the available human genomic sequence data exist in a fragmentary draft state following the completion of the initial high-volume sequencing performed by the International Human Genome Sequencing Consortium (IHGSC) and Celera Genomics (CG). We compared six draft genome assemblies over a region of chromosome 4p (D4S394-D4S403), two consecutive releases by the IHGSC at University of California, Santa Cruz (UCSC), two consecutive releases from the National Centre for Biotechnology Information (NCBI), the public release from CG, and a hybrid assembly we have produced using IHGSC and CG sequence data. This region presents particular problems for genomic sequence assembly algorithms as it contains a large tandem repeat and is sparsely covered by draft sequences. The six assemblies differed both in terms of their relative coverage of sequence data from the region and in their estimated rates of misassembly. The CG assembly method attained the lowest level of misassembly, whereas NCBI and UCSC assemblies had the highest levels of coverage. All assemblies examined included <60% of the publicly available sequence from the region. At least 6% of the sequence data within the CG assembly for the D4S394-D4S403 region was not present in publicly available sequence data. We also show that even in a problematic region, existing software tools can be used with high-quality mapping data to produce genomic sequence contigs with a low rate of rearrangements.

Chromosomes, Human, Pair 4↗

The automatic detection of homologous regions (ADHoRe) and its application to microcolinearity between Arabidopsis and rice.

It is expected that one of the merits of comparative genomics lies in the transfer of structural and functional information from one genome to another. This is based on the observation that, although the number of chromosomal rearrangements that occur in genomes is extensive, different species still exhibit a certain degree of conservation regarding gene content and gene order. It is in this respect that we have developed a new software tool for the Automatic Detection of Homologous Regions (ADHoRe). ADHoRe was primarily developed to find large regions of microcolinearity, taking into account different types of microrearrangements such as tandem duplications, gene loss and translocations, and inversions. Such rearrangements often complicate the detection of colinearity, in particular when comparing more anciently diverged species. Application of ADHoRe to the complete genome of Arabidopsis and a large collection of concatenated rice BACs yields more than 20 regions showing statistically significant microcolinearity between both plant species. These regions comprise from 4 up to 11 conserved homologous gene pairs. We predict the number of homologous regions and the extent of microcolinearity to increase significantly once better annotations of the rice genome become available.

Arabidopsis↗

Mapping expressed sequence tag sites on yeast artificial chromosome clones of Arabidopsis thaliana DNA.

We describe a method for efficient parallel mapping of expressed sequence tag (EST) sites onto yeast artificial chromosome (YAC) clones. The strategy involves an initial YAC clone pooling scheme that minimizes the number of required PCR amplifications. This is followed by parallel analysis of PCR amplicons of EST sequences. Using this method, we have screened 600 EST sites in combinatorial pools of 3449 YAC clones that contain Arabidopsis thaliana DNA inserts. The presence of these genes on YACs was detected by amplifying EST sequences with PCR and analyzing the reaction products by agarose gel electrophoresis. Of the 600 ESTs, 271 were found to map to individual YACs. Software tools are presented that allow for the automated analysis of this electrophoresis data. Suggestions for the scale-up of this method to map large genomes are discussed.

Arabidopsis↗

A liquid chromatography-mass spectrometry-based metabolome database for tomato.

For the description of the metabolome of an organism, the development of common metabolite databases is of utmost importance. Here we present the Metabolome Tomato Database (MoTo DB), a metabolite database dedicated to liquid chromatography-mass spectrometry (LC-MS)- based metabolomics of tomato fruit (Solanum lycopersicum). A reproducible analytical approach consisting of reversed-phase LC coupled to quadrupole time-of-flight MS and photodiode array detection (PDA) was developed for large-scale detection and identification of mainly semipolar metabolites in plants and for the incorporation of the tomato fruit metabolite data into the MoTo DB. Chromatograms were processed using software tools for mass signal extraction and alignment, and intensity-dependent accurate mass calculation. The detected masses were assigned by matching their accurate mass signals with tomato compounds reported in literature and complemented, as much as possible, by PDA and MS/MS information, as well as by using reference compounds. Several novel compounds not previously reported for tomato fruit were identified in this manner and added to the database. The MoTo DB is available at http://appliedbioinformatics.wur.nl and contains all information so far assembled using this LC-PDA-quadrupole time-of-flight MS platform, including retention times, calculated accurate masses, PDA spectra, MS/MS fragments, and literature references. Unbiased metabolic profiling and comparison of peel and flesh tissues from tomato fruits validated the applicability of the MoTo DB, revealing that all flavonoids and alpha-tomatine were specifically present in the peel, while several other alkaloids and some particular phenylpropanoids were mainly present in the flesh tissue.

Chromatography, Liquid↗

The lipopolysaccharide of Sinorhizobium meliloti suppresses defense-associated gene expression in cell cultures of the host plant Medicago truncatula.

In the establishment of symbiosis between Medicago truncatula and the nitrogen-fixing bacterium Sinorhizobium meliloti, the lipopolysaccharide (LPS) of the microsymbiont plays an important role as a signal molecule. It has been shown in cell cultures that the LPS is able to suppress an elicitor-induced oxidative burst. To investigate the effect of S. meliloti LPS on defense-associated gene expression, a microarray experiment was performed. For evaluation of the M. truncatula microarray datasets, the software tool MapMan, which was initially developed for the visualization of Arabidopsis (Arabidopsis thaliana) datasets, was adapted by assigning Medicago genes to the ontology originally created for Arabidopsis. This allowed functional visualization of gene expression of M. truncatula suspension-cultured cells treated with invertase as an elicitor. A gene expression pattern characteristic of a defense response was observed. Concomitant treatment of M. truncatula suspension-cultured cells with invertase and S. meliloti LPS leads to a lower level of induction of defense-associated genes compared to induction rates in cells treated with invertase alone. This suppression of defense-associated transcriptional rearrangement affects genes induced as well as repressed by elicitation and acts on transcripts connected to virtually all kinds of cellular processes. This indicates that LPS of the symbiont not only suppresses fast defense responses as the oxidative burst, but also exerts long-term influences, including transcriptional adjustment to pathogen attack. These data indicate a role for LPS during infection of the plant by its symbiotic partner.

Cells, Cultured↗

New developments in the Inorganic Crystal Structure Database (ICSD): accessibility in support of materials research and design.

The materials community in both science and industry use crystallographic data models on a daily basis to visualize, explain and predict the behavior of chemicals and materials. Access to reliable information on the structure of crystalline materials helps researchers concentrate experimental work in directions that optimize the discovery process. The Inorganic Crystal Structure Database (ICSD) is a comprehensive collection of more than 60,000 crystal structure entries for inorganic materials and is produced cooperatively by Fachinformationszentrum Karlsruhe (FIZ), Germany, and the US National Institute of Standards and Technology (NIST). The ICSD is disseminated in computerized formats with scientific software tools to exploit the content of the database. Features of a new Windows-based graphical user interface for the ICSD are outlined, together with directions for future development in support of materials research and design.

Journal Article↗

Decision support and disease management: a logic engineering approach.

This paper describes the development and application of PROforma, a unified technology for clinical decision support and disease management. Work leading to the implementation of PROforma has been carried out in a series of projects funded by European agencies over the past 13 years. The work has been based on logic engineering, a distinct design and development methodology that combines concepts from knowledge engineering, logic programming, and software engineering. Several of the projects have used the approach to demonstrate a wide range of applications in primary and specialist care and clinical research. Concurrent academic research projects have provided a sound theoretical basis for the safety-critical elements of the methodology. The principal technical results of the work are the PROforma logic language for defining clinical processes and an associated suite of software tools for delivering applications, such as decision support and disease management procedures. The language supports four standard objects (decisions, plans, actions, and enquiries), each of which has an intuitive meaning with well-understood logical semantics. The development toolset includes a powerful visual programming environment for composing applications from these standard components, for verifying consistency and completeness of the resulting specification and for delivering stand-alone or embeddable applications. Tools and applications that have resulted from the work are described and illustrated, with examples from specialist cancer care and primary care. The results of a number of evaluation activities are included to illustrate the utility of the technology.

Decision Support Systems, Clinical↗

Fast wavelet estimation of weak biosignals.

Wavelet-based signal processing has become commonplace in the signal processing community over the past decade and wavelet-based software tools and integrated circuits are now commercially available. One of the most important applications of wavelets is in removal of noise from signals, called denoising, accomplished by thresholding wavelet coefficients in order to separate signal from noise. Substantial work in this area was summarized by Donoho and colleagues at Stanford University, who developed a variety of algorithms for conventional denoising. However, conventional denoising fails for signals with low signal-to-noise ratio (SNR). Electrical signals acquired from the human body, called biosignals, commonly have below 0 dB SNR. Synchronous linear averaging of a large number of acquired data frames is universally used to increase the SNR of weak biosignals. A novel wavelet-based estimator is presented for fast estimation of such signals. The new estimation algorithm provides a faster rate of convergence to the underlying signal than linear averaging. The algorithm is implemented for processing of auditory brainstem response (ABR) and of auditory middle latency response (AMLR) signals. Experimental results with both simulated data and human subjects demonstrate that the novel wavelet estimator achieves superior performance to that of linear averaging.

Adult↗

A new insight into postsurgical objective voice quality evaluation: application to thyroplastic medialization.

This paper aims at providing new objective parameters and plots, easily understandable and usable by clinicians and logopaedicians, in order to assess voice quality recovering after vocal fold surgery. The proposed software tool performs presurgical and postsurgical comparison of main voice characteristics (fundamental frequency, noise, formants) by means of robust analysis tools, specifically devoted to deal with highly degraded speech signals as those under study. Specifically, we address the problem of quantifying voice quality, before and after medialization thyroplasty, for patients affected by glottis incompetence. Functional evaluation after thyroplastic medialization is commonly based on several approaches: videolaryngostroboscopy (VLS), for morphological aspects evaluation, GRBAS scale and Voice Handicap Index (VHI), relative to perceptive and subjective voice analysis respectively, and Multi-Dimensional Voice Program (MDVP), that provides objective acoustic parameters. While GRBAS has the drawback to entirely rely on perceptive evaluation of trained professionals, MDVP often fails in performing analysis of highly degraded signals, thus preventing from presurgical/postsurgical comparison in such cases. On the contrary, the new tool, being capable to deal with severely corrupted signals, always allows a complete objective analysis. The new parameters are compared to scores obtained with the GRBAS scale and to some MDVP parameters, suitably modified, showing good correlation with them. Hence, the new tool could successfully replace or integrate existing ones. With the proposed approach, deeper insight into voice recovering and its possible changes after surgery can thus be obtained and easily evaluated by the clinician.

Diagnosis, Computer-Assisted↗

Mixture model analysis of DNA microarray images.

In this paper, we propose a new methodology for analysis of microarray images. First, a new gridding algorithm is proposed for determining the individual spots and their borders. Then, a Gaussian mixture model (GMM) approach is presented for the analysis of the individual spot images. The main advantages of the proposed methodology are modeling flexibility and adaptability to the data, which are well-known strengths of GMM. The maximum likelihood and maximum a posteriori approaches are used to estimate the GMM parameters via the expectation maximization algorithm. The proposed approach has the ability to detect and compensate for artifacts that might occur in microarray images. This is accomplished by a model-based criterion that selects the number of the mixture components. We present numerical experiments with artificial and real data where we compare the proposed approach with previous ones and existing software tools for microarray image analysis and demonstrate its advantages.

Algorithms↗

Immediate-early and delayed cytokinin response genes of Arabidopsis thaliana identified by genome-wide expression profiling reveal novel cytokinin-sensitive processes and suggest cytokinin action through transcriptional cascades.

Cytokinins are hormones that regulate many developmental and physiological processes in plants. Recent work has revealed that the cytokinin signal is transduced by two-component systems to the nucleus where target genes are activated. Most of the rapid transcriptional responses are unknown. We measured immediate-early and delayed cytokinin responses through genome-wide expression profiling with the Affymetrix ATH1 full genome array (Affymetrix Inc., Santa Clara, CA, USA). Fifteen minutes after cytokinin treatment of 5-day-old Arabidopsis seedlings, 71 genes were upregulated and 11 genes were downregulated. Immediate-early cytokinin response genes include a high portion of transcriptional regulators, among them six transcription factors that had previously not been linked to cytokinin. Five plastid transcripts were rapidly regulated as well, indicating a rapid transfer of the signal to plastids or direct perception of the cytokinin signal by plastids. After 2 h of cytokinin treatment genes coding for transcriptional regulators, signaling proteins, developmental and hormonal regulators, primary and secondary metabolism, energy generation and stress reactions were over-represented. A significant number of the responding genes are known to regulate light (PHYA, PSK1, CIP8, PAT1, APRR), auxin (Aux/IAA), ethylene (ETR2, EIN3, ERFs/EREBPs), gibberellin (GAI, RGA1, GA20 oxidase), nitrate (NTR2, NIA) and sugar (STP1, SUS1) dependent processes, indicating intense crosstalk with environmental cues, other hormones and metabolites. Analysis of cytokinin-deficient 35S:AtCKX1 transgenic seedlings has revealed additional, long-lasting cytokinin-sensitive changes of transcript abundance. Comparative overlay-analysis with the software tool mapman identified previously unknown cytokinin-sensitive metabolic genes, for example in the metabolism of trehalose-6-phosphate. Taken together, we present a genome-wide view of changes in cytokinin-responsive transcript abundance of genes that might be functionally relevant for the many biological processes that are governed by cytokinins.

Arabidopsis↗

Usefulness of contrast-enhanced transabdominal ultrasonography in the diagnosis of intraductal papillary mucinous tumors of the pancreas.

BACKGROUND: The differentiation of benign from malignant intraductal papillary mucinous tumors (IPMT) is often difficult even by various examination methods. We evaluated the qualitative and quantitative diagnostic ability of contrast-enhanced transabdominal ultrasonography (CE-US), mainly in differentiating benign from malignant tumors in patients with IPMT. PATIENTS AND METHODS: There were 21 patients with IPMT who underwent CE-US and endoscopic ultrasonography (EUS). Surgery was performed in all 21 patients. Pathological findings were 4 with carcinoma and 17 with adenoma. CE-US was performed using a contrast agent (Levovist; Tanabe, Osaka, Japan) consisting of galactose microbubbles and a small (0.1%) admixture of palmitic acid, and the following items were evaluated by the following procedure. (1) Two reviewers with experienced sonographic and endosonographic ability evaluated CE-US images before and after contrast enhancement and classified the enhancement effects into three grades. In addition, the presence or absence of enhancement effects by CE-US was compared with that of mural nodules visualized by EUS. (2) In all 21 patients, changes in intensity after contrast enhancement were quantitatively measured using an HDI Lab. HDI Lab was provided by ATL (Philips; Bothell, WA) and these software tools rapidly quantify image characteristics within multiple ROI (regions of interest) and make comparisons between several areas or images. In both the early and late phases, the post-enhancement intensity, difference between pre- and post-enhancement intensity, and the percentage change ((post-enhancement value-pre-enhancement value)/pre-enhancement value) were compared between malignant and benign lesions, and the ability of CE-US to differentiate between benign and malignant lesions was evaluated in comparison with the ability of EUS to diagnose the degree of malignancy. RESULTS: (1) In both the early and the late phases, both reviewers observed enhancement effects in all 21 patients. And both reviewers observed mural nodules by EUS in all 21 patients. (2) In all 21 patients who underwent resection of IPMT, the intensity increased in both the early and late phases. When the patients with carcinoma were compared with those with adenoma, the post-enhancement intensity was significantly higher, and the difference between pre- and post-enhancement intensity and the percentage change in the early phase and the late phase was significantly more marked in the carcinoma group (p= 0.019, p= 0.002, p= 0.015, p= 0.012, and p= 0.039, respectively). CONCLUSIONS: CE-US was useful for qualitatively diagnosing tumor lesions in patients with IPMT. Moreover, quantitative changes in intensity can be a parameter for the differential diagnosis of benign and malignant tumors.

Abdomen↗

Computer-assisted mass spectrometric analysis of naturally occurring and artificially introduced cross-links in proteins and protein complexes.

A versatile software tool, VIRTUALMSLAB, is presented that can perform advanced complex virtual proteomic experiments with mass spectrometric analyses to assist in the characterization of proteins. The virtual experimental results allow rapid, flexible and convenient exploration of sample preparation strategies and are used to generate MS reference databases that can be matched with the real MS data obtained from the equivalent real experiments. Matches between virtual and acquired data reveal the identity and nature of reaction products that may lead to characterization of post-translational modification patterns, disulfide bond structures, and cross-linking in proteins or protein complexes. The most important unique feature of this program is the ability to perform multistage experiments in any user-defined order, thus allowing the researcher to vary experimental approaches that can be conducted in the laboratory. Several features of VIRTUALMSLAB are demonstrated by mapping both disulfide bonds and artificially introduced protein cross-links. It is shown that chemical cleavage at aspartate residues in the protease resistant RNase A, followed by tryptic digestion can be optimized so that the rigid protein breaks up into MALDI-MS detectable fragments, leaving the disulfide bonds intact. We also show the mapping of a number of chemically introduced cross-links in the NK1 domain of hepatocyte growth factor/scatter factor. The VIRTUALMSLAB program was used to explore the limitation and potential of mass spectrometry for cross-link studies of more complex biological assemblies, showing the value of high performance instruments such as a Fourier transform mass spectrometer. The program is freely available upon request.

Disulfides↗

Reconstruction of cardiac ventricular geometry and fiber orientation using magnetic resonance imaging.

An imaging method for the rapid reconstruction of fiber orientation throughout the cardiac ventricles is described. In this method, gradient-recalled acquisition in the steady-state (GRASS) imaging is used to measure ventricular geometry in formaldehyde-fixed hearts at high spatial resolution. Diffusion-tensor magnetic resonance imaging (DTMRI) is then used to estimate fiber orientation as the principle eigenvector of the diffusion tensor measured at each image voxel in these same hearts. DTMRI-based estimates of fiber orientation in formaldehyde-fixed tissue are shown to agree closely with those measured using histological techniques, and evidence is presented suggesting that diffusion tensor tertiary eigenvectors may specify the orientation of ventricular laminar sheets. Using a semiautomated software tool called HEARTWORKS, a set of smooth contours approximating the epicardial and endocardial boundaries in each GRASS short-axis section are estimated. These contours are then interconnected to form a volumetric model of the cardiac ventricles. DTMRI-based estimates of fiber orientation are interpolated into these volumetric models, yielding reconstructions of cardiac ventricular fiber orientation based on at least an order of magnitude more sampling points than can be obtained using manual reconstruction methods.

Animals↗

Autofluorescence removal, multiplexing, and automated analysis methods for in-vivo fluorescence imaging.

The ability to image and quantitate fluorescently labeled markers in vivo has generally been limited by autofluorescence of the tissue. Skin, in particular, has a strong autofluorescence signal, particularly when excited in the blue or green wavelengths. Fluorescence labels with emission wavelengths in the near-infrared are more amenable to deep-tissue imaging, because both scattering and autofluorescence are reduced as wavelengths are increased, but even in these spectral regions, autofluorescence can still limit sensitivity. Multispectral imaging (MSI), however, can remove the signal degradation caused by autofluorescence while adding enhanced multiplexing capabilities. While the availability of spectral "libraries" makes multispectral analysis routine for well-characterized samples, new software tools have been developed that greatly simplify the application of MSI to novel specimens.

Algorithms↗