Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “multimodal data fusion”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Demonstration of accuracy and clinical versatility of mutual information for automatic multimodality image fusion using affine and thin-plate spline warped geometric deformations.

This paper applies and evaluates an automatic mutual information-based registration algorithm across a broad spectrum of multimodal volume data sets. The algorithm requires little or no pre-processing, minimal user input and easily implements either affine, i.e. linear or thin-plate spline (TPS) warped registrations. We have evaluated the algorithm in phantom studies as well as in selected cases where few other algorithms could perform as well, if at all, to demonstrate the value of this new method. Pairs of multimodal gray-scale volume data sets were registered by iteratively changing registration parameters to maximize mutual information. Quantitative registration errors were assessed in registrations of a thorax phantom using PET/CT and in the National Library of Medicine's Visible Male using MRI T2-/T1-weighted acquisitions. Registrations of diverse clinical data sets were demonstrated including rotate-translate mapping of PET/MRI brain scans with significant missing data, full affine mapping of thoracic PET/CT and rotate-translate mapping of abdominal SPECT/CT. A five-point thin-plate spline (TPS) warped registration of thoracic PET/CT is also demonstrated. The registration algorithm converged in times ranging between 3.5 and 31 min for affine clinical registrations and 57 min for TPS warping. Mean error vector lengths for rotate-translate registrations were measured to be subvoxel in phantoms. More importantly the rotate-translate algorithm performs well even with missing data. The demonstrated clinical fusions are qualitatively excellent at all levels. We conclude that such automatic, rapid, robust algorithms significantly increase the likelihood that multimodality registrations will be routinely used to aid clinical diagnoses and post-therapeutic assessment in the near future.

Abdomen↗

[The multimodal integration, correlation and fusion of morphology and function: the methods and initial clinical applications].

Besides the pure visual diagnosis of morphology in comparison with function, digital multimodal image analysis is steadily gaining in importance. Digital image processing is used in integration, correlation and fusion of topographically identical image data. After defining these terms and describing the acquisition techniques and influence parameters, this article reviews the methods of multimodal image processing, with emphasis on correlation methods. It also gives a short description of special methods in nuclear medicine. The clinical part briefly reviews the clinical use as well as the progress achieved and the benefits expected for diagnostic applications.

Diagnostic Imaging↗

[Fusion of MRI, fMRI and intraoperative MRI data. Methods and clinical significance exemplified by neurosurgical interventions].

The aim of this work was to realize and clinically evaluate an image fusion platform for the integration of preoperative MRI and fMRI data into the intraoperative images of an interventional MRI system with a focus on neurosurgical procedures. A vertically open 0.5 T MRI scanner was equipped with a dedicated navigation system enabling the registration of additional imaging modalities (MRI, fMRI, CT) with the intraoperatively acquired data sets. These merged image data served as the basis for interventional planning and multimodal navigation. So far, the system has been used in 70 neurosurgical interventions (13 of which involved image data fusion--requiring 15 minutes extra time). The augmented navigation system is characterized by a higher frame rate and a higher image quality as compared to the system-integrated navigation based on continuously acquired (near) real time images. Patient movement and tissue shifts can be immediately detected by monitoring the morphological differences between both navigation scenes. The multimodal image fusion allowed a refined navigation planning especially for the resection of deeply seated brain lesions or pathologies close to eloquent areas. Augmented intraoperative orientation and instrument guidance improve the safety and accuracy of neurosurgical interventions.

Adult↗

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning↗

Periodic, multimodal distribution of granule volumes in mast cells.

The areas of 2327 mast cell granules in transmission electron micrographs of sections of peritoneal mast cells from adult rats were measured by digitized planimetry. A histogram constructed using equivalent volumes calculated from the measured areas assuming approximation of the granules to spheres showed a periodic multimodal distribution in which the modes fell at volumes that were successively larger integral multiples of the volume at the first mode. Application of a moving-bin technique to the data confirmed the presence of the modes. We propose a mechanism of fusion of unit sized granules to account for the multimodal distribution. The presence of pear- and dumbbell-shaped granules in mast cells is consistent with this mechanism.

Animals↗

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

Source propagation of interictal spikes in temporal lobe epilepsy. Correlations between spike dipole modelling and [18F]fluorodeoxyglucose PET data.

Source localization methods were applied to interictal spikes from scalp EEGs and correlated with metabolic (PET scan) data in eight patients suffering from drug-resistant temporal lobe epilepsy (TLE). Dipolar sources, [18F]fluorodeoxyglucose (18FDG)-PET data and anatomical images (MRI) were projected into the same three-dimensional coordinates system. Averaged spikes were adequately modelled by two or three dipolar sources with different onset time of activation but overlapping activity (mean residual variance 3.4 +/- 2.1%). Although, in all patients, spike modelling demonstrated dipolar sources in both mesial and lateral temporal cortex, dipole propagation was consistent with the early involvement of only one of these two areas (mesio-temporal, five patients; lateral and polar neocortex, three patients). Six patients showed a unilateral interictal decrease in glucose uptake, as measured with 18FDG-PET, in the temporal lobe ipsilateral to the EEG spike focus. Temporal hypometabolism was bilateral in one patient and absent in the remaining case. When projected onto PET-scan slices, the dipolar sources of these patients were always included within the hypometabolic area. However, within the hypometabolic zone, the decrease in glucose uptake was not found to be more pronounced in regions containing dipoles. Therefore the spatio-temporal spread of neuronal hyperactivity underlying interictal spiking suggests the presence of preferential epileptogenic networks inside the hypometabolic temporal lobe. Fusion of bioelectric, metabolic and anatomical data proves to be a convenient way of summarizing multimodal information from non-invasive investigations in TLE patients entering an epilepsy surgery programme, and suggests that both interictal spike dipole modelling and 18FDG-PET data might be useful, as a complement to ictal electro-clinical data, in the presurgical evaluation of such patients.

Adult↗

SPECT in the year 2000: basic principles.

OBJECTIVE: SPECT has become a routine procedure in most nuclear medicine departments. SPECT provides significant technical challenges for the nuclear medicine technologist, as compared with planar imaging, in the areas of SPECT acquisition, image reconstruction, and data processing. Many new advances in SPECT methodology are becoming available, such as iterative reconstruction, multimodality fusion, and advanced gated cardiac SPECT. SPECT imaging is demanding and requires careful attention to proper acquisition protocols, whether circular or noncircular orbits, and postprocessing is becoming more complex with the addition of iterative reconstruction and attenuation correction algorithms, among others. Understanding the principles of SPECT is essential not only to produce the highest quality scans but also to identify image artifacts. After reading this article, the nuclear medicine technologist should be able to: (a) describe the historical development and benefits of SPECT imaging; (b) state the impact of image matrix size, number of projections, and arc of rotation on final SPECT image quality; (c) discuss the trade-offs between image noise content and spatial and contrast resolution in SPECT reconstruction; (d) discuss SPECT filters and their impact on image quality; (e) explain the differences between filtered backprojection and iterative reconstruction; and (f) describe the impact of attenuation and scatter in SPECT imaging and the advantages and pitfalls of attenuation correction methods.

Humans↗

EWS::WT1 Isoform-Dependent Regulation of Neogenes in Desmoplastic Small Round Cell Tumors.

Desmoplastic small round cell tumor (DSRCT) is a rare, aggressive sarcoma characterized by the pathognomonic EWS::WT1 fusion protein (FP), an oncogenic chimeric transcription factor (OCTF) resulting from the t(11;22)(p13;q12) translocation. Recent studies have identified "neogenes" (NGs), genes normally silent in normal tissues but transcriptionally activated by OCTFs, as potential tumor-specific markers in fusion-driven cancers. In this study, we investigated the expression and regulation of DSRCT-specific NGs (DSRCT_NGs) using multimodal data across different cohorts of patients, PDX, and cell line data. We evaluated bulk and single-nucleus RNA sequencing of patient specimens from MD Anderson Cancer Center, revealing the robust ability for DSRCT_NGs to distinguish FP-positive DSRCT from samples failing detection of the EWS::WT1 FP. To elucidate the regulatory role of the EWS::WT1 FP in driving NG expression, we performed knockdown experiments in four DSRCT cell lines. This consistently resulted in a reduction of DSRCT_NG expression. Isoform-specific expression of EWS::WT1 in LP9 and MeT-5A mesothelial cells revealed that the E-KTS isoform of EWS::WT1 predominantly drives DSRCT_NG expression. Mechanistically, ATAC-seq and ChIP-seq analyses demonstrated that EWS::WT1 directly binds to accessible chromatin regions near NG transcription start sites, enriched for WT1 motifs and active histone marks. Integration of Hi-ChIP data further revealed that EWS::WT1 facilitates long-range enhancer-promoter looping at DSRCT_NG loci, promoting the expression of nearby genes. Collectively, these findings establish DSRCT_NGs as direct transcriptional outputs of the EWS::WT1 FP and implicate their loci as regulatory regions of the DSRCT transcriptome. Their fusion-dependent expression, chromatin accessibility, and promoter-enhancer connectivity underscore their potential utility as highly specific biomarkers and therapeutic targets in DSRCT.

DSRCT↗

Implementation of three-dimensional EEG brain mapping.

The electroencephalogram (EEG) visualization software was developed containing two-dimensional (2D) and three-dimensional (3D) brain mapping modules. The input to the program is standard clinical individual patient data recorded using digital EEG and magnetic resonance imaging (MRI). The software utilizes several techniques, such as heuristic triangulation, ray casting, Gouraud shading, and image fusion to form multimodal 3D images. The program has been applied to the 3D visualization of various EEG signals, "cortical" EEG signals, and potential fields generated by a computer model. The developed program appears to operate efficiently and intuitively in PC/Windows environment.

Algorithms↗

Development of a new reporter gene system--dsRed/xanthine phosphoribosyltransferase-xanthine for molecular imaging of processes behind the intact blood-brain barrier.

We report the development of a novel dual-modality fusion reporter gene system consisting of Escherichia coli xanthine phosphoribosyltransferase (XPRT) for nuclear imaging with radiolabeled xanthine and Discosoma red fluorescent protein for optical fluorescent imaging applications. The dsRed/XPRT fusion gene was successfully created and stably transduced into RG2 glioma cells, and both reporters were shown to be functional. The level of dsRed fluorescence directly correlated with XPRT enzymatic activity as measured by ribophosphorylation of [14C]-xanthine was in vitro (Ki = 0.124 +/- 0.008 vs. 0.00031 +/- 0.00005 mL/min/g in parental cell line), and [*]-xanthine octanol/water partition coefficient was 0.20 at pH = 7.4 (logP = -0.69), meeting requirements for the blood-brain barrier (BBB) penetrating tracer. In the in vivo experiment, the concentration of [14C]-xanthine in the normal brain varied from 0.20 to 0.16 + 0.05% dose/g under 0.87 + 0.24% dose/g plasma radiotracer concentration. The accumulation in vivo in the transfected flank tumor was to 2.4 +/- 0.3% dose/g, compared to 0.78 +/- 0.02% dose/g and 0.64 +/- 0.05% dose/g in the control flank tumors and intact muscle, respectively. [14C]-Xanthine appeared to be capable of specific accumulation in the transfected infiltrative brain tumor (RG2-dsRed/XPRT), which corresponded to the 585 nm fluorescent signal obtained from the adjacent cryosections. The images of endogenous gene expression with the "sensory system" have to be normalized for the transfection efficiency based on the "beacon system" image data. Such an approach requires two different "reporter genes" and two different "reporter substrates." Therefore, the novel dsRed/XPRT fusion gene can be used as a multimodality reporter system in the biological applications requiring two independent reporter genes, including the cells located behind the BBB.

Amino Acid Sequence↗

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 ± 0.0994, with a Log-rank testp-value of 1.6553×10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 ± 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 ± 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 ± 0.1211) and discrete-time survival models such as DeepHit (0.7655 ± 0.1041) and Nnet-surv (0.7694 ± 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 ± 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 ± 0.0818) and Multimodal Co-Attention Transformer (0.8102 ± 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

[Image fusion of MRI and immunoscintigraphy with MAb-170 in ovarian tumors].

In recent years multimodality imaging achieved growing importance. It is mostly performed by means of quite expensive software and hardware solutions. In the present pilot study a simple and low-cost procedure was developed to achieve image fusion in the pelvis. The image data of immunoscintigraphy (SPECT) and MRI were transferred to a personal computer and combined by standard software for image manipulation. The results in eleven patients with space-occupying lesions in the pelvis showed that adequate anatometabolic slices could be achieved. The results show a tendency to increased specificity and precision of multimodality imaging in comparison with SPECT and MRI alone. In conclusion, the low-cost solution, as developed by us, is feasible in clinical practice. Its results are reliable in clinical decision making.

Adult↗

The space of senses: impaired crossmodal interactions in a patient with Balint syndrome after bilateral parietal damage.

Balint syndrome after bilateral parietal damage involves a severe disturbance of space representation including impaired oculomotor behaviour, optic ataxia, and simultanagnosia. Binding of object features into a unique spatial representation can also be impaired. We report a patient with bilateral parietal lesions and Balint syndrome, showing severe spatial deficits in several visual tasks predominantly affecting the left hemispace. In particular, we tested whether a loss of spatial representation would affect crossmodal interactions between simultaneous visual and tactile events occurring at the same versus different locations. A tactile discrimination task, where spatially congruent or incongruent visual cues were delivered near the patient's hands, was used. Following stimulation of the left hand in the left side of space, we observed visuo-tactile interactions that were not modulated by spatially congruent conditions. In contrast, performance following stimulation of the right hand in the right side of space was affected in a spatially selective manner--facilitated for congruent stimuli and slowed for incongruent stimuli. To dissociate effects on somatotopic and spatiotopic coordinates, we crossed the patient's hands during unimodal tactile discriminations. Tactile performance of the left hand improved when it was positioned in the right hemispace, whereas placing the right hand in left space produced no significant changes, suggesting that left-sided tactile inputs are coded with respect to a combination of limb- and trunk-centred coordinates. These data converge with recent findings in animals and healthy humans to indicate a critical role of the posterior parietal cortex in multimodal spatial integration, and in the fusion of different coordinates into a unified representation of space.

Analysis of Variance↗

Minimally invasive transforaminal lumbar interbody fusion with unilateral pedicle screw fixation.

OBJECT: Posterior lumbar interbody fusion (PLIF) has been shown to be effective in the treatment of axial low-back pain. Minimally invasive spine surgery for arthrodesis has several advantages, including quicker patient recovery, less postoperative pain, and less destruction of adjacent tissue. The purpose of this paper is to evaluate the clinical outcomes after PLIF procedures in which unilateral pedicle screw fixation was used. METHODS: Prospective data were collected in 34 patients undergoing a one-level minimally invasive transforaminal lumbar interbody fusion (TLIF) in 2003. Conservative therapy, including physical therapy and aggressive multimodality pain management, had failed in all patients. Selection was based on magnetic resonance imaging studies demonstrating degenerative disc disease. All patients underwent a unilateral TLIF procedure in conjunction with posterior unilateral pedicle screw fixation. Twenty patients in whom the follow-up duration was longer than 6 months were included in this study. The follow-up duration in all patients ranged from 6 to 12 months. Seventeen (85%) of 20 patients had a good result, which was defined as a greater than 20-point reduction in the Oswestry Disability Index (ODI) score. The other three patients had no improvement. The mean preoperative ODI score of 57 improved to 25 after surgery (p < 0.005). In the 17 patients who demonstrated improvement, the mean ODI score improved from 57 to 18. The patients' visual analog scale pain scores improved from 8.3 to 1.4 (p < 0.005) after surgery. In patients who received Workers' Compensation, three (75%) of four improved. Follow-up computerized tomography scans were obtained in all 20 patients at 6 months. At that time, 13 of the patients demonstrated some degree of fusion, and no symptomatic pseudarthrosis was noted. CONCLUSIONS: Minimally invasive TLIF in conjunction with unilateral pedicle screw instrumentation is an effective treatment for axial low-back pain in appropriately selected patients.

Adult↗

Direct quantitative in vivo comparison of calcified atherosclerotic plaque on vascular MRI and CT by multimodality image registration.

PURPOSE: To investigate direct volumetric in vivo correspondence of calcified atherosclerotic plaque lesions in MRI and CT images of the thoracic aorta by multimodality image registration and fusion. MATERIALS AND METHODS: Twelve CT (11 noncontrast and one contrast) and MRI (TruFISP, contrast T1-weighted volumetric interpolated breath-hold examination (VIBE)) data sets were co-registered by approximate segmentation of the aorta and subsequent automatic co-registration by maximization of mutual information (MI). We quantitatively assessed 22 co-registered calcified plaque lesions on CT and MRI. RESULTS: The three-dimensional registration consistency and accuracy were 1.74 +/- 1.3 mm, and 2.42 +/- 1.65 mm, respectively. The ratio of CT/MRI calcified plaque volume decreased asymptotically with MRI volume, and correlated with average CT lesion density (r = 0.72) for small lesions (<25 mm(3)). The average calcified plaque volume, circumferential extent, and maximal radial width by MRI were significantly smaller compared to CT (35%, 68%, and 53%, respectively; P < 0.05). CONCLUSION: Software co-registration allowed precise, direct, and voxel-based comparison of calcified atherosclerotic plaque lesions imaged by MRI and CT. In comparison with co-registered MRI, overestimation of calcified plaque in aortic CT due to "blooming" correlates with the average lesion density for small plaques, and is greater for small plaques.

Aged↗

Automated CEAP Classification of Venous Duplex Reports Using Multimodal Artificial Intelligence.

OBJECTIVE: To develop and internally validate a prototype multimodal artificial intelligence system for automated CEAP (Clinical, Etiological, Anatomical and Pathophysiological) classification of venous duplex ultrasound (VDUS) reports, integrating natural language processing of free-text components with computer vision analysis of hand-drawn anatomical diagrams. METHODS: Single centre retrospective observational study using routinely collected clinical data. One thousand consecutive venous duplex ultrasound reports from Cambridge University Hospitals NHS Foundation Trust, UK (July 2024 - May 2025) were labelled according to the CEAP classification, excluding the Etiological component, which could not be reliably determined from duplex reports alone. Transfer learning was applied using ClinicalBERT for text and MobileNetV3 for diagrammatic data. Clinical classes were predicted from request line text. Text- and image-based pathophysiological models were developed for four anatomical territories (Great Saphenous Vein, Small Saphenous Vein, Deep system, Perforators), combined using late fusion with probability averaging. RESULTS: The clinical CEAP model achieved accuracy of 0.91, macro-F1 of 0.82, and macro-AUC of 0.98. Pathophysiological prediction varied, with text models broadly outperforming image models. Fusion yielded heterogeneous benefits, improving SSV performance but reducing Deep system accuracy. The performance of the final pathophysiological CEAP fusion models varied across anatomical territories: accuracy ranged from 0.70-0.92 and macro-AUC from 0.80-0.92. CONCLUSION: This study demonstrates the feasibility of automated CEAP classification from VDUS reports. Despite class imbalance affecting minority class predictions, the strong discriminatory performance validates this multimodal ML model for extracting clinically meaningful information from real-world data. This approach offers potential, pending external validation, to streamline vascular services through automated triage and guideline-compliant decision making.

Artificial intelligence↗