ViPSAM: Medical Image Segmentation with Visual Prompting

Learn how ViPSAM leverages SAM and cross-modality visual prompts to accurately segment liver lesions in non-contrast CT for proton therapy planning.

lunes, 20 de julio de 2026 • 5 min read • Q2BSTUDIO Team

Prompts visuales cruzados para imágenes médicas

Planning high-precision oncological treatments represents one of the most demanding technological challenges within today's healthcare ecosystem. In particular, proton therapy requires millimetre-level delineation of tumour volumes to maximise dosage in cancerous tissue while preserving surrounding organs. In this scenario, non-contrast computed tomography constitutes the reference modality for dosimetry, since it avoids the attenuation distortions introduced by contrast media. Nevertheless, manual segmentation of hepatic lesions and other soft tissues on these images presents added difficulty stemming from poor differentiation between tumour and anatomical background, which increases inter-observer variability and can compromise treatment plan accuracy.

Recent advances in artificial intelligence have opened new avenues for automating anatomical structure delineation, reducing planning times and standardising results. Generative models and conventional architectures have demonstrated notable capabilities in controlled environments, but their performance drops significantly when facing low-contrast radiological studies, especially those acquired under respiratory gating where diaphragmatic motion introduces artefacts and deformations. This limitation has driven the scientific community to explore hybrid strategies that transcend the information contained within a single imaging modality.

In this context an innovative proposal emerges that redefines the concept of prompting applied to foundational segmentation models. The methodology, known as ViPSAM, starts from the premise that a pretrained vision model can reorient its attention toward low-contrast regions if it receives visual cues derived from a complementary modality. Specifically, the architecture leverages contrast-enhanced magnetic resonance sequences, where lesions appear clearly differentiated, to generate guidance signals that condition the segmentation process on non-contrast computed tomography. This cross-modality knowledge transfer materialises through cross-attention mechanisms and specialised encoders that extract relevant features from the magnetic study, injecting them into the mask decoder layers without needing to retrain the base model's entire parameter set.

From a technical perspective, the approach proves particularly elegant because it preserves the computational efficiency inherent to foundational segmentation models while adapting their behaviour to a highly specialised medical domain. The introduction of a visual prompt encoder allows translating high-quality contrast information into embeddings that highlight lesion contours in the CT latent space. Simultaneously, the visual-guided cross-attention module regulates the influence of these guides, preventing morphological differences between modalities from generating segmentation hallucinations. The decoder adaptation, performed via parameter-efficient fine-tuning techniques, ensures the integration remains lightweight and scalable, two indispensable attributes for deployment in real clinical environments.

Experimental results focused on hepatic lesions for respiratory-gated proton therapy evidence substantial improvement over classical U-Net-based methods and direct implementations of universal segmenter models. The system's ability to maintain anatomical coherence in low-contrast zones demonstrates that multimodal visual prompting is not a mere academic curiosity, but a tool with transformative potential for daily radiotherapeutic practice. By reducing dependence on the medical oncologist's subjective experience during initial delineation, these solutions pave the way toward more reproducible protocols and better risk stratification.

Adopting such technologies in hospitals and proton therapy centres depends not only on the algorithm itself, but on the digital infrastructure enabling its integration into clinical workflows. This is where the development of custom software takes a leading role. Healthcare institutions require platforms connecting radiological information systems with artificial intelligence inference engines, respecting DICOM and HL7 FHIR standards. At Q2BSTUDIO, as a company specialised in software and technology development, we design bespoke solutions that encapsulate advanced segmentation models within intuitive interfaces for medical professionals, facilitating the transition from research to production environments without operational friction.

Furthermore, processing high-resolution medical imaging volumes demands a robust and elastic AWS/Azure cloud architecture. The ability to scale GPU-accelerated compute during therapeutic planning hours, combined with tiered storage strategies for image archives, constitutes a fundamental pillar for the economic viability of these projects. Migrating segmentation pipelines to cloud environments not only optimises hardware resources, but also enables deployment of AI agents capable of orchestrating preprocessing, inference and quality control tasks autonomously, freeing specialists to focus their expertise on complex clinical decision-making.

However, intensive digitisation of health data carries the inescapable obligation to protect sensitive patient information. Cybersecurity in hospital settings must extend beyond regulatory compliance, adopting a security-by-design approach encompassing encryption in transit and at rest, network segmentation and continuous threat monitoring. Any artificial intelligence solution applied to medical imaging must integrate within a trust perimeter that guarantees model integrity and data confidentiality, aspects that at Q2BSTUDIO we consider non-negotiable in every implementation.

Parallelly, exploiting clinical results generated by these systems opens extraordinary opportunities in advanced analytics. BI/Power BI platforms enable consolidation of segmentation accuracy metrics, planning times and dosimetric outcomes into interactive dashboards facilitating continuous audit of care quality. Combining computer vision models with business intelligence capabilities allows hospital executives to identify bottlenecks, compare performance across units and ground strategic decisions in real quantitative evidence.

The ViPSAM case illustrates a broader trend in contemporary medicine: the convergence between imaging modalities and foundational algorithms is redefining the boundaries of what technology can achieve without replacing the human specialist, but rather enhancing their judgement. The next generation of oncological planning systems will likely incorporate by default visual prompting mechanisms integrating magnetic resonance, computed tomography and potentially molecular data into a single shared representation space. For this vision to materialise, it will be essential to have technology partners capable of navigating the complexity of regulated medical software, large-scale data engineering and interoperability between heterogeneous systems.

In summary, robust segmentation in non-contrast computed tomography through cross-modal visual prompting represents a qualitative leap toward precision radiotherapy medicine. Organisations betting on integrating these capabilities within their digital ecosystems will not only improve patient safety, but also optimise operational processes and position themselves at the forefront of technological oncology. From Q2BSTUDIO we accompany medical centres and health sector companies in this transformation, contributing expertise in specialised software development, cloud infrastructures and data governance so that algorithmic innovation translates into tangible and sustainable clinical impact.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.