OTCache: Geometric Caching with Optimal Transport for Diffusion

OTCache accelerates diffusion models up to 4.7x without training, using optimal transport to plan geometric caches. Improves fidelity and speed.

miércoles, 1 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Extreme acceleration without training: OTCache achieves 4.7x

The generation speed in diffusion models has long been a bottleneck for their adoption in enterprise environments. When a company needs to create high-quality images, videos, or synthetic content in real time, every millisecond counts. Traditionally, graph-based caching solutions sought to reduce redundant computation, but they assumed that the contributions of each step were independent of each other—an assumption that breaks down when the number of function evaluations (NFE) is extremely low. This is where OTCache is born, an innovative approach that models cache planning as a smooth evolution in the policy space, inspired by optimal transport.

OTCache operates in three distinct phases. First, it obtains a reliable reference plan using graph-based caching methods under a conservative budget, ensuring high fidelity. Second, it performs a lightweight search for an anchor point under extremely low budget conditions, employing optimization with end-to-end perceptual objectives. Finally, it predicts plans for any intermediate budget by interpolating between the quantiles of the reference policy and the anchor policy, using continuous deformation representations. This approach avoids the need to retrain the model and dynamically adapts to different levels of acceleration.

The experimental results are compelling: on models such as FLUX.1 [dev], Qwen-Image, and HunyuanVideo, OTCache achieves speedups of 4.5x, 4.7x, and 3.66x respectively, while also improving generation fidelity compared to previous caching techniques. For companies working with generative artificial intelligence, this translates into a significant reduction in inference costs and the ability to offer interactive experiences without perceptible delays. At Q2BSTudio, we understand that computational efficiency is key to integrating artificial intelligence into production processes. That is why we develop AI solutions for businesses that incorporate the latest innovations in model acceleration, ensuring optimal performance in production environments.

Beyond speed, intelligent cache planning opens the door to new architectures where AI agents can run diffusion models in real time on cloud infrastructures such as AWS or Azure. Combined with business intelligence services like Power BI, it is possible to process and visualize large volumes of synthetically generated data. Additionally, cybersecurity benefits from the ability to generate simulated environments for penetration testing without exposing real data. At Q2BSTudio, we offer custom applications and custom software to integrate these capabilities securely and scalably, always adapting to the specific needs of each client.

Research in diffusion models is advancing rapidly, and OTCache represents a paradigm shift: instead of optimizing each step separately, the entire sequence is optimized as a continuous flow. For organizations looking to make the most of the potential of generative AI, understanding and adopting these geometric caching techniques can make the difference between a pilot project and a production-ready solution. At Q2BSTudio, we are prepared to advise and implement these innovations, helping companies transform their data into real value.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.