EvoCUA-1.5: Online RL for Multi-turn Computer-Use Agents

EvoCUA-1.5 achieves 63.2% on OSWorld-Verified via online RL for multi-turn computer agents. Step-level optimization and dynamic curriculum boost performance.

martes, 28 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Aprendizaje por refuerzo online multiturno con EvoCUA-1.5

The advancement of artificial intelligence agents capable of autonomously interacting with computer environments has taken a qualitative leap with EvoCUA-1.5, an online reinforcement learning system designed specifically for multi-turn tasks on virtual desktops. Unlike previous approaches that relied on static trajectory imitation, this model introduces a continuous improvement cycle where the agent learns directly from interaction with executable environments, receiving verifiable rewards upon completing objectives. This paradigm not only overcomes the limitations of offline data but also poses unique technical challenges: observations that change with each step, sparse rewards, and delays in environment feedback. To address these, EvoCUA-1.5 incorporates three key innovations: step-level policy optimization (STEPO), policy-aware filtering with pass-rate calibration, and a dynamic tri-adaptive curriculum (DTAC). All of this is supported by a fully asynchronous infrastructure that manages latency and data staleness.

From a business perspective, understanding how these mechanisms work is crucial for any company seeking to automate complex processes — such as system management, technical support, or navigating legacy applications — through intelligent agents. At Q2BSTUDIO, we develop custom AI solutions that integrate reinforcement learning techniques for dynamic environments. The ability of an agent to operate over multiple steps, remember context, and recover from errors is exactly what distinguishes a useful assistant from a fragile one. Our team applies principles similar to those of EvoCUA-1.5 when designing automated workflows that require interaction with multiple windows, web forms, or administration consoles.

The first component, STEPO (Step-Level Policy Optimization), solves a fundamental problem in multi-turn reinforcement learning: how to decompose a long trajectory into individual steps without losing the notion of accumulated advantage. Instead of treating each episode as a monolithic block, STEPO assigns to each step an advantage calculated from the value of future states, preserving the balance between exploration and exploitation. This allows the agent to learn from every action, even when the final reward is delayed. For a company implementing a customer service agent, this translates into gradual improvement of responses without constant supervision.

The second component is a filtering and calibration mechanism over verifiable synthetic tasks. The system generates thousands of test scenarios — such as opening a file, changing a setting, or extracting data from an application — and retains only those trajectories that exceed a success probability threshold. Subsequently, the success rate is calibrated to prevent the agent from over-specializing in easy tasks. This strategy is analogous to regression testing in custom software development: each functionality is verified in isolation before being integrated into the main flow. In our custom software work, we apply a similar approach when validating agent behaviors in controlled environments before deploying them to production.

The third pillar is DTAC (Dynamic Tri-Adaptive Curriculum), an adaptive curriculum that combines three strategies: learnable tasks (those where the agent can still improve), positive replay of difficult tasks (to consolidate skills), and controlled exposure to infeasible tasks (to prevent the agent from wasting resources on impossible objectives). This dynamic balance resembles how engineers at cloud on AWS and Azure design CI/CD pipelines: prioritizing changes that provide the most value without saturating the system. For a cybersecurity agent, for example, DTAC allows it to first face moderate simulated threats before being exposed to real attacks, improving robustness without compromising environment security.

The infrastructure supporting the entire process is an asynchronous architecture with staleness control and mini-group batching. In a business environment, this is equivalent to having a distributed system that handles multiple agent instances in parallel, collecting experiences and updating the policy without bottlenecks. Companies integrating Business Intelligence with Power BI or process automation solutions need precisely this scalability: the ability to launch hundreds of simultaneous tests, process results in real time, and adjust models without disrupting service.

Experimental results for EvoCUA-1.5 are compelling: it achieves a 63.2% success rate on the OSWorld-Verified benchmark, outperforming open-weight models with 32B-35B parameters and approaching much larger models. This demonstrates that online learning efficiency can compensate for model size, a relevant finding for companies seeking to deploy agents without incurring exorbitant computational costs. At Q2BSTUDIO, we believe the key lies not only in parameter volume but in the quality of the training process and the ability to adapt to changing environments.

Finally, it’s worth noting that such technologies are not limited to virtual desktops. Any system requiring sequential interactions — from a CRM to an industrial control panel — can benefit from EvoCUA-1.5 principles. The combination of online reinforcement learning with adaptive curricula opens the door to assistants that truly learn from experience, improving with each completed task. Our experience in cybersecurity and pentesting has taught us that AI agents must be trained in realistic conditions to be effective; this multi-turn approach is the right path.

In conclusion, EvoCUA-1.5 represents a milestone in reinforcement learning for multi-turn agents, offering a practical framework that combines theory with computational efficiency. For companies wishing to adopt these capabilities, having a technology partner who understands both the algorithmic side and integration into cloud and on-premise infrastructures is essential. At Q2BSTUDIO, we offer custom software development, artificial intelligence, cybersecurity, cloud computing, and business intelligence services to help organizations build agents that not only execute tasks but learn and adapt as a human collaborator would.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.