RL forgets! Towards continuous policy optimization

Continuous reinforcement can suffer from catastrophic forgetting. Discover CPO, a new method that reduces forgetting in visual language models without replaying data.

martes, 7 de julio de 2026 • 1 min read • Q2BSTUDIO Team

CPO: new framework to prevent forgetting in continuous RL

In the fast-paced world of artificial intelligence, the ability of models to adapt to new tasks without forgetting what they have learned is a central challenge. Traditionally, it has been assumed that reinforcement learning (RL) suffers less from catastrophic forgetting than supervised fine-tuning, but recent research shows that this belief does not always hold. A new analysis of language and vision models (such as Qwen3-VL-8B) reveals that even with RL, forgetting can be severe when models face changing domains and updated benchmarks. To mitigate this, Continuous Policy Optimization (CPO) has been proposed, a replay-free framework that uses KL divergence-based regularization to limit model drift. This approach not only reduces forgetting by 13.7%, but also improves pre-trained capabilities by 7%, opening new possibilities for continuous adaptation in multimodal systems.

For companies integrating artificial intelligence into their processes, understanding these dynamics is key. At Q2BSTUDIO, we develop AI for businesses that need robust and durable solutions, avoiding performance degradation when faced with new tasks. Our approach ranges from creating custom AI agents to implementing tailor-made applications that scale in cloud environments. For example, we combine AWS and Azure cloud services with continuous optimization techniques to ensure models stay up-to-date without sacrificing prior knowledge. Additionally, we integrate cybersecurity and business intelligence services with tools like Power BI to offer a complete view of the AI lifecycle.

The lesson is clear: reinforcement learning is not immune to forgetting, but strategies like CPO show that it is possible to move towards more resilient models. At Q2BSTUDIO, we help companies design and implement these strategies, ensuring their AI systems not only learn but retain what they have learned. If you are looking for custom software with advanced continuous adaptation capabilities, our team is ready to transform your data into sustainable competitive advantages.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.