CLEANER: Self-purifying trajectories improve reinforcement learning

CLEANER purifies erroneous trajectories in small models (4B-7B) to improve RL. Achieves up to 6% more accuracy with fewer training steps.

martes, 7 de julio de 2026 • 1 min read • Q2BSTUDIO Team

Improve accuracy of 4B-7B models with CLEANER

Training small language models, such as those with 4 to 7 billion parameters, presents a recurring challenge when integrated into reinforcement learning systems with tool-use capabilities. During the exploration phase, these models tend to generate erroneous trajectories due to execution failures, introducing noise into the training data. This noise causes that, when using rewards based solely on final outcomes, incorrect actions are inadvertently reinforced alongside correct ones, creating a credit assignment problem. Previous solutions, such as dense rewards or oversampling, often lead to overfitting or prohibitive computational costs. In response, the CLEANER proposal —based on the SAAR (Similarity-Aware Adaptive Rollback) mechanism— exploits the model's intrinsic self-correction ability to remove contaminated context during data collection, replacing failures with successful corrections and generating self-purifying trajectories. The resulting model internalizes correct reasoning patterns, avoiding error recovery cycles. This approach has demonstrated improvements of up to 6% on benchmarks such as AIME24/25 and GPQA, in addition to achieving state-of-the-art performance using only one-third of the usual training steps. In the current business context, optimizing AI agents for resource-constrained environments is key. At Q2BSTUDIO, as a company specialized in AI for businesses, we understand the relevance of purifying learning trajectories to build more robust models. Additionally, we offer custom applications that integrate these artificial intelligence principles with AWS and Azure cloud services, cybersecurity, and business intelligence services tools such as Power BI. Our team develops custom software that enhances automation and predictive analytics, enabling organizations to leverage more reliable and efficient AI agents without relying on massive infrastructures. Self-purification of trajectories is not just an academic innovation: it is a practical path to scaling reinforcement learning in real products, minimizing noise and maximizing reasoning quality.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.