Training small language models, such as those with 4 to 7 billion parameters, presents a recurring challenge when integrated into reinforcement learning systems with tool-use capabilities. During the exploration phase, these models tend to generate erroneous trajectories due to execution failures, introducing noise into the training data. This noise causes that, when using rewards based solely on final outcomes, incorrect actions are inadvertently reinforced alongside correct ones, creating a credit assignment problem. Previous solutions, such as dense rewards or oversampling, often lead to overfitting or prohibitive computational costs. In response, the CLEANER proposal —based on the SAAR (Similarity-Aware Adaptive Rollback) mechanism— exploits the model's intrinsic self-correction ability to remove contaminated context during data collection, replacing failures with successful corrections and generating self-purifying trajectories. The resulting model internalizes correct reasoning patterns, avoiding error recovery cycles. This approach has demonstrated improvements of up to 6% on benchmarks such as AIME24/25 and GPQA, in addition to achieving state-of-the-art performance using only one-third of the usual training steps. In the current business context, optimizing AI agents for resource-constrained environments is key. At Q2BSTUDIO, as a company specialized in AI for businesses, we understand the relevance of purifying learning trajectories to build more robust models. Additionally, we offer custom applications that integrate these artificial intelligence principles with AWS and Azure cloud services, cybersecurity, and business intelligence services tools such as Power BI. Our team develops custom software that enhances automation and predictive analytics, enabling organizations to leverage more reliable and efficient AI agents without relying on massive infrastructures. Self-purification of trajectories is not just an academic innovation: it is a practical path to scaling reinforcement learning in real products, minimizing noise and maximizing reasoning quality.

.jpg)



