Fine adjustment by reinforcement with checkers for thermal storage control

Optimises thermal storage in buildings with RLVR: emissions reduced by 13% compared to models without specific training.

miércoles, 15 de julio de 2026 • 4 min read • Q2BSTUDIO Team

RLVR for heat pump programming in thermal storage

Efficient energy management in commercial buildings has become a priority for both economic and environmental reasons. Thermal storage systems (TES) allow cooling demand to be shifted to lower-cost or lower-carbon footprint hours, but scheduling them optimally requires anticipating future conditions under capacity constraints. Model-based predictive control (MPC) or reinforcement learning (RL) has traditionally been employed, however, scaling these solutions to hundreds of buildings remains a computational and operational challenge. A recent study proposes a novel approach: tuning an open-source reasoning model using verifiable reward reinforcement learning (RLVR), using the action values of exact dynamic programming (DP) as a dense reward signal. This method, trained with only thirty indications, achieves emissions very close to the theoretical optimum, demonstrating that the combination of formal verifiers and reasoning models can simplify the optimization of thermal storage without the need for complex control infrastructure.

The building industry faces a dilemma: HVAC systems account for a significant portion of urban energy consumption, and their flexibility is key to integrating renewables and stabilizing the grid. However, classic approaches such as MPC require accurate dynamic models of the building and storage system, the calibration of which requires time and experts. On the other hand, pure RL needs long simulations to converge and often lacks guarantees of optimality. The aforementioned study overcomes these limitations by turning a sequential control problem into an autonomous reasoning task: the model receives a textual state with weather and price forecasts, and generates hourly setpoints for the heat pump. The dense DP-based reward allows the model to learn planning patterns such as alternative comparison, prediction horizon, and feasibility verification, without the need to explicitly program those rules.

The results are revealing. With fine-tuning by reinforcement (RFT), emissions from the open-source model were reduced from 70.5 to 61.2 kg of CO₂, very close to the optimum of PD (60.8 kg). Even generalist models such as GPT-5, without specific training, achieved near-optimal performance, while non-reasoning models such as GPT-4o generated emissions higher than those of the baseline without storage. This underscores the importance of reasoning ability in sequential decision-making. In addition, trace analysis showed that the RFT does not create an entirely new strategy, but rather stabilizes and reinforces observable patterns of planning: the model learns to look ahead, compare options, and check capacity constraints. These patterns are transferred to scenarios with forecast errors and unseen storage conditions, and even to a battery-electric task, albeit with lower gains due to the different structure of the problem.

From a business and technology perspective, this work offers a practical roadmap for implementing smart control systems in buildings without the need for cumbersome physical models or large volumes of training data. Verifiable DP-based rewards provide a reliable and scalable learning signal, allowing open-source reasoning models to be adapted to multiple buildings with just a few examples. For a software development company like Q2BSTUDIO, this opens the door to creating bespoke applications that integrate these types of decision engines directly into energy management systems. Custom software can incorporate trained AI models with verifiable rewards, deploying on cloud infrastructures such as those we offer with our AWS and Azure cloud services to ensure scalability and low latency.

The implementation of these systems not only requires the control algorithm, but also a robust cybersecurity layer to protect consumer data and communications between the building and the cloud. Q2BSTUDIO includes in its cybersecurity projects the protection of these information flows, avoiding intrusions that could compromise the operation of thermal storage. In addition, visibility into the performance of the system is achieved through business intelligence services such as Power BI, which allow facility managers to monitor in real time the emissions saved, the state of charge of the TES and the effectiveness of the model's decisions. The integration of AI agents that reason about the state of the building and adjust setpoints automatically represents a qualitative leap compared to traditional programmable thermostats.

The research also points out that reinforcement learning with verifiers can be extended to other energy storage applications, such as electric batteries or hydrogen, although with adaptations needed due to different dynamics. For companies looking to reduce their carbon footprint and optimize operational costs, combining AI for business with a reasoning-based approach to control is a promising avenue. Q2BSTUDIO, as a software and technology development company, offers the necessary capabilities to design, implement and maintain these solutions: from the creation of the AI model to its deployment in cloud environments, including cybersecurity and data visualization with Power BI. The future of building energy management lies in autonomous systems that learn efficiently with little data, and dynamically program-based verifiers provide the mathematical basis for that autonomy to be reliable and transferable.

In conclusion, reinforcement fine-tuning with verifiers represents a paradigm shift in thermal storage control: it simplifies optimization, reduces reliance on complex models, and can be generalized to different storage types. For companies in the sector, adopting this technology not only implies economic savings and reduced emissions, but also a competitive advantage in an increasingly regulated and sustainability-conscious market. Q2BSTUDIO is prepared to accompany its customers in this process, offering tailor-made applications that integrate artificial intelligence, cloud and business analytics, all with the highest cybersecurity standards. The combination of automatic reasoning and mathematical verification opens up a new era for smart energy efficiency.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.