Energy savings with DVFS on GPU for fine-tuning embedded SLMs

Discover how DVFS on GPU optimizes fine-tuning of SLMs in embedded devices, achieving up to 26.73% energy savings. Ideal for

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Energy optimization in embedded devices with DVFS and SLMs

Fine-tuning small language models (SLMs) on embedded GPU devices has become an imperative need for organizations seeking to customize their applications without relying on the cloud, ensuring privacy and low latency. However, this iterative optimization process of forward and backward passes over multiple mini-batches consumes significantly more energy than simple inference, posing a critical challenge on resource-constrained platforms. To address this, techniques such as Dynamic Voltage Frequency Scaling (DVFS) allow dynamically modulating the GPU voltage and frequency, reducing consumption without sacrificing the performance needed for local training. In this context, at Q2BSTUDIO we develop artificial intelligence solutions that integrate energy efficiency strategies, facilitating the sustainable deployment of SLMs in edge environments.

Characterizing energy behavior during fine-tuning reveals substantial differences between encoder-only architectures (such as BERT) and decoder-only ones (such as Pythia). While the former tend to be more stable in consumption, the latter exhibit spikes that can be mitigated with intelligent selection of DVFS configurations. A machine learning-based approach to choose the optimal voltage and frequency combination achieves average savings of over 13% and even peaks of nearly 27% compared to the unlimited power mode. These results demonstrate that it is possible to maintain model accuracy while drastically reducing the energy footprint, a key factor for companies seeking to scale their custom applications on constrained hardware.

In a landscape where enterprise artificial intelligence is moving toward decentralization, energy optimization becomes a competitive differentiator. Embedded architectures are no longer just for prototypes; they now host AI agents capable of processing sensitive data locally. To manage the lifecycle of these models, many companies combine local fine-tuning with AWS and Azure cloud services to synchronize versions or store metadata, requiring a robust hybrid architecture. At Q2BSTUDIO we offer custom software that connects these worlds, ensuring both local efficiency and cloud scalability.

Furthermore, monitoring energy consumption and SLM performance can be integrated into business intelligence platforms such as Power BI, allowing technical teams to visualize the impact of each DVFS configuration and make data-driven decisions. This analytical capability is essential for R&D departments seeking to iterate quickly without skyrocketing operational costs. Cybersecurity also plays a relevant role: by keeping fine-tuning on local devices, the attack surface is reduced, but it is vital to implement access controls and encryption on training data. Our cybersecurity services help secure these hybrid workflows.

Ultimately, the combination of DVFS with automatic configuration selection via ML opens a pragmatic path for the efficient deployment of SLMs on embedded GPUs. This approach not only extends device autonomy but also allows companies to adopt more sustainable artificial intelligence models. At Q2BSTUDIO, as a technology partner, we accompany organizations on this journey, from initial consulting to the implementation of AI for businesses with a focus on resource optimization and customization without compromising privacy.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.