Different Teachers, Different Capabilities: Sub-1B On-Device Distillation

Distilling an 8B reasoning teacher into a 0.6B on-device model recovers 58% quality, runs 48x faster, and reveals per-task capability splits.

miércoles, 29 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Destilación en dispositivo: capacidades por tipo de profesor

Artificial intelligence is advancing rapidly, but one of the biggest challenges remains bringing powerful models to resource-constrained devices. Model distillation, a technique that transfers knowledge from a large 'teacher' to a smaller 'student,' promises to make this vision a reality. However, not all teachers teach the same way. A recent study shows that distilling from a reasoning model (8B parameters) to a tiny one (0.6B) not only speeds up inference — from 39 seconds to 0.8 seconds per article — but the student inherits very specific capabilities depending on the teacher type. This finding is not merely academic; it has direct implications for businesses seeking to deploy intelligent applications on mobile, IoT, or edge environments.

In the experiment, three teachers — a reasoning model (deepseek-r1:8b), a same-size instruction model, and a managed pipeline — trained students based on Qwen3-0.6B via QLoRA. The results reveal interesting splits: the reasoning teacher's student excels in summary quality, recovering 58% of the gap between base and teacher, outperforming the constrained decoding baseline by +16.8 points. In contrast, the managed pipeline's student shines in label diversity. But there is a warning: the reasoning teacher's students tend to fabricate content when the source is sparse (22 short articles out of 93), while the instruction teacher's student remains more faithful to real data.

This divergence in capabilities is a reminder that distillation is not a universal process. The choice of 'teacher' must align with the final application goal. For a company developing custom software requiring creative yet accurate summaries, a reasoning teacher may be ideal. Conversely, if data fidelity is the priority — for instance, in cybersecurity systems analyzing logs — an instruction teacher or managed pipeline offers greater reliability. Here, the expertise of Q2BSTUDIO, a specialist in AI and enterprise software, comes into play, knowing how to select the right distillation strategy for each client.

Infrastructure also matters. Running large models in the cloud incurs costs and latencies that are not always sustainable. Distillation allows moving intelligence to the device, reducing connectivity dependencies and improving privacy. Imagine a sales assistant that, running offline on a tablet, categorizes products with the same accuracy as a cloud model. Or a cloud AWS/Azure system combining local and remote inference to optimize costs. Q2BSTUDIO integrates these paradigms into its BI / Power BI solutions, where local response speed is critical for real-time dashboards.

Nevertheless, distillation does not solve all problems. The study shows that a same-size non-reasoning teacher hardly improves the student over the baseline, revealing that the gain comes from reasoning nature, not scale. This forces companies to invest in more sophisticated models as teachers, increasing training costs. The key is to balance: use a large teacher only for distillation, then deploy the student in production. Here, process automation, another Q2BSTUDIO service, can orchestrate the distillation, evaluation, and continuous deployment pipeline.

Moreover, the emergence of AI agents making autonomous decisions makes teacher selection even more critical. A cybersecurity agent that must block threats in milliseconds needs a fast and reliable student, trained with a teacher that prioritizes truthfulness over creativity. Q2BSTUDIO helps define these profiles through technical consulting and custom software development, tailoring distillation to latency, privacy, and accuracy requirements.

In short, the study confirms that there is no single teacher for all tasks. Each distillation is a trade-off between speed, response quality, and data fidelity. For businesses looking to leverage AI on devices without sacrificing performance, the solution lies in understanding these dynamics and applying a customized approach. Q2BSTUDIO, with its expertise in AI, cloud, and application development, offers precisely that: tailor-made distillation paths so that each client gets the best possible student for their specific problem.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.