Small VLMs Know They Are Wrong but Cannot Say: Confidence Study

Small VLMs encode self-knowledge their verbal output doesn't express. Internal token probability detects errors with AUROC 0.92-0.99, but fails under severe

martes, 28 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Confianza verbal vs interna en modelos visión-lenguaje

Small vision-language models (VLMs) are increasingly deployed on consumer devices where input images suffer from compression, camera shake, or poor lighting. A recent study reveals a troubling paradox: these systems know when they are wrong but do not say it. While verbalized confidence barely varies and detects errors at chance level, internal token probability separates correct from incorrect answers with near-perfect accuracy. For a software development company like Q2BSTUDIO, which builds robust and reliable AI solutions, this gap has direct consequences for designing systems that must defer or act autonomously.

The study compares two open-weight models — Qwen2-VL-2B-Instruct and SmolVLM-Instruct — facing six realistic photographic degradations at three severity levels. It analyzes two confidence signals: the model’s natural language expression (e.g., 'I am 90% sure') and the mean probability of generated tokens. Across 3,800 predictions, the pattern is clear. In Qwen2-VL, verbalized confidence remains nearly constant (0.87-0.90) and error detection is at chance level (AUROC 0.39-0.75, typically ~0.50), while internal probability achieves AUROC 0.92-0.99. In SmolVLM, obtaining verbal confidence proved nearly impossible: out of five attempts with different prompts, only one produced a parseable value, while internal probability again outperformed chance (AUROC 0.54-0.92).

Both models fail in the same place: under severe underexposure, accuracy collapses (0.99 to 0.22 for Qwen2-VL; 0.97 to 0.42 for SmolVLM) while confidence signals barely move, and internal error detection falls to chance. This indicates that small VLMs encode self-knowledge they fail to verbalize, and that internal probability is therefore the better signal for deferral in resource-constrained deployments. But neither signal should be trusted under very low light conditions.

From a business perspective, this finding is critical. An application that trusts verbalized confidence risks acting with false certainty in degraded situations. For instance, a visual classification system in logistics might accept an underexposed image and make an error that costs the supply chain dearly. That is why at Q2BSTUDIO we design custom software that integrates internal uncertainty mechanisms, such as prediction entropy or model consistency. We also deploy these solutions on AWS and Azure cloud infrastructure to scale processing and monitor model reliability in real time.

Cybersecurity also plays a key role. A VLM that verbalizes incorrect confidence can be exploited by adversarial attacks that degrade the input to force wrong decisions. Therefore, at Q2BSTUDIO we embed cybersecurity and pentesting practices into every AI deployment, ensuring systems are not only accurate but also robust against manipulation. Additionally, business analytics with Power BI allows companies to visualize the evolution of internal model confidence and detect drifts before they affect operations.

The AI agents we build at Q2BSTUDIO embody this lesson exactly. Instead of asking the model 'Are you sure?' and expecting a textual answer, we directly extract internal probabilities and combine them with dynamic thresholds. If uncertainty exceeds a limit, the agent defers the decision to a human or triggers an additional verification process. This is especially relevant in environments like healthcare, autonomous driving, or industrial inspection, where a false positive or negative can have serious consequences.

The study also highlights the importance of validating models under real-world conditions. Many companies test their VLMs with pristine images and then are surprised by failures in production. Severe underexposure, for example, is common in low-light environments or nighttime shots. Our development team always recommends including a validation set with realistic degradations and measuring not just accuracy but the quality of confidence signals. Only then can a reliable deferral system be designed.

In summary, small VLMs know they fail but do not say it. Verbalized confidence is a mirage; the true knowledge lies in internal probabilities. For companies wanting to deploy AI responsibly, it is essential to adopt architectures that capture that hidden signal. Q2BSTUDIO is ready to help: from initial consulting to AI agent implementation, including cloud integration and cybersecurity. If your organization needs a system that knows when to stay silent and when to act, contact us. The difference between a model that gets it right and one that deceives can be as subtle as an internal probability that is never verbalized.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.