Propose and Attend: Confidence without Training in MLLM with Localized Attention

MTLA improves confidence in multimodal localization without training, reducing hallucinations and doubling precision in object detection.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Improved temporal localization and zero-shot object detection

In the rapid advancement of artificial intelligence, multimodal language models (MLLMs) have demonstrated an impressive ability to locate objects in images, mark temporal windows in video or audio, and even generate bounding boxes with apparent precision. However, a persistent problem in these architectures is region hallucination: the model 'invents' locations that do not correspond to the actual content. Until now, confidence mechanisms based on token probabilities have proven unreliable, as they confuse anchoring quality with input ambiguity. In response, a novel proposal has emerged that dispenses with additional training and uses localized attention between multiple tokens to measure the robustness of a prediction. This approach, called MTLA (Multi-Token Localized Attention), analyzes how the prediction tokens attend to the region they claim to describe, summing the attention within that area rather than over the entire input. The results are compelling: it improves hallucination detection by 7% to 38% across multiple MLLM families and three modalities, and when used as a confidence score to re-rank candidates, it boosts zero-shot object detection performance from 20.4 to 37.0 AP on COCO, approaching supervised detectors.

This technique has enormous practical implications for companies integrating artificial intelligence into their products, especially when robustness is required in critical applications. At Q2BSTUDIO we develop AI for businesses that needs not only precision but also transparency in its decisions. For example, when implementing a video event detection system for logistics or security, knowing when the model is hallucinating allows filtering false positives and generating reliable reports. Our experience in custom applications enables us to adapt these localized attention techniques to production environments, whether in the cloud or on edge devices.

Furthermore, we combine this type of innovation with AWS and Azure cloud services to scale AI solutions without compromising security. Cybersecurity is a pillar in our developments, especially when handling sensitive data such as surveillance videos or audio recordings. Likewise, we integrate AI agents that, by having confidence metrics like MTLA, can autonomously prioritize their actions. For business decision-making, our business intelligence services with Power BI visualize the reliability of predictions, allowing executives to adjust processes with solid data. In short, techniques like multi-token localized attention not only represent an academic advancement but also a real opportunity to build more reliable and transparent custom software, where artificial intelligence ceases to be a black box and becomes a strategic ally.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.