EgoSafetyBench: evaluating VLMs as safety guards in robots

EgoSafetyBench: benchmark of 1,200 videos to evaluate VLMs as safety guards. Do they detect real dangers or are they deceived?

jueves, 2 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Diagnosis of VLMs as robotic safety guardians

The emergence of autonomous robots in industrial and domestic environments poses a critical challenge: how to ensure they act safely without generating false alarms that paralyze production processes. Vision-language models (VLMs) have been proposed as real-time safety guardians, but traditional benchmarks, by categorizing scenes as safe or dangerous in a binary manner, fail to capture the complexity of ambiguous situations. This is where EgoSafetyBench comes in, an egocentric dataset with 1,200 videos annotated at half-second intervals, designed to evaluate these systems' ability to identify truly critical moments without reacting to deceptive appearances.

The benchmark is divided into two tracks. The first covers 800 situational scenarios ranging from safe routines to contextual dangers, including suspicious but harmless actions. The second, with 400 scenarios, focuses on misleading visual cues (signs, labels, warnings) that can distort the perception of physical safety. Each scenario includes a contrastive pair where only one decisive visible detail changes, forcing the model to rely on that clue rather than the global context. The results reveal that, although VLMs correctly detect dangerous videos, they fail to identify the exact moment of risk, especially in contextual dangers. More seriously: misleading cues degrade the performance of all tested models; some overlook up to a third of the dangers, while others intervene excessively in safe scenes, confusing apparent robustness with a tendency to alarm indiscriminately.

This finding underscores the need for more sophisticated approaches to autonomous supervision. At Q2BSTUDIO, we understand that robotic safety cannot depend on binary classifications or models that react to visual noise. That is why we offer AI for businesses that combines computer vision, natural language processing, and contextual reasoning, integrated into custom applications deployed in real industrial environments. Our AI agents are trained to discriminate between genuine and deceptive signals, using AWS and Azure cloud services to scale real-time processing. Furthermore, the cybersecurity of these systems is reinforced through continuous audits, and performance data is visualized with Power BI to provide dashboards that allow human supervisors to identify patterns of false positives or negatives.

The path toward truly safe robots goes through benchmarks like EgoSafetyBench, but also through custom software development that incorporates robust decision logic, adversarial training, and the ability to explain its decisions. Our business intelligence services help companies measure the real impact of these systems, adjusting intervention thresholds according to the operational context. Robotic safety is not a binary problem; it is an engineering challenge that requires a combination of vision, language, reasoning, and cloud deployment. At Q2BSTUDIO, we work so that artificial intelligence not only detects dangers but understands when to intervene and when to let the robot continue its task.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.