Evaluation of generative agents with Incognita in distributed social environments

Incognita evaluates generative agents on social tasks. Results: improvement in behavior but low reliability. Learn more.

miércoles, 8 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Testing generative agents on distributed social tasks

The evaluation of generative artificial intelligence agents has taken a qualitative leap with the emergence of environments that combine social interaction and grounded actions. Instead of testing models only on isolated tasks, new benchmarks require agents to know when to seek knowledge, whom to ask, and how to act based on information obtained from multiple sources. This approach is essential for developing AI for companies operating in real-world contexts, where knowledge is distributed among different roles and the consequences of each decision depend on precise communications and verifiable executions.

The concept of distributed social environments precisely describes this reality: the knowledge relevant to a task is distributed among participants isolated by roles, and impactful actions can only be carried out through them. Communication thus becomes an exploration mechanism over partitioned knowledge, while grounded action represents the exploitation of the environment state. This dichotomy forces agents to balance inquiry and execution, a challenge that traditional benchmarks did not address.

Incognita, a framework built on Concordia, addresses this need by separating social interaction from practical execution. In its architecture, an evaluating agent routes messages to specialized users or entities; these specialists mediate allowed operations; a deterministic sub-environment executes accepted actions on a canonical state; and an offline evaluator scores results with inherited rewards. For example, Incognita-Retail transforms the tau-bench retail benchmark into a multi-entity environment, preserving the final reward semantics. This design allows measuring behaviors such as elicitation of hidden knowledge, selection of appropriate sources, grounded writing attempts, and detection of premature terminations.

Experiments with three generative models on 18 tasks stratified by social breadth, across 540 trials, reveal significant progress but also limitations. The success rate rose from 0% to 8.9% and 17.2% in the most advanced models, while premature terminations fell from 100% to 87% and 58%. The more powerful models elicited more hidden knowledge, contacted more entities, and attempted more grounded writings, but overall reliability remains low. This confirms that distributed social environments expose emergent behaviors —such as information elicitation, source selection, and action attempts— before consistent success is achieved. For companies seeking to implement AI agents in critical processes, these results underscore the importance of evaluating not only the final outcome but also the interaction and decision-making process.

At Q2BSTUDIO, we understand that building robust multi-agent systems requires a combination of development expertise, infrastructure, and data analysis. That is why we offer artificial intelligence services that enable designing, evaluating, and deploying agents capable of operating in complex social environments. Furthermore, our ability to create custom applications and custom software ensures that each solution is tailored to the specific needs of the business, integrating communication and execution modules similar to those proposed by Incognita. To support these deployments, we have AWS and Azure cloud services that provide the necessary scalability and reliability, as well as business intelligence services based on Power BI to monitor and optimize agent performance in real time. And we do not forget security: our solutions include cybersecurity and pentesting to protect interactions and sensitive data in multi-agent environments.

Research with Incognita demonstrates that rigorous evaluation in distributed social environments is key to advancing toward more reliable and effective artificial intelligence. At Q2BSTUDIO, we help companies translate these findings into practical solutions, combining technological innovation with a focus on business value. Whether you need to develop AI for companies, implement AI agents, or strengthen your infrastructure with cloud and cybersecurity services, our team is ready to accompany you every step of the way.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.