The evaluation of Theory of Mind (ToM) in Large Language Models (LLMs) has traditionally relied on static tests like the Sally-Anne task, where the model is expected to infer false beliefs. However, recent research indicates that these tests can be gamed through memorization of patterns in training data, without the model developing deep social reasoning. To go beyond this approach, the Epistemic Asymmetry Schelling Task (EAST) emerges as a two-player dialogue game that forces LLM dyads to coordinate on semantic Schelling points under varying levels of epistemic transparency. This design measures whether models can apply ToM functionally, rather than merely answering standard benchmark questions. Results reveal a significant gap: only frontier models successfully navigate the epistemic demands, while coordination failures stem primarily from errors in tracking shared versus private knowledge.
This challenge has direct implications for enterprise AI development. If models cannot distinguish what an interlocutor knows from what the group knows, they will struggle to collaborate effectively in real-world settings. For example, in autonomous agent systems that must negotiate terms or share sensitive information, an LLM with limited ToM might leak private data or misinterpret contextual instructions. This is where custom software engineering becomes crucial: at Q2BSTUDIO we design solutions that integrate language models with epistemic control mechanisms, ensuring AI understands not only language but also the social and technical context of the organization.
The concept of “Schelling points”, borrowed from game theory, refers to obvious solutions where two parties converge without explicit communication. In EAST, LLMs must find these points under asymmetric information conditions, a key requirement for applications like process automation where different systems need to synchronize. For instance, an AI agent managing orders on AWS or Azure cloud must coordinate with another agent handling inventory, knowing which information is common and which is exclusive to each party. A failure in this distinction can lead to duplicate orders or data loss, compromising cybersecurity and operational efficiency.
From a technical perspective, EAST results suggest that functional ToM requires not only a large model but also an architecture that enables explicit tracking of epistemic states. This parallels the development of Business Intelligence (BI) and Power BI, where report reliability depends on the system distinguishing between public, private and aggregated data. At Q2BSTUDIO we combine these capabilities with AWS and Azure cloud services, offering environments where AI can operate with privacy and accuracy guarantees. Furthermore, incorporating AI agents into enterprise workflows demands that these agents recognize when to share information and when to withhold it, a skill that traditional benchmarks do not measure.
The gap between frontier models and smaller ones indicates there is still a path ahead to achieve truly collaborative artificial intelligence. Companies investing in custom applications can benefit from integrating coordination tasks similar to EAST into their quality tests, ensuring that selected LLMs not only pass static assessments but demonstrate robust understanding of social dynamics. At Q2BSTUDIO we offer consulting and development to implement these controls, from model selection to integration in cloud infrastructures, including cybersecurity of exchanged data. The future of AI lies in moving beyond the Sally-Anne test and achieving operational theory of mind, and we are ready to accompany that leap.




