In the field of corporate and educational training, tabletop exercises (TTXs) have become essential tools for preparing teams for critical situations, such as cybersecurity incidents or operational crises. However, assessing team performance in these exercises represents a significant challenge due to their open-ended and complex nature. Traditionally, instructors spend hours or days analyzing recorded interactions and actions, delaying feedback and limiting learning. In this context, combining clustering techniques and large language models (LLMs) offers an innovative approach that not only accelerates evaluation but also provides valuable insights for continuous improvement. At Q2BSTUDIO, as a company specializing in software development and technology, we understand that digitizing these processes is key to scaling training without sacrificing quality.
Recent studies, such as the one published on arXiv with reference 2607.19209, compare two post-TTX assessment methods: clustering and LLMs. Using a dataset of 81 participants from two countries, researchers analyzed how these techniques can replace or complement instructor-assigned scores based on standardized rubrics. Clustering groups teams according to their approach to task resolution, allowing instructors to deliver faster, more specific feedback to each group. This method proved valid and reliable, with low computational requirements. On the other hand, LLMs—such as GPT-4o and GPT-5.2—evaluate team communications using the same rubrics. Although GPT-4o frequently disagreed with instructors, GPT-5.2 significantly reduced errors, approaching human-level accuracy.
This research highlights the importance of integrating technological tools into educational platforms like INJECT, an open-source TTX system that already incorporates these methods. For companies looking to implement similar solutions, developing custom software applications allows clustering algorithms and AI models to be adapted to each organization's specific needs. At Q2BSTUDIO, we design simulation platforms that record actions, communications, and decisions, providing a structured database for automated assessment. Additionally, integrating cloud AWS/Azure services ensures scalability and real-time processing of large data volumes—essential when managing multiple teams simultaneously.
Cybersecurity is another critical pillar in these exercises. TTXs often simulate security incidents, and evaluating team responses requires detailed analysis of technical and communicative decisions. Based on our experience as a technology company, we recommend that TTX platforms incorporate security modules to protect simulation data, especially when using AI models trained on sensitive information. Artificial intelligence, particularly AI agents, can play an active role in assessment: not only classifying teams but also generating personalized improvement recommendations based on behavior patterns. This aligns with current trends in Business Intelligence, where tools like Power BI allow visualization of assessment results and detection of global performance trends across teams.
The impact of these methodologies extends beyond the classroom. In corporate environments, tabletop exercises are used to train incident response teams—from IT departments to executive boards. The ability to obtain objective and rapid assessments enables companies to adjust their business continuity plans and improve cross-departmental coordination. For example, a well-implemented clustering approach can identify which teams adopt similar strategies, facilitating the creation of homogeneous training programs. Meanwhile, LLMs offer a more granular evaluation of communication—a key aspect in crisis management where clarity and precision are vital.
From a technical perspective, implementing these systems requires deep knowledge of machine learning, natural language processing, and cloud architectures. At Q2BSTUDIO, we combine these competencies to deliver comprehensive solutions. Our team develops custom software applications that integrate clustering engines and LLMs into educational platforms, ensuring instructors receive automatic reports with actionable recommendations. Furthermore, integration with cloud AWS/Azure services allows flexible deployment that scales on demand. We also incorporate AI agents that act as virtual assistants during exercises, providing hints or real-time alerts, and later participating in post-exercise evaluation.
The reliability of LLMs in team assessment remains an active research area. Study results show that newer models like GPT-5.2 approach human accuracy, but challenges persist in multilingual contexts or with technical jargon. To mitigate these risks, companies can adopt a hybrid approach: use clustering for initial categorization, then apply LLMs for detailed, instructor-supervised analysis. This strategy maximizes efficiency without sacrificing quality. At Q2BSTUDIO, we help define these workflows, adapting them to each organization's specific rubrics.
Finally, adopting these tools must be accompanied by sound data governance practices. Tabletop exercises generate sensitive information about team response capabilities, making robust cybersecurity measures essential. Cloud-hosted platforms with encryption and access control protocols are critical. Additionally, integration with BI systems like Power BI allows training managers to visualize key metrics—such as average response time, decision accuracy, or team cohesion—facilitating strategic decision-making.
In conclusion, team assessment in educational tabletop exercises is evolving thanks to artificial intelligence and data analysis. Methods like clustering and LLMs, validated in recent research, offer a path toward faster, more accurate, and personalized feedback. For companies seeking to implement these capabilities, partnering with a technology provider like Q2BSTUDIO—specializing in custom software applications, AI, cybersecurity, cloud AWS/Azure, BI/Power BI, and AI agents—is key to transforming training into a competitive advantage. The datasets and tools shared by the academic community, such as those from the INJECT study, provide a solid foundation, but customization is what makes the difference in the real world.



