Tabletop exercises (TTXs) have become a cornerstone of cybersecurity training in computer science education. These simulations allow student teams to face simulated incidents where they must coordinate actions, communicate effectively, and make decisions under pressure. However, evaluating their performance poses significant challenges due to the open-ended and non-linear nature of solutions. Instructors often rely on subjective rubrics or manual observations, which delays feedback and limits the scalability of training programs.
Recent research, such as that published on arXiv (2607.19209), has proposed two automated methods for evaluating teams in TTXs: clustering and large language models (LLMs). The first groups teams according to their strategies and behavior patterns, allowing instructors to provide personalized feedback to each cluster without reviewing every team individually. This approach proved valid and reliable, with low computational requirements. On the other hand, LLMs like GPT-4o and GPT-5.2 were used to score team communication based on standardized rubrics. While GPT-4o showed frequent discrepancies with instructor scores, GPT-5.2 achieved considerably lower error, moving closer to human evaluation.
These methods have been integrated into INJECT, an open-source platform for TTXs, facilitating their adoption in educational settings. Automating assessment not only speeds up the process but also provides objective data that can be analyzed with Business Intelligence tools. This is where companies like Q2BSTUDIO come into play. With experience in developing custom software and cloud solutions, Q2BSTUDIO can help build personalized TTX platforms that efficiently integrate clustering and LLMs.
Technological infrastructure is key to scaling these evaluations. Using cloud services like AWS or Azure allows handling large volumes of data generated by teams during exercises, ensuring availability and performance. Furthermore, implementing AI agents can enrich the experience: autonomous agents that monitor communication in real time, detect deviations in strategy, or even simulate roles within the scenario. Artificial intelligence not only improves assessment accuracy but also opens the door to adaptive feedback, adjusting exercise difficulty based on team performance.
Another relevant aspect is analyzing results through dashboards based on Power BI. Visualizing metrics such as response time, communication frequency, or effectiveness of decisions helps instructors identify patterns and areas for improvement at both individual and group levels. Q2BSTUDIO, with its expertise in BI and software development, can design customized dashboards that integrate TTX data and facilitate pedagogical decision-making.
The future of tabletop exercise assessment lies in combining quantitative methods (clustering) and qualitative methods (LLMs), supported by a robust cloud infrastructure and the ability of AI agents to adapt dynamically. Technology development companies like Q2BSTUDIO are in a privileged position to create turnkey solutions that not only automate assessment but transform the educational experience. Collaboration between researchers and developers is essential to bridge the gap between theory and practice, allowing students to receive immediate and relevant feedback that truly enhances their cybersecurity learning.





