Comparison of language models on Scrum certification questions

LLM comparison in Scrum: Which is more accurate and stable? We analyze GPT-5, Gemini 3, and DeepSeek on 993 PSM I questions. Discover their error patterns.

jueves, 2 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Accuracy, stability, and error patterns in three LLMs

Evaluating language models in technical certification environments, such as Scrum questions, reveals much more than simple right or wrong answers. When an artificial intelligence faces a formal exam, it demonstrates not only its memorization capacity but also its ability to interpret rules, discern nuances, and apply strict definitions without falling into market overinterpretations. This type of analysis is especially relevant for technology companies seeking to automate training or internal evaluation processes, where regulatory precision is critical. At Q2BSTUDIO, we understand that custom software development is not just about writing functional code, but about building systems that understand the business context and specific domain rules—something that large language models still handle unevenly.

In a recent study of three contemporary models evaluated with over 900 PSM I-style questions, notable differences in accuracy were observed depending on the topic and question format. For example, normatively explicit areas such as Artifacts or Empiricism showed homogeneous performance, while more abstract concepts like Scrum Values or Self-Managing Teams generated systematic errors. These failures were not random: patterns of overgeneralization and conflict between commercial interpretations and official definitions emerged. This finding underscores the importance of training and fine-tuning models with curated and contextualized data, something we do when implementing AI for businesses that require reliable virtual assistants in regulatory or certification environments.

Intra-model stability was high, but inter-model variability was significant. While one model excelled in overall accuracy, another showed weaknesses in advanced multiple-choice and true/false questions. This reinforces the need to choose the right architecture based on the use case. In custom application projects, we often combine different AI agents to cover complementary strengths, for example, using models specialized in structured reasoning for analytical tasks and more generalist models for natural interaction. Additionally, integration with cloud services AWS and Azure allows deploying these solutions with scalability and security, key aspects when handling sensitive data or requiring response auditing.

From a business perspective, these results offer practical guidance: one should not assume that a language model will equally master all areas of knowledge. For critical applications such as study assistants, technical support chatbots, or certification preparation tools, it is advisable to conduct domain- and format-specific evaluations. At Q2BSTUDIO, we apply similar testing methodologies when developing business intelligence services where data accuracy and interpretation of business rules are fundamental. We also explore the use of AI agents that can reason step by step, improving decision traceability.

Ultimately, comparing models on Scrum questions is not just an academic exercise: it is a thermometer of how artificial intelligence understands complex regulatory frameworks. For companies looking to innovate with AI for businesses, these insights help design more robust solutions, avoiding common biases and ensuring that technology not only responds but understands the spirit of the rules. Cybersecurity also comes into play when these systems handle responses that could become automated decisions; therefore, at Q2BSTUDIO we integrate cybersecurity practices from the design stage, protecting both training data and generated results.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.