In the world of large language model (LLM) development, traditional evaluation often focuses on measuring current performance but rarely explains why a model fails on certain tasks. Most evaluation pipelines identify weak examples or categories but leave the underlying cause implicit. This is where CRAFT comes in, an innovative method that converts any rubric-based evaluation dataset into a specific diagnosis of a model’s weak capabilities. Instead of just saying “where” it fails, CRAFT reveals the “why”.
CRAFT treats each grading criterion as a capability probe. It extracts a capability description from each prompt-rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at each node, and dynamically selects low-performing nodes across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised fine-tuning data. Experiments with four open-source models in finance and legal domains, using 13 held-out benchmarks, show that CRAFT outperforms prompt-level clustering and untargeted random generation.
From a technical and business perspective, CRAFT represents a paradigm shift. Instead of investing resources in areas that already perform well, organizations can focus their efforts on real capability gaps. This is especially relevant in regulated sectors like finance and legal, where a subtle failure can have serious consequences.
At Q2BSTUDIO, we understand the importance of precise evaluation and fine-grained diagnosis of AI models. As a software development and technology company, we offer services that integrate these advanced methodologies. Our team can implement customized AI solutions, using techniques like CRAFT to identify and correct weaknesses in language models. We combine this with custom software tailored to each client’s specific needs.
CRAFT’s ability to generate a hierarchical diagnosis allows development teams to prioritize improvements more effectively. For example, if a model shows low performance on criteria related to legal reasoning, specific training data can be generated for that capability. This process aligns perfectly with modern software engineering principles, where continuous improvement is driven by data and concrete metrics.
In the cloud context, CRAFT diagnoses can run on scalable infrastructures like AWS or Azure, enabling companies to analyze large volumes of evaluation data without worrying about computational capacity. At Q2BSTUDIO we offer cloud AWS/Azure services that facilitate such processes.
Cybersecurity also plays a crucial role. A model with weak capabilities in areas like bias detection or unsafe content generation can be a risk. CRAFT’s diagnosis helps identify these blind spots, and at Q2BSTUDIO we provide cybersecurity services to protect AI-based systems.
Another area where CRAFT has a direct impact is business intelligence (BI). Language models are increasingly used to analyze financial, legal or market data. A precise diagnosis of their capabilities enables more reliable reports. Our BI / Power BI services benefit from these improvements.
Finally, AI agents, which combine LLMs with tools and actions, require thorough quality control. CRAFT can be applied to diagnose the capabilities of each agent component, ensuring the entire system operates robustly. At Q2BSTUDIO we develop custom AI agents that integrate these practices.
In summary, CRAFT not only improves model performance but provides a framework for evidence-based continuous improvement. For companies looking to optimize their AI investments, this approach is an indispensable tool. Contact Q2BSTUDIO to learn how we can apply these techniques to your software, AI and digital transformation projects.





