Auto-DSM Under the Lens: Black-Box Evaluation for DSM with LLM

We evaluate DSM generation with LLM using a reproducible black-box framework. Discover metrics, results, and limitations.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Reference framework for evaluating LLM-generated DSMs

In the current model-based engineering ecosystem, the ability to structure complex systems from technical documentation has become a critical challenge. The automatic generation of design structure matrices (DSM) using large language models (LLM) promises to streamline this process, but its black-box nature demands robust and reproducible evaluation frameworks. A recent study proposes a novel approach that combines structural metrics, classification, and stability to measure the quality of generated DSMs against manually validated reference matrices. The results reveal that, while LLMs can produce plausible and repeatable matrices with well-defined inputs, ambiguity in dependencies and prompt formulation remains a significant source of hallucinations and abstention failures.

This type of analysis not only highlights the potential of artificial intelligence for businesses in automating engineering processes, but also underscores the need for transparent auditing tools. At Q2BSTUDIO, we understand that integrating LLMs into MBSE workflows requires robust and customized solutions. Therefore, we offer custom applications that allow adapting these systems to the specific needs of each organization, from implementing AI agents to managing cloud infrastructures with AWS and Azure.

The systematic evaluation of LLM-generated DSMs opens the door to new validation methodologies that combine metrics such as Completeness, Correctness, coupling, and entropy. For companies seeking to advance in digital transformation, having a technology partner that masters both artificial intelligence and cybersecurity or business intelligence services with Power BI is essential. At Q2BSTUDIO, we integrate these capabilities into custom software solutions that improve data traceability and quality, ensuring that each generated model is reliable and actionable.

The study's findings reinforce the importance of carefully designing prompts and establishing composite metrics such as the Composite Quality Score (Q) to synthesize overall performance. As AI for businesses matures, collaboration between domain experts and software developers becomes indispensable. Drawing on our experience in AWS and Azure cloud services and process automation, we accompany organizations in adopting these technologies, ensuring that innovation translates into tangible value without sacrificing security or precision.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.