Beyond the prompt: jailbreak attacks on LLMs with functions

Discover the SMT attack: exploit vulnerabilities in LLMs with functions through simulated moderation traces to bypass prompt-based defenses.

jueves, 2 de julio de 2026 • 3 min read • Q2BSTUDIO Team

SMT: vulnerability in LLMs with function calls

Security in large language models (LLMs) has been a central topic in the development of AI-based systems. Traditionally, protection efforts have focused on the prompt level, assuming that filtering malicious inputs is sufficient to prevent unwanted responses. However, recent research reveals a deeper structural vulnerability when these models operate in environments that allow the execution of external functions, such as API calls, developer-defined schemas, and third-party tools. This new attack vector, known as 'Simulated Moderation Jailbreak' (SMT), demonstrates that the attack surface expands dramatically by intertwining trusted control logic with unverified data in a shared context across multiple conversation turns. Instead of a single prompt, the adversary distributes the malicious intent through an execution trajectory that mimics a legitimate moderation audit flow, progressively weakening the model's security barriers.

This finding has direct implications for companies seeking to integrate artificial intelligence into their workflows. Many organizations build custom applications that combine language models with external tools, such as AI agents that manage inventories, automate processes, or generate reports. While these solutions offer enormous value, the underlying architecture must consider not only prompt hygiene but also the contextual validation of every element entering the model: schemas, function arguments, tool outputs, and the accumulated conversation state. Ignoring this layer exposes companies to cybersecurity risks that go beyond simple offensive responses, potentially compromising sensitive data or executing unauthorized actions on internal systems.

From a technical perspective, the SMT attack operates in black-box mode, without needing to know the model's internal weights. It simulates a moderation flow where a red teaming test request is presented as a pretext to generate harmful content. When the model refuses for security reasons, subsequent validation interprets that refusal as an execution error, forcing the attacker to refine the request until restrictions erode. This approach has demonstrated a high success rate in commercial models from five different vendors, far surpassing methods based solely on prompts. Its efficiency in terms of query count makes it a practical threat that is difficult to detect with superficial filters.

For companies developing LLM-based solutions, the lesson is clear: security must be integrated from the architectural design, not as a later layer. At Q2BSTUDIO, as a software and technology development company, we understand that protecting artificial intelligence systems requires a holistic approach. Our cybersecurity and pentesting services include risk assessments in AI agent flows, identifying attack vectors like those described, and proposing countermeasures such as conversational state validation and tool output sanitization. Additionally, when implementing custom applications with LLM capabilities, we offer solutions that integrate AI for businesses with robust access controls and continuous monitoring.

The issue presented also connects with the underlying cloud infrastructure. Many LLM implementations are deployed in environments such as AWS and Azure cloud services, where identity management, data encryption in transit, and network segmentation are critical. A jailbreak attack that manages to execute malicious functions could escalate to greater compromises without a zero-trust architecture. Therefore, at Q2BSTUDIO, we accompany our clients in adopting business intelligence and automation services with a security-by-design approach, using tools like Power BI for real-time security metric visualization, and developing custom software that audits every interaction between the model and corporate systems.

In conclusion, the evolution of jailbreak attacks reminds us that artificial intelligence cannot be treated as a magic box. Integrating LLMs with external functions opens new opportunities but also new vulnerabilities. The response lies not only in improving prompts but in rethinking the entire interaction architecture, with contextual validation, state monitoring, and principles of least privilege. Companies that adopt this vision, supported by technology partners like Q2BSTUDIO, will be better prepared to harness the potential of AI without compromising their security or operations.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.