What limits LLMs in cybersecurity? Study with HexStrike-AI

Study with HexStrike-AI reveals factors limiting LLMs in security. 774 tests show 72% improvement after corrections.

martes, 7 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Analysis of 774 tests reveals key factors in LLM agent effectiveness

In the realm of modern cybersecurity, agents based on large language models (LLMs) are gaining ground as orchestrators of security tools. However, a recent study with HexStrikeAI, a system that exposes over 150 tools, highlights the real limitations of these artificial intelligences. Beyond the model itself, factors such as the client driving it, tool availability, and reasoning ability determine success in complex tasks like capture the flag (CTF) challenges. The research revealed that the resolution rate increased from 55.4% to 72.0% after adjusting the orchestrator and adding new capabilities, but residual failures were mainly due to reasoning or environmental limitations, not a lack of tools. This underscores the importance of having customized cybersecurity solutions that intelligently integrate AI agents, adapting software to each organization's real needs.

For companies looking to implement these systems, the lesson is clear: a powerful model is not enough; a robust infrastructure and tailored applications that optimize workflow are required. The choice of the orchestrator client can double performance, as observed with two versions of DeepSeek. Therefore, at Q2BSTUDIO we develop custom software that integrates artificial intelligence for businesses, allowing our clients to make the most of LLMs in security environments. Additionally, the scalability offered by AWS and Azure cloud services is essential for efficiently running these workloads, while business intelligence tools like Power BI help visualize penetration testing results and security posture.

Another relevant finding is that the capability of AI agents is not limited only to available tools, but to their logical reasoning. This reinforces the need to design orchestrators that not only provide functions but also guide the model in decision-making. In this context, we offer business intelligence services to monitor agent performance and detect failure patterns. If your organization seeks to effectively implement AI for businesses, a multidisciplinary approach combining cybersecurity, cloud, and data analysis is key to overcoming current LLM limitations and moving toward truly autonomous security.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.