The adoption of artificial intelligence agents in production environments has opened a new front in cybersecurity. Tools like Claude Code and Codex operate over untrusted content, files, commands, and workspace state, turning any security flaw into a direct threat. Traditional red-teaming optimizes attack success rates but rarely captures the underlying conditions that enable unsafe behaviors. This is where the need for an automated approach arises, one that not only discovers vulnerabilities but makes them reusable across different models and agents.
The concept of 'reusable vulnerabilities' starts from a simple idea: instead of generating one-off attacks, build a knowledge graph that links attack surfaces with unsafe trajectories. Each node in the graph contains a hypothesis, a falsifier, an enabling condition, and transfer evidence. This allows a security team, when validating a patch, to query the graph and determine whether the same vulnerability might appear in another agent or scenario. This methodology, which we can call a 'falsifiable discovery loop', runs in a sandboxed environment, reflects on the attack trajectory, and promotes confirmed findings into a vulnerability concept.
From a business perspective, this approach transforms how organizations address the security of their AI agents. Companies like Q2BSTUDIO, specialized in custom software development, integrate these techniques into their solutions to ensure that AI-based systems are not only functional but also secure. The ability to reuse vulnerability knowledge across different models —even between Claude Code and Codex— represents significant time and resource savings in red-teaming cycles.
One of the most relevant findings of this approach is that a frozen vulnerability graph, without further search, outperforms the strongest baselines by 14.2 percentage points under the same single-shot protocol. This means that once built, the graph acts as an auditable knowledge asset for production security teams. The traceability of each concept allows inspecting how a failure occurs, validating patches, and accumulating reusable security knowledge.
Automated red-teaming is not limited to direct attacks; it also covers indirect vectors such as malicious prompt injection or agent state manipulation. For companies deploying agents in the cloud, integrating this methodology with AWS/Azure cloud services is natural. At Q2BSTUDIO we help our clients design sandbox testing environments that replicate real cloud infrastructure, allowing controlled attacks and extracting lessons that are then applied to production systems.
The cybersecurity dimension is key. An LLM agent handling sensitive data or executing commands on an operating system can become an attack vector if its operating conditions are not properly audited. The use of a vulnerability concept graph provides a solid foundation for continuous pentesting and for designing proactive defenses. Companies that already have cybersecurity solutions can extend their capabilities to agent security without starting from scratch.
Another relevant aspect is the link with business intelligence. The data generated during automated red-teaming —trajectories, falsifiers, conditions— can be integrated into BI platforms like Power BI to visualize vulnerability trends, correlate failures across model versions, and prioritize patches. Q2BSTUDIO offers Business Intelligence with Power BI services that turn this data into actionable dashboards for security teams.
Reusable vulnerability knowledge also impacts process automation. With a concept graph, teams can create scripts that automate patch validation against all known vulnerabilities, accelerating the deployment cycle of new agent versions. At Q2BSTUDIO we integrate these practices into our process automation solutions, ensuring that security does not become a bottleneck.
In short, automated red-teaming with reusable vulnerability generation represents a paradigm shift. It moves from being a reactive, one-off activity to a continuous process of knowledge accumulation. Companies that adopt this approach, supported by technology partners like Q2BSTUDIO, can deploy AI agents with greater confidence, knowing that every discovered vulnerability is recorded, analyzed, and ready to be transferred to future scenarios. The cybersecurity of LLM agents is not a destination, but a path of continuous improvement based on reusable evidence.





