In the world of artificial intelligence, language models (LLMs) have become strategic assets for companies across all sectors. However, a new type of attack known as LeakyLMs has demonstrated that it is possible to steal critical information from these models simply by measuring token generation times through remote APIs. This finding, published on arXiv, reveals vulnerabilities that go beyond what was previously thought possible, affecting both the internal architecture and deployment strategies of models. In this article, we analyze the scope of this technique, its implications for business security, and how organizations can protect themselves with custom cybersecurity and AI solutions.
The LeakyLMs attack is based on a simple yet devastating principle: the time a model takes to generate each token is not constant, but depends on factors such as the number of layers, hidden dimension, number of attention heads, and the inference optimizations used. For example, researchers identified that Google Gemini Flash 2.5 uses speculative decoding with a draft context of approximately 128K tokens. This information, in the hands of competitors or malicious actors, can reveal trade secrets and competitive advantages.
From a technical perspective, the attack is divided into two main phases. The first focuses on inference optimizations and deployment strategies. By measuring response latency, it is possible to detect whether a provider uses speculative decoding, how many draft tokens it handles, and other optimizations like dynamic batching. The second phase targets the model architecture itself, allowing deduction of the number of transformer layers, hidden dimension, and number of attention heads. To achieve this, the authors built a detailed model of token generation time on modern NVIDIA GPUs, correlating latency with configuration parameters and hardware. Then, by searching the architecture space, they succeeded in having the correct configuration appear among the top ten guesses more than 90% of the time with Llama models.
For companies using proprietary or custom models, this vulnerability represents a direct threat to their intellectual property. An attacker could, with only standard API access, extract enough information to replicate or improve the model, bypassing years of R&D investment. Moreover, the attack requires no prior knowledge of the model or access to its weights, making it particularly difficult to detect. Traditional security measures, such as communication encryption or strong authentication, do not mitigate this risk because the leak occurs through a timing side channel inherent to the inference process.
How can organizations defend themselves? The answer lies in a comprehensive cybersecurity approach that includes techniques such as introducing controlled noise into response times, randomizing decoding strategies, and continuously monitoring latency patterns for anomalies. However, these solutions must be implemented without degrading user experience or model performance. This is where custom software development and specialized consulting come into play. A team like Q2BSTUDIO can design personalized protection systems that maintain AI efficiency while blocking leakage channels.
Q2BSTUDIO is a software and technology development company offering advanced solutions in multiple key areas. For instance, in cybersecurity, we perform pentesting and security audits to identify vulnerabilities like those exploited by LeakyLMs. Our specialists implement countermeasures at both the infrastructure and application levels, using AWS and Azure cloud technologies to deploy secure and scalable environments. Additionally, we integrate Business Intelligence (Power BI) systems to monitor model performance and security metrics in real time, detecting anomalies before they become incidents.
Artificial intelligence is at the core of our offering. We develop AI agents and automation systems that not only improve productivity but also incorporate protections against side-channel attacks. Our team understands transformer architectures and inference optimizations, so we can advise on best practices to hide sensitive details without sacrificing speed. Whether you need a custom model, cloud integration, or a complete security solution, at Q2BSTUDIO we combine technical knowledge with business vision to protect your most valuable asset: your intellectual property.
In conclusion, the LeakyLMs attack is not a futuristic threat but a reality already being researched and potentially used by malicious actors. Companies investing in language models must take proactive steps to harden their systems. The key is to understand that security is not an add-on but an essential component of any AI solution design. With partners like Q2BSTUDIO, it is possible to stay ahead of risks, leveraging expertise in cybersecurity, cloud, BI, and AI agents to build a robust and reliable technological ecosystem. Do not wait for your model to be leaked; act today.





