Black-box inference of LLM architecture with restrictive APIs

Discover how NightVision extracts the hidden architecture of LLMs even with the most restrictive APIs. Learn about parameter inference and

viernes, 3 de julio de 2026 • 2 min read • Q2BSTUDIO Team

NightVision: extracting the hidden architecture of LLMs

In the current landscape of artificial intelligence, large language models (LLMs) have become a strategic asset for companies across all sectors. However, the internal architecture of these models is often a closely guarded secret by commercial providers. Recent research demonstrates that, even with the most restrictive APIs —those that only return a single logit per token without the possibility of bias— it is possible to estimate key parameters such as hidden dimension, depth, and total parameter count of the model. This finding opens important debates about security, intellectual property, and transparency in the AI ecosystem.

The described method, known as NightVision, employs a common prompting technique on a fixed set of tokens, combined with spectral analysis of the obtained log-probabilities. From there, using time-to-first-token (TTFT) measurements and the estimated hidden dimension, depth and parameter count can be inferred with relative errors ranging from 9% to 23% depending on the model type. This demonstrates that current API restrictions are not sufficient to completely hide architectural details.

For companies developing or deploying LLMs, this vulnerability has direct implications for the cybersecurity of their systems. If a competitor or malicious actor can reconstruct critical aspects of the architecture, it facilitates reverse engineering, unauthorized knowledge extraction, and potentially the exploitation of weaknesses. Therefore, it is essential to have protection strategies that go beyond simple API restrictions.

At Q2BSTUDIO, we understand the complexity of these challenges. As a software development company, we offer artificial intelligence solutions for businesses that integrate best security practices from the design phase. Additionally, our custom application services allow the implementation of additional layers of obfuscation and access control, reducing the attack surface against techniques like NightVision. We also develop custom AI agents that, when running in managed cloud environments, limit the exposure of internal metadata.

In parallel, monitoring performance metrics such as TTFT can be leveraged not only for malicious inference but also for legitimate optimization. With our AWS and Azure cloud services, we help companies deploy language models with configurations that minimize information leaks while maintaining high performance. And for data-driven decision-making, we offer business intelligence services with Power BI, allowing visualization of key model indicators without exposing sensitive details.

Research into inferring LLM architectures from restrictive APIs reminds us that AI security is a constantly evolving field. Companies that adopt a proactive approach —combining custom software, granular access control, and periodic audits— will be better prepared to protect their intellectual property. At Q2BSTUDIO, we accompany our clients on this path, integrating cybersecurity, artificial intelligence, and cloud services into robust solutions tailored to their needs.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.