Cloud computing has transformed the way large-scale artificial intelligence models are deployed, especially those combining edge processing with centralized servers. This architecture, known as cloud-edge inference, allows distributing the computational load between local devices and the cloud, reducing latency and optimizing resource usage. However, this division introduces a new attack vector that has received little attention until now: the manipulation of visual tokens generated by large language and vision models (LVLMs).
A team of researchers has highlighted a critical vulnerability in this type of system. In a man-in-the-middle attack scenario, an adversary can intercept communication between the edge device and the cloud server, modifying a small fraction of the transmitted visual tokens. Although the attack is carried out under budget constraints (i.e., it can only alter a limited number of tokens), the results are alarming: by manipulating just 10% of the tokens, model accuracy can be reduced by up to 88%. These findings demonstrate that security in distributed inference cannot be taken for granted.
To understand the severity of the problem, it is necessary to consider how these models work. LVLMs process images and text jointly, generating a sequence of tokens that represent visual information. In a cloud-edge architecture, the edge device extracts these tokens and sends them to the cloud server for final processing. If an attacker intercepts and modifies those tokens, they can induce the model to make serious errors, from incorrect classifications to completely misguided responses. Attack strategies range from clever optimization-based methods to more direct approaches such as random or gradient-based substitution.
This type of threat is not isolated. In the field of cybersecurity, new attack surfaces emerge daily that require proactive solutions. Companies deploying artificial intelligence in cloud environments must consider not only the security of data at rest or in transit, but also the integrity of intermediate processes. At Q2BSTUDIO, as a software development company, we understand the importance of protecting every layer of the technological ecosystem. That is why we offer cybersecurity and pentesting services that help identify vulnerabilities before they are exploited.
The solution to this problem involves several lines of action. On one hand, it is essential to implement integrity verification mechanisms in token transmission, such as digital signatures or cryptographic checksums. On the other hand, inference architectures can be designed to tolerate some token corruption, through regularization techniques or adversarial training. Additionally, network segmentation and the use of encrypted channels are basic practices that any cloud-edge deployment should adopt. In this context, companies working with AWS and Azure cloud services must ensure that their network configurations and APIs are properly protected.
Beyond security, cloud-edge inference has enormous potential for applications such as real-time computer vision, remote assistance, or industrial automation. However, its adoption must be accompanied by maturity in development practices. At Q2BSTUDIO, we help organizations create AI solutions for businesses that are not only efficient but also robust against attacks. We work with AI agents and language models that can be integrated into secure workflows, and we also offer Power BI to visualize security monitoring data. The combination of business intelligence with cybersecurity makes it possible to detect anomalous patterns in token traffic and respond automatically.
The aforementioned study reveals a security gap affecting models ranging from 3 billion to 72 billion parameters, evaluated across four different benchmarks. This indicates that the vulnerability is transversal, regardless of the model's size or architecture. Therefore, it is not a minor problem that can be ignored. Organizations adopting cloud-edge architectures for their AI systems must include specific penetration tests for this attack vector. From custom application development to custom software implementation, it is crucial to integrate security from the design phase.
In summary, visual token manipulation in cloud-edge inference is a real and quantifiable threat. The security community and AI developers must collaborate to create more resilient models and more secure communication channels. At Q2BSTUDIO, as a technology partner, we offer consulting and development that address these challenges comprehensively, from the cloud infrastructure to the application layer, ensuring that innovation does not compromise security.

.jpg)



