Comparison of privacy risks according to tokenizers in federated learning

The study reveals that up to 44% of radiology reports can be reconstructed from gradients in federated learning, even with tokenizers

sábado, 18 de julio de 2026 • 3 min read • Q2BSTUDIO Team

How do tokenizers affect data leakage in radiology reports?

Federated learning has become a fundamental promise for collaboration between institutions that handle sensitive data, such as hospitals and research centers. Instead of centralizing information, this approach trains distributed models and only shares gradient updates, which in theory protects patient privacy. However, recent research shows that these gradients can be exploited to reconstruct original texts with alarming accuracy, especially when working with clinical reports. One factor that has so far received little attention is the design of the tokenizer, the component that divides text into manageable units for the model. Different tokenizers can significantly influence the amount of information that is filtered through gradients, making this technical decision a critical security aspect.

In a realistic scenario, a malicious server could modify the model's architecture before distributing it to clients, allowing it to apply gradient inversion techniques to retrieve complete sentences. Experiments controlled with GPT-2 models and various tokenizers – such as those used in GPT-2, RadBERT or LLaMA-2 – reveal that, even with large batch sizes of 256, the exact reconstruction of sentences reaches between 30% and 44%. When the batch is smaller, the accuracy exceeds 40%. This poses a huge risk to clinical data, where a single sentence may contain diagnoses, procedures, or patient identifying information. The RadBERT tokenizer, specialized in medical domains, showed the highest reconstruction fidelity, managing to recover up to 18% of a reference vocabulary with clinical terms. This indicates that, paradoxically, a tokenizer designed to improve the utility of the model can increase vulnerability.

From a business perspective, any organization that implements federated learning to process sensitive text—whether in health, finance, or legal—should consider these findings as a wake-up call. The choice of tokenizer is not just a matter of performance; it is a privacy decision that may have legal implications under regulations such as HIPAA or GDPR. Companies looking for secure solutions and complying with these standards can benefit from cybersecurity services that assess these attack vectors and propose countermeasures. In addition, the integration of techniques such as secure aggregation, homomorphic encryption or differential privacy becomes essential to mitigate information leakage. In this context, having a technological ally that understands both the model and the infrastructure is key.

Q2BSTUDIO, as a software and technology development company, offers a comprehensive approach to addressing these challenges. Our team develops AI for enterprises that incorporates protection mechanisms by design, including the careful selection of tokenizers and the implementation of security protocols at the communication layer. In addition, our custom applications allow these solutions to be tailored to the specific needs of each customer, integrating AWS and Azure cloud services to scale securely, business intelligence tools such as Power BI to monitor performance, and AI agents that operate under strict privacy controls. The combination of these capabilities ensures that organizations not only harness the potential of federated learning, but also protect their most valuable assets: data.

On the other hand, the current regulatory context requires transparency and traceability in data processing. Companies implementing AI solutions must demonstrate that they have adopted appropriate technical measures to prevent the reconstruction of personal information. In this sense, the development of process automation with artificial intelligence must be accompanied by periodic security audits and the formation of multidisciplinary teams. Tokenizers are not the only weak point: the architecture of the model, the size of the batch, the number of customers, and the frequency of updates also play a role. Therefore, we recommend a holistic analysis that includes penetration testing and simulations of gradient inversion attacks to validate the resilience of the system.

In conclusion, federated learning is still a powerful tool for collaboration without compromising privacy, but it is not inherently secure. Evidence shows that tokenizers can act as leakage channels, and that no current design eliminates them entirely. Organizations must take a multi-layered approach: from the choice of tokenizer to the implementation of cryptographic and differential privacy safeguards. With the support of experts in AWS and Azure cloud services and artificial intelligence, it is possible to build robust systems that meet the highest security standards. Q2BSTUDIO is prepared to guide companies on this path, offering tailor-made software and specialized consulting that turns risk into opportunity.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.