MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers

MLPs store facts optimally like Hebbian networks. This post reveals a closed-form construction achieving up to 104x parameter efficiency for transformers.

martes, 28 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Construcción cerrada para almacenar hechos con eficiencia óptima

In the fast-paced advancement of artificial intelligence, large language models (LLMs) have become fundamental pillars for tasks ranging from text generation to knowledge retrieval. However, one of the deepest challenges lies in how these models store and retrieve facts efficiently. Recently, a theoretical study has revealed that MLP (multilayer perceptron) networks within Transformers can achieve optimal fact storage from an information-theoretic perspective. This finding, which has direct implications for the development of custom applications in artificial intelligence, opens new avenues for building lighter, faster, and more editable systems.

The key to this innovation is the concept of Hebbian MLPs, an architecture that leverages Hebbian learning principles — 'cells that fire together, wire together' — to store associations between keys and values optimally. Unlike previous constructions, these MLPs achieve three crucial properties: they reach optimal scaling in fact storage, handle arbitrary input and output geometries, and work inside Transformer blocks without compromising performance. This means that, for the first time, it is possible to design MLPs that store information at the same theoretical limit as a communication channel, using 10 to 104 times fewer parameters than prior alternatives for the same number of facts.

From a technical perspective, the work analyzes the decoding margin of MLPs — a metric that measures the separation between stored representations — and demonstrates that under isotropic embeddings, storage capacity reaches the Shannon limit. When embeddings are not isotropic, capacity is maintained up to certain penalty factors that depend on the geometry of the representations. This level of detail is fundamental for companies developing custom software with AI components, as it allows them to predict exactly how many parameters are needed for a given set of facts, optimizing computational resources and reducing costs.

But the impact does not stop at theory. The researchers demonstrated that these constructed MLPs can be directly integrated into Transformer blocks for factual retrieval tasks, achieving optimal scaling and requiring 15 to 63 times fewer parameters than prior constructions. Moreover, as a proof of concept, they showed that modular fact editing is possible: simply replacing a Transformer's MLP with a new one allows updating knowledge without retraining the entire model. This is immensely valuable in business environments where data changes constantly, such as in systems of AI agents or virtual assistants that require frequent information updates.

For a software development and technology company like Q2BSTUDIO, these advances represent a direct opportunity to enhance its artificial intelligence solutions. For example, when designing custom applications that integrate factual reasoning capabilities, this architecture can be adopted to minimize memory consumption and speed up inference. Furthermore, the ability to edit facts modularly fits perfectly with the cybersecurity services offered by the company, as it allows patching or updating knowledge bases in critical systems without regression risks. Also in the cloud computing domain, deploying models with Hebbian MLPs on AWS or Azure reduces infrastructure costs by requiring fewer parameters and thus less compute resources.

Another field where this technology can make a difference is Business Intelligence (BI). Power BI systems often rely on language models to generate automatic reports or answer natural language questions. With Hebbian MLPs, the factual knowledge layer would be lighter and updatable, allowing BI dashboards to reflect the most recent information in real time without reloading the entire model. This translates into a more agile user experience and more accurate data-driven decision-making.

Of course, the practical implementation of these MLPs is not without challenges. It requires deep knowledge of Transformer architecture and embedding optimization techniques. This is where companies like Q2BSTUDIO bring their expertise in software engineering and technology consulting. By offering services ranging from AI agent integration to process automation, the company can guide its clients in adopting these cutting-edge solutions, ensuring they align with business objectives and technical constraints.

In summary, Hebbian MLPs represent a conceptual leap in optimal fact storage for Transformers. By reaching the theoretical capacity limit with a fraction of the parameters, they not only improve LLM efficiency but also pave the way for more sustainable and adaptable systems. For Q2BSTUDIO, this innovation integrates seamlessly into its portfolio of AI, cloud, cybersecurity, and BI services, reinforcing its commitment to delivering cutting-edge technological solutions that maximize value for its clients. The future of knowledge storage in artificial intelligence is more efficient, modular, and above all, intelligent.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.